Claude Opus 5 Scores 61 on Intelligence Index, Beats Fable 5 on Cost Across Most Benchmarks
Anthropic's Claude Opus 5 posts a 61 on the Artificial Analysis Intelligence Index, narrowly beating Claude Fable 5 (60) and GPT-5.6 Sol (59) while costing less per task. The model leads in coding and knowledge-work benchmarks but shows a rising hallucination rate of 50 percent.
Anthropic's Claude Opus 5 has posted the highest score on the Artificial Analysis Intelligence Index among current frontier models, hitting 61 points while undercutting rival Claude Fable 5 on cost across most reasoning tiers, according to benchmark data from Artificial Analysis, Epoch AI, and Vals.ai.
Intelligence Index and coding results
Opus 5 scored 61 on the Artificial Analysis Intelligence Index, a composite of nine tests spanning knowledge work, coding, scientific reasoning, and factual accuracy. That places it just ahead of Claude Fable 5 (60), GPT-5.6 Sol (59), Kimi K3 (57), and Anthropic's own Claude Opus 4.8 (56). Artificial Analysis says it worked with Anthropic to test the model ahead of its public release.
On coding tasks, Opus 5 paired with Claude Code shares the top spot on the Artificial Analysis Coding Index at the "xhigh" reasoning tier. On Terminal-Bench v2.1, which evaluates autonomous agents in real terminal environments, Opus 5 scored 89 percent at the "max" tier, matching the previous leader GPT-5.6 Sol. On the Coding Agent Index, Opus 5 with Claude Code ties GPT-5.6 Sol with Codex at 67 points.
Epoch AI's independent testing puts Opus 5 slightly behind Fable 5 overall — 159 versus 161 on its Capability Index — but the two models tie at 161 on the software engineering-specific SWE-ECI benchmark, ahead of GPT-5.6 Terra and Opus 4.8. GPT-5.6 Sol leads both Epoch AI categories.
Accuracy trade-offs
Opus 5's factual reliability shows mixed results. On Humanity's Last Exam, a difficult cross-disciplinary knowledge test, it scored 53 percent, tying Fable 5. On the CritPt physics benchmark developed with Argonne National Laboratory and UIUC, it again matches Fable 5 but trails GPT-5.6 Sol, GPT-5.5 Pro, and GPT-5.6 Terra.
On AA-Omniscience, which measures the accuracy of a model's knowledge claims, Opus 5 improved 7 points over Opus 4.8 but still trails Fable 5. According to Artificial Analysis, the model's tendency to answer more often under uncertainty pushed its hallucination rate up 14 points to 50 percent — a figure that raises reliability concerns for high-stakes use cases.
Pricing and reasoning tiers
Anthropic prices Opus 5 at $5 per million input tokens and $25 per million output tokens. Cache writes cost $6.25 per million tokens with a five-minute lifetime; cache hits cost $0.50 per million tokens.
The average Intelligence Index task costs $2.03 with Opus 5 — cheaper than Fable 5's $2.75 but more expensive than Opus 4.8 ($1.80) and Sonnet 5 ($1.53). At the "high" and "xhigh" reasoning tiers specifically, Opus 5 outperforms both Opus 4.8 and Sonnet 5 while costing less.
Vals.ai testing across all five reasoning tiers (low, medium, high, xhigh, max) using Vibe Code Bench found scores climbing from 76.7 percent (low) to 82 percent (medium) and 89.8 percent (high), before dipping to 88.3 percent (xhigh) and 88.4 percent (max) despite higher cost. Vals.ai attributes this to higher tiers producing more complex but more error-prone solutions. Anthropic has set "high" as the default tier in both its API and Claude Code, a choice the data appears to support.
Knowledge work and office tasks
On AA-Briefcase, which scores AI performance on office tasks like report writing and spreadsheet analysis using an Elo-style rating, Opus 5 at "max" reaches 1720 — 146 points ahead of Fable 5 (1574). Its top three reasoning tiers sweep the top three AA-Briefcase rankings. Cost per task at "max" is $17.79, a 20 percent drop from Fable 5's $22.30; the "high" tier costs just $10.41.
On GDPval-AA v2, Opus 5 at "max" reaches an Elo of 1861, far above the human baseline of 1000. Its Analytical Quality Elo hits 2016 at "max," nearly 300 points ahead of Fable 5. Presentation quality lags, however, with Opus 5 scoring 1628 versus GPT-5.6 Sol's 1666.
The gains come with a time cost: at "max," Opus 5 needs over 36 minutes per task and averages 103 passes, roughly 50 percent more than Opus 4.8's 24 minutes and 55 passes.
What this means
The benchmark spread between Opus 5, Fable 5, and GPT-5.6 Sol is narrow — a few points on most indices — which supports the argument that frontier model capability is converging rather than any single lab pulling ahead. Anthropic's real edge here isn't raw intelligence score but price-to-performance: Opus 5 matches or beats larger rivals at the "high" tier while costing less than half as much as Fable 5 on knowledge-work tasks. The rising hallucination rate (50 percent) is the clearest liability, suggesting Anthropic traded some caution for higher coverage on uncertain questions — a trade worth watching closely in production deployments where wrong-but-confident answers carry real cost.
Related Articles
OpenAI Says GPT-5.6 Sol Scores 38.3% on ARC-AGI-3, Beating Claude Opus 5 — But Only With a Custom Harness
OpenAI says GPT-5.6 Sol scores 38.3 percent on ARC-AGI-3 using a custom Responses API configuration, edging out Claude Opus 5's 30.2 percent. Under the official benchmark harness, GPT-5.6 Sol's score drops to 7.8 percent, raising questions about what the comparison actually measures.
Anthropic Adds Cross-Session Messaging to Claude Code v2.1.224
Claude Code v2.1.224 introduces cross-session messaging, letting separate Claude Code instances on macOS and Linux send each other summaries to coordinate work. The feature does not support approving permissions or executing commands remotely.
Anthropic Sets Claude Code Auto Mode as Default Starting August 14
Anthropic will switch Claude Code's default permission setting to auto mode on August 14 for Pro, Max, and Team users. The company says its safety classifier caught 89% of dangerous commands in testing, compared to 13.6% for human reviewers, and will no longer charge extra tokens for the classifier itself.
Composio Benchmark: Claude Code Fastest Agent Framework, But Costs Nearly 3x More Than OpenCode
Composio benchmarked DeepSeek V4 Flash across four agent frameworks—Claude Code, Codex, OpenCode, and Oh My Pi—on 30 real-world tasks. Claude Code finished fastest at 122 seconds per task but cost $0.195, nearly three times OpenCode's $0.073, while Oh My Pi had the highest success rate at 17/30 but took 272 seconds per task.
Comments
Loading...