Claude Opus 5 benchmarks show cheaper gains over Fable 5
Artificial Analysis and Epoch AI put Claude Opus 5 near the top of frontier AI tests, with lower task costs than Fable 5 in key tiers.
By Wei-Lin Zhao · AI Correspondent
· 3 min read
Anthropic’s Claude Opus 5 benchmarks put the new model at or near the front of current frontier AI systems, according to tests from Artificial Analysis and Epoch AI. The practical point for AI buyers is cost performance: Artificial Analysis found Opus 5 beating or matching Fable 5 across several benchmark areas while costing less on some high-value task settings, though reliability remains uneven.
Artificial Analysis gave Claude Opus 5 a score of 61 on its Intelligence Index, which combines nine evaluations across knowledge work, coding, scientific reasoning and factual accuracy. That placed it ahead of Claude Fable 5 at 60, GPT-5.6 Sol at 59, Kimi K3 at 57 and Claude Opus 4.8 at 56. Artificial Analysis said it worked with Anthropic to test the model before public release, a useful disclosure for readers comparing vendor-linked benchmark results.
The model’s pricing did not change from the stated Opus line: $5 per million input tokens and $25 per million output tokens. Cache writes cost $6.25 per million tokens for a five-minute lifetime, while cache hits cost $0.50 per million tokens.
How does Claude Opus 5 compare with Fable 5?
On Artificial Analysis’ average Intelligence Index task, Opus 5 cost $2.03, below Fable 5 with fallback at $2.75 but above Opus 4.8 at $1.80 and Sonnet 5 at $1.53. At the “high” and “xhigh” settings, Artificial Analysis found Opus 5 outperforming both Opus 4.8 and Sonnet 5 while costing less.
In coding, Artificial Analysis said Opus 5 at “xhigh” with Claude Code tied for first place on its Coding Index. On Terminal-Bench v2.1, which tests agent performance in real terminal environments, Opus 5 scored 89 percent at “max,” matching GPT-5.6 Sol.
Vals.ai reported that higher reasoning settings were not necessarily better for coding. On Vibe Code Bench, Opus 5 rose from 76.7 percent at “low” to 82 percent at “medium” and 89.8 percent at “high,” then dipped to 88.3 percent at “xhigh” and 88.4 percent at “max,” despite higher costs. Vals.ai attributed that pattern to more complex solutions that more often contained errors. Terminal-Bench 2.1 showed a similar pattern, with “high” beating “max” because top-tier runs spent more time per attempt and completed fewer attempts within the time limit.
Where Opus 5 looks strongest
Artificial Analysis reported Opus 5’s clearest advantage in knowledge work. On AA-Briefcase, a benchmark covering tasks such as research reports, presentations and spreadsheet analysis from large sets of input files, Opus 5 reached an Elo score of 1720 at max reasoning. That was 146 points ahead of Fable 5 at 1574. Its max, xhigh and high tiers took the top three positions in that benchmark, according to Artificial Analysis.
The cost gap was also visible there. Artificial Analysis said Opus 5 cost $17.79 per task at “max,” down from $22.30 for Fable 5. Opus 5 cost $14.26 at “xhigh” and $10.41 at “high,” with both settings still ranking above Fable 5 on Elo.
Opus 5’s gains were concentrated in analytical quality. Artificial Analysis said the model reached an Analytical Quality Elo of 2016 at “max,” nearly 300 points above Fable 5. Its Rubric Pass Rate was 58 percent at “max,” 57.2 percent at “xhigh” and 56 percent at “high.” Presentation quality was less favorable: Opus 5 scored 1628, behind GPT-5.6 Sol at 1666.
Reliability remains a problem
Artificial Analysis also found a tradeoff in factual accuracy. Opus 5 improved by 7 points over Opus 4.8 on AA-Omniscience, but still trailed Fable 5. Because it answered more often when uncertain, its hallucination rate rose 14 points to 50 percent, a material issue for legal, finance, medical and other high-stakes deployments.
Epoch AI’s results were closer. It gave Claude Opus 5 an overall Epoch Capability Index score of 159, just below Fable 5 at 161, while GPT-5.6 Sol led. On Epoch’s software engineering subset, Opus 5 tied Fable 5 at 161 and beat GPT-5.6 Terra and Claude Opus 4.8. The combined picture is a tight frontier market, with no single model holding a broad, uncontested lead across capability, price and reliability.
This story draws on original reporting from The Decoder.