Claude Opus 5 Quadruples the Toughest AI Reasoning Score, and Does It Cheaper Than Rivals

The Core · TL;DR
- Claude Opus 5 scored 30.2% on ARC-AGI-3, nearly quadrupling the prior record of 7.8% held by GPT-5.6 Sol.
- The model solved five previously unsolved ARC-AGI-3 environments, four at or above human level, and topped the Artificial Analysis Intelligence Index at 61.
- Opus 5 undercuts rivals on cost, averaging $2.03 per Intelligence Index task versus $2.75 for Fable 5 with fallback.
- Results are mixed elsewhere: Opus 5 trails Fable 5 on factual accuracy (AA-Omniscience) and its hallucination rate rose 14 points to 50%.
A score of 30.2 percent might not sound impressive, until you learn the previous best was 7.8 percent. That leap, achieved by Anthropic's newly released Claude Opus 5 on the ARC-AGI-3 benchmark, is the standout figure in a wave of independent testing that places the model at or near the top of nearly every major reasoning evaluation currently in use.
ARC-AGI-3 was built specifically to resist the kind of pattern-matching shortcuts language models often exploit, using interactive puzzle environments designed to probe genuine problem-solving rather than memorized skills. Opus 5 solved five environments that no model had cracked before, four of them at a level matching or exceeding human performance, and pushed the total number of solved public demo environments to six out of 25.
Researchers observed the model independently deriving reflection equations and converting puzzle logic into algebraic notation while working through ARC-AGI-3 tasks, a problem-solving approach not previously documented in these evaluations. Anthropic's earlier Fable-class Claude models had scored only around 20 percent on the same benchmark, according to ARC Prize, underscoring how large a jump this generation represents.
The gains extend beyond ARC-AGI-3. Opus 5 posted 90.4 percent on ARC-AGI-2 and 97.5 percent on ARC-AGI-1, and topped the Artificial Analysis Intelligence Index with a score of 61, edging out Fable 5 (60), GPT-5.6 Sol (59), Kimi K3 (57), and its own predecessor Opus 4.8 (56).
Not every benchmark tilts in Anthropic's favor. On CritPt, a physics reasoning test, Opus 5 matches Fable 5 but falls behind GPT-5.6 Sol, GPT-5.5 Pro, and GPT-5.6 Terra. On AA-Omniscience, a factual accuracy measure, it improved seven points over Opus 4.8 yet still trails Fable 5, and its hallucination rate on that same test climbed 14 points to 50 percent. On Epoch AI's Capability Index, Opus 5 scored 159 against Fable 5's 161, a near-tie rather than a win.
Pricing undercuts the competition
Perhaps the more consequential detail for developers is cost. At the "high" and "xhigh" reasoning tiers, Opus 5 outperforms both Opus 4.8 and Sonnet 5 while charging less to run. Artificial Analysis found the average Intelligence Index task costs $2.03 with Opus 5, compared to $2.75 for Fable 5 when fallback routing is included.
On Terminal-Bench v2.1, Opus 5 at maximum reasoning tied the prior leader GPT-5.6 Sol at 89 percent, and it matched Fable 5's score of 161 on the SWE-ECI software engineering benchmark, though GPT-5.6 Sol still leads both categories outright.
One caveat worth flagging: Opus 5 was trained after ARC-AGI-3's puzzle formats became public, raising the possibility that Anthropic could have optimized specifically for the benchmark's structure rather than for general reasoning gains alone. That doesn't erase the results, but it's a factor independent evaluators will likely scrutinize as more third-party testing accumulates.
Original reporting and research used to synthesize this article.
WAKIB Editorial Team
This review was prepared and summarized by the WAKIB AI intelligence engine and vetted by our editorial board for accuracy and reliability.
Subscribe to Newsletter
Get a weekly summary of the most promising AI research and tools delivered to your inbox.
Telegram Channel
Join our active community on Telegram for real-time tracking of AI models and trends.
