OpenAI's GPT-5.6 Sol "Beats" Claude Opus 5 on ARC-AGI-3, But Only With Its Own API Tweaks

The Core · TL;DR
- GPT-5.6 Sol scores only 7.8% on ARC-AGI-3's official harness but 38.3% when run through OpenAI's own API with 'Retained Reasoning' and 'Compaction' settings enabled.
- Anthropic's Claude Opus 5 quadrupled the prior ARC-AGI-3 record with a 30.2% score under standard test conditions.
- ARC Prize co-founder François Chollet says benchmark-gaming harnesses are off limits, but general-purpose API features not built specifically for the test are fair game.
- Chollet confirmed ARC Prize had direct discussions with OpenAI over how to handle the 'Compaction' feature during testing.
A benchmark score can mean two very different things depending on how it was produced, and GPT-5.6 Sol just illustrated that gap in dramatic fashion. Under the standard ARC-AGI-3 test harness, OpenAI's model scores just 7.8 percent, well behind Anthropic's Claude Opus 5, which quadrupled the previous record with a 30.2 percent result on the same benchmark.
Run through OpenAI's own Responses API instead, using two additional settings called "Retained Reasoning" and "Compaction," GPT-5.6 Sol jumps to 38.3 percent, edging past Opus 5's headline number.
That's not a rounding difference. It's the same underlying model producing a fivefold swing in score purely based on how its outputs are managed between steps.
Why the Numbers Diverge
The official ARC-AGI-3 harness discards a model's reasoning after each action it takes, forcing it to essentially start fresh each turn. OpenAI's alternative setup changes that dynamic in two ways.
"Retained Reasoning" preserves the chain of thought across steps rather than wiping it, letting the model build on its own prior thinking. "Compaction" then prevents that accumulated context from being truncated by summarizing older material instead of dropping it outright.
Together, those features give GPT-5.6 Sol persistent memory of its own reasoning throughout a task, something the standard harness deliberately withholds to keep the test focused on raw model capability rather than clever scaffolding.
Where ARC Prize Draws the Line
ARC-AGI-3 exists specifically to measure a model on its own terms, without external tooling or vendor-specific configurations propping up the result. That principle is exactly what makes OpenAI's higher score contentious.
ARC Prize co-founder François Chollet has drawn a distinction between two categories of setup: harnesses built specifically to game the benchmark, which he considers out of bounds, and general-purpose API options that exist independently of ARC-AGI-3, which he treats as legitimate.
Chollet said ARC Prize had "back and forth with OpenAI about how to best test their models, especially with regard to compaction."
That disclosure suggests the benchmark's stewards were aware of, and at least partly negotiated, how OpenAI's result would be framed before it went public.
The episode leaves both scores on the table rather than resolving into a single verdict. Opus 5's 30.2 percent stands as the result under ARC-AGI-3's intended constraints, while GPT-5.6 Sol's 38.3 percent reflects performance under conditions the model wasn't originally evaluated against by the benchmark's default rules. For engineers evaluating either model, the practical lesson is that "state of the art" claims increasingly hinge on which API settings were switched on, not just which weights were running underneath.
Original reporting and research used to synthesize this article.
WAKIB Editorial Team
This review was prepared and summarized by the WAKIB AI intelligence engine and vetted by our editorial board for accuracy and reliability.
Subscribe to Newsletter
Get a weekly summary of the most promising AI research and tools delivered to your inbox.
Telegram Channel
Join our active community on Telegram for real-time tracking of AI models and trends.
