Jul 26, 2026
AI

Claude Opus 5 ARC-AGI-3 result tops GPT-5.6 Sol

ARC Prize says Anthropic’s model scored 30.2% on ARC-AGI-3, far above GPT-5.6 Sol’s prior 7.8% record.

Wei-Lin Zhao

By Wei-Lin Zhao · AI Correspondent

· 3 min read

Claude Opus 5 ARC-AGI-3 result tops GPT-5.6 Sol
Photo: The Decoder

The Claude Opus 5 ARC-AGI-3 result gives Anthropic a clear lead on one of the more watched tests for interactive reasoning: 30.2%, according to ARC Prize. OpenAI’s GPT-5.6 Sol (Max) previously led the benchmark with 7.8%, while Anthropic’s Fable-class models were around 20%, ARC Prize said.

ARC Prize attributed the jump to stronger logical reasoning, saying the model was better able to explore, plan and execute in unfamiliar environments. During the evaluation, the group said Opus 5 translated tasks into algebraic notation and independently derived reflection equations, behaviors it had not previously observed from an AI model.

The model also solved five environments that had not been solved before. ARC Prize said four of those were at or above human level. Six of the 25 public demo environments have now been completed, and ARC Prize has published the results, replays and benchmarking code.

What is ARC-AGI-3?

ARC-AGI-3 is a benchmark meant to test whether AI systems can solve new tasks they did not see during training. The current version is structured like an interactive game: a model has to infer the rules of an environment, choose actions and carry them out over multiple steps.

That design makes the test more relevant to agent-style systems than static Q&A benchmarks. ARC Prize’s official scores count the language model’s own performance, rather than giving credit for external software harnesses that can add planning or tool-use scaffolding. ARC Prize has argued that future AGI systems should not depend on outside systems to handle unfamiliar tasks.

Opus 5 also matched existing top results on older ARC tests. ARC Prize reported a 90.4% score on ARC-AGI-2 and 97.5% on ARC-AGI-1, with costs slightly above the prior best runs. The group did not frame those older results as the main change; the gap on ARC-AGI-3 is the outlier.

Did Anthropic train directly for this benchmark?

Anthropic has not explained what produced the gain. The public timing leaves room for a narrower explanation: Opus 5 was developed after ARC-AGI-3 and its format were available, so Anthropic may have been able to train toward the skills and puzzle styles the benchmark rewards. That does not show that Anthropic trained on the exact benchmark tasks.

Guanghan Ning’s private Witness benchmark points to a more mixed picture. Ning reported that Opus 5 scored 43.4 on the interactive puzzle-game test, statistically tied with Kimi K3 and Fable 5. He said the improvement over Opus 4.8 was much smaller than the improvement shown on ARC-AGI-3.

Ning also said Opus 5 identified hidden rules in a conventional puzzle before taking action, but lagged Opus 4.8 on a game with less familiar mechanics. He said that pattern fits the possibility of training on genre-specific data, while adding that Witness cannot determine what data Anthropic used.

Greg Kamradt, one of the researchers behind ARC-AGI-3, said the Witness results do not rule out broader reasoning gains. He argued that a familiar game may not test adaptation to novelty, and that one weaker task is not enough to outweigh the overall improvement without more detailed scoring. Ning later said Opus 5 did generalize to Witness, though by a smaller margin than on ARC-AGI-3.

This story draws on original reporting from The Decoder.

More from AI

All AI →