Jul 30, 2026
AI

GPT-5.6 Sol tops Opus 5 on ARC-AGI-3 in OpenAI test

OpenAI says GPT-5.6 Sol reached 38.3% on ARC-AGI-3 using its Responses API, but the result depends on nonstandard settings.

Colin Brandt

By Colin Brandt · Enterprise Reporter

· 3 min read

GPT-5.6 Sol tops Opus 5 on ARC-AGI-3 in OpenAI test
Photo: The Decoder

OpenAI says GPT-5.6 Sol scored 38.3% on ARC-AGI-3 when tested through its Responses API with two settings enabled, putting it ahead of Anthropic’s Claude Opus 5 score of 30.2%. The claim matters because the same model posted 7.8% in the official ARC Prize harness, making the result as much a dispute over evaluation setup as model capability.

OpenAI published the result after Anthropic’s Opus 5 set a new mark on the logic benchmark. According to OpenAI, GPT-5.6 Sol’s lower official score came from the benchmark environment discarding the model’s reasoning after each action. The company said its own test used Retained Reasoning and Compaction, two Responses API features it argues better reflect how developers can run the model in production.

How did GPT-5.6 Sol score higher on ARC-AGI-3?

Retained Reasoning preserves the model’s intermediate reasoning state across steps instead of forcing each action to start without that context. Compaction compresses older context into summaries rather than cutting it off, which can matter in long-running tasks where the model has to carry forward what it has already inferred.

OpenAI’s position is that a benchmark score reflects the model plus the execution system around it, including API behavior and state management. That argument is commercially relevant for AI vendors because frontier model comparisons increasingly depend on harnesses, tools, memory behavior and inference-time controls rather than a raw prompt-response call.

ARC-AGI-3, however, is designed to test model performance under a standardized setup. ARC Prize said its official scores use a common approach without provider-specific configuration so that comparisons remain consistent. The disagreement turns on whether that common harness gave Anthropic access to capabilities through its API that OpenAI’s older completions-style setup did not expose.

ARC Prize co-founder François Chollet later drew a line between two categories of evaluation. He said harnesses built specifically to solve the benchmark, or ones that contain knowledge of its format, should not be allowed. General-purpose API settings that were not created for ARC-AGI-3 and are available to all API customers are acceptable, he said.

Chollet also said ARC Prize had discussed testing details with OpenAI, including compaction, and welcomed the company’s progress on the issue. He acknowledged that different providers using different settings can create a parity problem, but said that is acceptable if the settings and costs are disclosed clearly.

The result leaves two numbers attached to GPT-5.6 Sol: 7.8% in the official ARC Prize harness and 38.3% in OpenAI’s Responses API setup. For buyers and builders comparing frontier models, the gap is a reminder to read the test conditions before treating benchmark rankings as procurement guidance.

This story draws on original reporting from The Decoder.

More from AI

All AI →