OpenAI has found that enabling retained reasoning and compaction in its Responses API close to tripled GPT-5.6 Sol’s score on the ARC-AGI-3 puzzle benchmark, though the result remains self-reported.
Picture a model that can, according to OpenAI, help crack open unsolved problems in mathematics — yet stumbles badly when asked to learn a simple 2D puzzle game from scratch. That apparent contradiction sits at the heart of a public debate that broke out this week over how AI systems are tested, and what benchmark scores actually tell us.
The model in question is GPT-5.6 Sol, OpenAI’s frontier reasoning system. The test is ARC-AGI-3, a benchmark built by the ARC Prize Foundation that presents AI agents with unfamiliar grid-based puzzle games. The catch: the AI must learn each game through trial and error, not by drawing on anything it has seen before. It’s closer to watching someone play a board game for the first time than it is to answering a multiple-choice question.
A Puzzling Gap
When GPT-5.6 Sol was first run against ARC-AGI-3 using the standard ARC harness, it scored about 7.8%. That was, a state-of-the-art result. Its predecessor, GPT-5.5, managed only around 0.4%. Anthropic’s Claude Opus 4.8 had scored about 1.5%. Most frontier models score close to zero on this benchmark. So 7.8% was a genuine step forward — and GPT-5.6 Sol became the first verified frontier model to beat a specific ARC-AGI-3 game, clearing environment ft09 with an 87% score.
But critics were quick to point out the tension. A model capable of contributing to frontier mathematics research was failing to learn the rules of a novel puzzle game in about nine attempts out of ten. Something, they argued, didn’t add up.
OpenAI looked into it. Their conclusion was straightforward: the harness was getting in the way.
What the Harness Was Missing
The ARC harness is the scaffolding around the model — the software that manages how the AI receives information, takes actions, and moves between steps in a game. OpenAI’s internal analysis found that the standard ARC harness wasn’t allowing GPT-5.6 Sol to carry forward what it had learned from one move to the next, or from one episode to another.
Think of it like playing chess with no memory of the last ten moves. You could be a grandmaster and still look ordinary.
OpenAI identified two settings in its Responses API as the missing pieces. The first, retained reasoning, allows the system to carry internal reasoning traces forward between calls rather than starting fresh each time. The second, compaction, compresses the conversation history and game state so the model doesn’t drown in its own context as a session grows longer.
After enabling both settings on the public ARC-AGI-3 task set, OpenAI reports that GPT-5.6 Sol’s score rose from around 13.3% to close to 38.3%. Output tokens used dropped by roughly a factor of six. That’s a meaningful efficiency gain alongside the performance jump.
The Leaderboard Complication
Here’s where it gets complicated. OpenAI’s 38.3% figure is self-reported, applies only to the public subset of ARC-AGI-3, and uses a custom harness configuration rather than the official ARC Prize setup. It does not appear on the official leaderboard.
The current official leader on the semi-private ARC-AGI-3 benchmark is Anthropic’s Claude Opus 5, recorded at 30.2%. GPT-5.6 Sol’s official entry remains at 7.8% under the standard harness. There is also a separate Benchmark Registry entry listing a 0.3% “Overall” result for GPT-5.6 Sol, which appears to reflect an earlier or differently configured run and conflicts with the later 7.8% figure — a reminder that even the record-keeping around these benchmarks can be messy.
François Chollet, the creator of the ARC-AGI benchmark series, has previously been clear that the benchmark is designed to test abstraction and reasoning on genuinely novel problems, not memorised patterns. The ARC Prize Foundation has not yet incorporated OpenAI’s retained reasoning and compaction configuration into its official evaluation protocol.
Independent analysts have pointed out that this episode illustrates something broader. As one external commentator put it, ARC-AGI-3 is testing the complete agent system — model, memory, context management, harness design — not the raw neural network in isolation. Changing those surrounding components can shift scores sharply, which raises real questions about what it means to compare models on the same benchmark.
Still Far From Solving It
Even at 38.3% on the public tasks, GPT-5.6 Sol is failing most ARC-AGI-3 games. The benchmark’s design means that scoring 100% would require something close to human-level flexible reasoning on entirely novel problems — a bar no current system is near.
OpenAI has not claimed otherwise. But the gap between “solving open mathematics problems” and “learning a new puzzle game” does make clear that AI capability can be highly domain-specific. A system can be extraordinarily powerful in one area and genuinely weak in another, sometimes for reasons that turn out to be about configuration rather than fundamental intelligence.
That’s an important distinction for anyone trying to assess what these systems can actually do.
What This Means for Kent Residents
For Kent businesses and developers using OpenAI’s APIs, this story is a practical reminder that getting the best out of AI tools often comes down to configuration — settings like retained reasoning and compaction can matter as much as which model you choose. Anyone in Kent evaluating AI vendors or tools should treat self-reported benchmark figures, such as OpenAI’s 38.3% claim, with caution until they appear on independently verified leaderboards. Universities and colleges in the county teaching computer science or AI will find the GPT-5.6 Sol episode a useful case study in how evaluation design shapes results — and why headline scores don’t always tell the full story.
Source: @OpenAI
OpenAI Says Two API Settings Tripled GPT-5.6 Sol's ARC-AGI-3 Benchmark Score Quiz
5 questions