OZZZER · AI NEWS2 of 3 free stories opened
← Back to AI News

Research · 29 Jul 2026 · 17:00 CEST

How enabling two settings tripled our scores on the ARC-AGI-3 benchmark

OpenAI · 29 Jul 2026 · 17:00 CESTRead original at OpenAI ↗
Share
LinkedInX
How enabling two settings tripled our scores on the ARC-AGI-3 benchmark

Publisher preview · OZZZER analysis pending editorial review.

PUBLISHER ARTICLE PREVIEW

From the original article

A sped-up video of GPT‑5.6 Sol attempting to solve puzzles in the ARC-AGI-3 benchmark, with the official harness (left) and our Responses API harness (right), which retains reasoning and enables compaction. On the leaderboard for this game⁠, no frontier model solves any level beyond the first. With our harness, GPT‑5.6 Sol solves all six.

When we first saw GPT‑5.6 Sol’s low scores on the ARC-AGI-3⁠ benchmark, we were puzzled.

GPT‑5.6 Sol has solved longstanding open problems in mathematics like the cycle double cover conjecture⁠ and beaten games like Pokémon FireRed. But on ARC-AGI-3, a benchmark of 2D puzzle games, GPT‑5.6 Sol scored just 7.8%, and GPT‑5.5 could barely play the games at all, scoring a paltry 0.4%.

Were 2D puzzle games unusually difficult for our models? Or was something else going on?

Benchmarks rarely measure AI models in isolation. They also measure less visible choices about API settings, harness design, and prompting. In the case of ARC-AGI-3, we discovered that turning on two API settings we use in ChatGPT and Codex—retained reasoning and compaction—tripled scores and cut output tokens by 6x on the public task set.

With the official harness, GPT‑5.6 Sol scored 13.3% on the ARC-AGI-3 public set. With retained reasoning and compaction, it scored 38.3%. Scores measure Relative Human Action Efficiency (RHAE⁠)—a metric comparing model performance to a human baseline. Based on official gameplay logs⁠, we estimate the average human tester scored 48%. Models are not told how they will be scored, and cannot see their score throughout—actions only return a text representation of each frame and what level they are on.

ARC-AGI-3 is a benchmark designed to measure how well AI agents learn and reason. Agents explore unfamiliar 2D games and infer how they work without explicit instructions. You can play 25 demo games at arcprize.org/tasks⁠.

ARC-AGI-3 uses an intentionally generic harness, without tools or special features. ARC’s reasoning was that a simple harness makes model shortcomings more visible and makes model comparisons more fair. Commercial developers, by contrast, optimize harnesses for each model’s features and quirks.

In gaming, GPT‑5.6 Sol

Source

OpenAI · 29 Jul 2026 · 17:00 CEST

Open the original at OpenAI ↗