GPT-6 Astra's ARC-AGI-3 Split: 62.7% vs 99.9%, Same Model
On ARC-AGI-3, OpenAI's GPT-6 Astra scores 62.7% on the Standard harness and 99.9% on the Provider Adapter harness — both certified SOTA by ARC Prize, which now reports both. We break down what each harness measures, why the gap is so large, and what it means for benchmark credibility.
Same model, same benchmark, same set of tasks — and a 37-point gap. On September 3, 2026, ARC Prize published GPT-6 Astra's results on ARC-AGI-3 under two different harnesses, and both numbers are official.
Two Harnesses, Two Scores
The headline pair: Astra (max effort) scored 62.7% for $26,098 on the Standard harness, while Astra (high effort) hit 99.9% for $18,817 on the Provider Adapter harness. ARC Prize certified both as state-of-the-art and announced it will report the two tracks side by side going forward — the first time a single leaderboard entry has split like this.
The difference is not the model. The Standard harness is provider-neutral: a minimal, uniform interface where the model itself decides how to keep notes between steps. The Provider Adapter harness preserves Astra's opaque chain-of-thought across turns and applies context compaction — letting the model carry its own scratchpad through long tasks. Astra used both freedoms aggressively: it invented its own algebraic shorthand for ARC's puzzle notation ("L8: hub q2 (8↓). Lengths: 14=1…") and, in longer runs, built its own tools from scratch — including a maze_solver.py — inside the PRO-LONG harness.
The Effort Curve, Both Ways
Across the full effort sweep the pattern holds. Standard: max 62.7%, xhigh 59.3%, high 54.8%, medium 38.6%, low 17.5%, none 35.2%. Provider Adapter: max 98.6%, xhigh 98.4%, high 99.9%, medium 98.4%, low 98.0%, none 96.7% — the top score actually lands on high, not max, and higher reasoning levels generally cost less per run, not more (Provider Adapter: max $17,332 vs none $23,457), because extra reasoning trims the action count. On 167 jointly-solved game-reasoning pairs, the Provider Adapter track finished ~3.66x faster while using 49% fewer tokens. Efficiency per action beat the human baseline too: on 96.0% of levels Astra used fewer actions than human testers, averaging 51.7% fewer — while human solving costs roughly $12.78 per game ($115 for 90 minutes plus $5 per game).
What This Does to Benchmark Credibility
The uncomfortable takeaway: the harness is now a first-class variable, sometimes bigger than the model. A leaderboard number without its harness context has stopped being a fact. ARC Prize's own framing is careful — it does not claim Astra is AGI, and it cautions that a saturated score is not proof of general intelligence. But it does call the Provider Adapter results a step-function change in frontier capability. For everyone watching the AGI race, the rule going forward is simple: read the harness before you read the score.