ARC-AGI-3 is the first interactive reasoning benchmark designed for human-comparable agent evaluation. Without natural-language instructions, agents must explore unfamiliar environments, infer goals on the fly, build and update world models, plan over long horizons, and learn from sparse feedback. A 100% score means completing every game as efficiently as humans.
Browse the latest scores, model modes, release dates, and parameter sizes for ARC-AGI-3.
Data sourced primarily from official releases (GitHub, Hugging Face, papers), then benchmark leaderboards, then third-party evaluators. Learn about our data methodology
| Rank | Model | License | |||
|---|---|---|---|---|---|
— | ![]() GPT-6 Astra Thinking Level · Max | 62.70 | 2026-09-03 | Unknown | Closed |
— | ![]() Claude Opus 5 Thinking Level · High | 30.20 | 2026-07-24 | Unknown | Closed |
— | ![]() GPT-5.6 Sol Thinking Level · Max | 7.80 | 2026-06-26 | Unknown | Closed |
— | ![]() GPT-5.6 Sol Thinking Level · Extra High | 7.00 | 2026-06-26 | Unknown | Closed |
— | ![]() GPT-5.6 Sol Thinking Level · High | 2.10 | 2026-06-26 | Unknown | Closed |
— | Grok 4.6 Thinking Level · Extra High | 2.10 | 2026-08-12 | Unknown | Closed |
— | ![]() GPT-5.6 Sol Thinking Level · Medium | 1.10 | 2026-06-26 | Unknown | Closed |
— | ![]() GPT-5.6 Terra Thinking Level · Max | 0.8 | 2026-06-26 | Unknown | Closed |
— | ![]() GPT-5.6 Sol Thinking Level · Low | 0.3 | 2026-06-26 | Unknown | Closed |
— | ![]() GPT-5.6 Luna Thinking Level · Max | 0.2 | 2026-06-26 | Unknown | Closed |
— | ![]() Claude Opus 4.6 Thinking Level · Max | 0.0045 | 2026-02-05 | Unknown | Closed |
— | ![]() GPT-5.5 Thinking Level · High | 0.0043 | 2026-04-23 | Unknown | Closed |
— | ![]() Gemini 3.1 Pro Preview Thinking Level · High | 0.004 | 2026-02-20 | Unknown | Closed |
— | ![]() GPT-5.4 Thinking Level · High | 0.002 | 2026-03-05 | Unknown | Closed |
— | ![]() Opus 4.7 Thinking Level · High | 0.0018 | 2026-04-16 | Unknown | Closed |
— | Grok 4.20 Thinking Enabled | 0.001 | 2026-03-09 | Unknown | Closed |
As of September 1, 2026, the verified high is 30.2% from Claude Opus 5 (High), followed by 7.8% from GPT-5.6 Sol (Max). The newest ARC-AGI-3 result is 2.1% from Grok 4.6 (XHigh), while the human baseline is 100%.