ARC-AGI-2 is the second-generation abstract-reasoning benchmark for frontier reasoning systems. It emphasizes symbolic interpretation, compositional reasoning, contextual rule application, and computational efficiency. Its public data includes 1,000 training tasks and 120 public evaluation tasks, plus calibrated semi-private and private evaluation sets of 120 tasks each.
Browse the latest scores, model modes, release dates, and parameter sizes for ARC-AGI-2.
Data sourced primarily from official releases (GitHub, Hugging Face, papers), then benchmark leaderboards, then third-party evaluators. Learn about our data methodology
| Rank | Model | License | |||
|---|---|---|---|---|---|
![]() GPT-6 Astra Thinking Level · Max | 95.00 | 2026-09-03 | Unknown | Closed | |
![]() GPT-5.6 Sol Thinking Level · Max | 92.50 | 2026-06-26 | Unknown | Closed | |
![]() Claude Opus 5 Thinking Level · Max | 90.40 | 2026-07-24 | Unknown | Closed | |
4 | ![]() GPT-5.6 Sol Thinking Level · Extra High | 90.00 | 2026-06-26 | Unknown | Closed |
5 | ![]() Claude Fable 5.1 Thinking Level · Max | 90.00 | 2026-09-01 | Unknown | Closed |
6 | ![]() Claude Fable 5.1 Thinking Level · Extra High | 90.00 | 2026-09-01 | Unknown | Closed |
7 | ![]() Claude Fable 5 Thinking Level · Max | 89.20 | 2026-06-09 | Unknown | Closed |
8 | ![]() Claude Fable 5.1 Thinking Level · High | 88.80 | 2026-09-01 | Unknown | Closed |
9 | ![]() Claude Fable 5 Thinking Level · Extra High | 88.30 | 2026-06-09 | Unknown | Closed |
10 | ![]() Claude Fable 5 Thinking Level · High | 87.50 | 2026-06-09 | Unknown | Closed |
11 | ![]() Claude Fable 5.1 Thinking Level · Medium | 86.30 | 2026-09-01 | Unknown | Closed |
12 | ![]() GPT-5.6 Sol Thinking Level · High | 85.40 | 2026-06-26 | Unknown | Closed |
13 | ![]() GPT-5.5 Thinking Level · Extra High | 85.00 | 2026-04-23 | Unknown | Closed |
14 | ![]() GPT-5.5 Thinking Level · High | 85.00 | 2026-04-23 | Unknown | Closed |
15 | ![]() Gemini 3 Deep Think - 2620 Thinking Enabled | 84.60 | 2026-02-13 | Unknown | Closed |
16 | ![]() GPT-5.5 Pro Thinking Level · High | 84.60 | 2026-04-23 | Unknown | Closed |
17 | ![]() Gemini 3.7 Flash Thinking Level · High | 84.60 | 2026-08-13 | Unknown | Closed |
18 | ![]() GPT-5.5 Pro Thinking Level · Extra High | 84.20 | 2026-04-23 | Unknown | Closed |
19 | ![]() GPT-5.6 Terra Thinking Level · Max | 83.90 | 2026-06-26 | Unknown | Closed |
20 | ![]() GPT-5.4 Pro Thinking Level · High | 83.30 | 2026-03-05 | Unknown | Closed |
21 | ![]() Claude Fable 5 Thinking Level · Medium | 82.50 | 2026-06-09 | Unknown | Closed |
22 | ![]() Claude Fable 5.1 Thinking Level · Low | 78.30 | 2026-09-01 | Unknown | Closed |
23 | ![]() Gemini 3.1 Pro Preview Thinking Level · High | 77.10 | 2026-02-20 | Unknown | Closed |
24 | ![]() GPT-5.4 Standard Mode | 77.10 | 2026-03-05 | Unknown | Closed |
25 | ![]() Claude Fable 5 Thinking Level · Low | 76.80 | 2026-06-09 | Unknown | Closed |
26 | ![]() Opus 4.7 Thinking Level · Max | 75.80 | 2026-04-16 | Unknown | Closed |
27 | ![]() GPT-5.4 Thinking Level · Extra High | 74.00 | 2026-03-05 | Unknown | Closed |
28 | ![]() Gemini 3.5 Flash Thinking Level · HighTools | 72.10 | 2026-06-20 | Unknown | Closed |
29 | ![]() GPT-5.5 Thinking Level · Medium | 70.40 | 2026-04-23 | Unknown | Closed |
30 | ![]() Opus 4.7 Thinking Level · High | 68.30 | 2026-04-16 | Unknown | Closed |
As of September 1, 2026, the newest verified entry is Claude Fable 5.1: 90.0% at Max and XHigh, 88.8% at High, 86.3% at Medium, and 78.3% at Low. The overall verified high is 92.5% from GPT-5.6 Sol (Max).