Scale Labs and Reflection's refreshed public software engineering agent evaluation: 642 tasks across 11 repositories, with corrected tasks, updated verifiers and a network-locked evaluation protocol. The original 731-task public scores use a different dataset and are not directly comparable.
Browse the latest scores, model modes, release dates, and parameter sizes for SWE-Bench Pro V2.
Data sourced primarily from official releases (GitHub, Hugging Face, papers), then benchmark leaderboards, then third-party evaluators. Learn about our data methodology
| Rank | Model | License | |||
|---|---|---|---|---|---|
— | ![]() Claude Opus 5 Thinking Level · Extra HighTools | 99.40 | 2026-07-24 | Unknown | Closed |
— | ![]() Claude Fable 5.1 Thinking Level · HighTools | 99.10 | 2026-09-01 | Unknown | Closed |
— | ![]() Kimi K3 Thinking Level · MaxTools | 97.70 | 2026-07-16 | 2800B | Conditional |
— | ![]() GPT-6 Astra Thinking Level · HighTools | 96.90 | 2026-09-03 | Unknown | Closed |
— | ![]() GLM-5.3 Thinking Level · MaxTools | 95.60 | 2026-08-14 | 744B | Conditional |
— | ![]() GPT-5.6 Sol Thinking Level · Extra HighTools | 95.50 | 2026-06-26 | Unknown | Closed |
— | ![]() Gemini 3.8 Flash Thinking Level · HighTools | 94.86 | 2026-09-02 | Unknown | Closed |
— | ![]() Claude Sonnet 5 Thinking Level · Extra HighTools | 93.15 | 2026-06-30 | Unknown | Closed |
— | ![]() GPT-5.6 Terra Thinking Level · Extra HighTools | 92.37 | 2026-06-26 | Unknown | Closed |
— | Inkling Thinking Level · Extra HighTools | 89.88 | 2026-07-15 | 975B | Free Commercial |
Official Scale Labs results, September 22, 2026. Scores use pass@1 resolve rate under a network-locked agent protocol. Agent harnesses differ; public task exposure may affect generalization claims.