Agents' Last Exam evaluates AI agents on long-horizon, economically valuable professional workflows with verifiable outcomes across a broad set of industries.
Browse the latest scores, model modes, release dates, and parameter sizes for Agents' Last Exam.
Data sourced primarily from official releases (GitHub, Hugging Face, papers), then benchmark leaderboards, then third-party evaluators. Learn about our data methodology
| Rank | Model | License | |||
|---|---|---|---|---|---|
![]() GPT-5.6 Sol Thinking Level · Extra HighTools | 52.70 | 2026-06-26 | Unknown | Closed | |
![]() GPT-5.6 Terra Thinking Level · Extra HighTools | 50.40 | 2026-06-26 | Unknown | Closed | |
![]() GPT-5.6 Luna Thinking Level · Extra HighTools | 50.30 | 2026-06-26 | Unknown | Closed | |
4 | ![]() GLM-5.3 Thinking Level · MaxTools | 28.50 | 2026-08-14 | 753.3B | — |
5 | ![]() Kimi K3 Thinking Level · MaxTools | 28.30 | 2026-07-16 | 2800B | Free Commercial |
6 | ![]() DeepSeek-V4-Flash-Vision-Exp Thinking Level · MaxTools | 27.30 | 2026-08-21 | Unknown | — |
7 | ![]() Qwen3.8-Max Thinking Level · Extra HighTools | 27.00 | 2026-08-03 | 2400B | Free Commercial |
8 | ![]() Gemini 3.7 Flash Thinking Level · MediumTools | 26.30 | 2026-08-13 | Unknown | Closed |
9 | ![]() DeepSeek-V4-Pro Thinking Level · Extra HighTools | 25.70 | 2026-08-13 | 1600B | Free Commercial |
10 | ![]() DeepSeek-V4-Flash Thinking Level · MaxTools | 25.20 | 2026-04-24 | 284B | Free Commercial |
11 | ![]() Qwen3.8-27B Thinking EnabledTools | 20.40 | 2026-08-14 | 27B | Free Commercial |