AA-AnalystAgent is a private research-agent evaluation from Artificial Analysis with 80 questions across 14 domains. Each task is attempted five times, and the headline pass^5 score requires all five attempts to be correct. The Stirrup harness provides code execution, web fetch, image viewing, and answer-submission tools, so results include the named reasoning setting and internet-enabled tool environment.
Browse the latest scores, model modes, release dates, and parameter sizes for AA-AnalystAgent.
Data sourced primarily from official releases (GitHub, Hugging Face, papers), then benchmark leaderboards, then third-party evaluators. Learn about our data methodology
| Rank | Model | License | |||
|---|---|---|---|---|---|
![]() Gemini 3.7 Flash Thinking Level · HighToolsInternet | 60.00 | 2026-08-13 | Unknown | Closed | |
![]() Claude Opus 5 Thinking Level · MaxToolsInternet | 53.75 | 2026-07-24 | Unknown | Closed | |
![]() GPT-5.5 Thinking Level · Extra HighToolsInternet | 50.00 | 2026-04-23 | Unknown | Closed | |
4 | ![]() Claude Fable 5 Thinking Level · MaxToolsInternet | 48.75 | 2026-06-09 | Unknown | Closed |
5 | ![]() GPT-5.6 Sol Thinking Level · MaxToolsInternet | 47.50 | 2026-06-26 | Unknown | Closed |
6 | Grok 4.6 Thinking Level · HighToolsInternet | 41.25 | 2026-08-12 | Unknown | Closed |
7 | ![]() Kimi K3 Thinking Level · MaxToolsInternet | 38.75 | 2026-07-16 | 2800B | Conditional |
8 | Inkling Thinking EnabledToolsInternet | 23.75 | 2026-07-15 | 975B | Free Commercial |
9 | ![]() Haiku 4.5 Thinking EnabledToolsInternet | 15.00 | 2025-10-15 | Unknown | Closed |
10 | ![]() Mistral Medium 3.5 Standard ModeToolsInternet | 12.50 | 2026-05-01 | Unknown | Closed |
11 | ![]() MiniMax M3 Thinking EnabledToolsInternet | 10.00 | 2026-06-01 | 428B | Non-Commercial |
12 | ![]() Nemotron 3 Ultra Thinking EnabledToolsInternet | 6.25 | 2026-06-04 | 550B | Free Commercial |
In the official leaderboard reviewed on August 30, 2026, Gemini 3.7 Flash at high reasoning leads with 60.0% pass^5, followed by Claude Opus 5 at max with 53.75% and GPT-5.5 at xhigh with 50.0%.