Toolathlon-Verified is an AI benchmark used to evaluate model capabilities. Review its overview, metrics, official resources, and model leaderboard results on DataLearnerAI.
Browse the latest scores, model modes, release dates, and parameter sizes for Toolathlon-Verified.
Data sourced primarily from official releases (GitHub, Hugging Face, papers), then benchmark leaderboards, then third-party evaluators. Learn about our data methodology
| Rank | Model | License | |||
|---|---|---|---|---|---|
— | ![]() GLM-5.3-Flash Thinking Level · MaxTools | 78.40 | 2026-08-26 | 320B | Free Commercial |
— | ![]() Claude Fable 5.1 Thinking Level · MaxTools | 77.80 | 2026-09-01 | Unknown | Closed |
— | ![]() Kimi K3 Thinking Level · MaxTools | 76.50 | 2026-07-16 | 2800B | Conditional |
— | ![]() DeepSeek-V4-Flash-Vision-Exp Thinking Level · MaxTools | 75.90 | 2026-08-21 | 305B | Free Commercial |
— | ![]() DeepSeek-V4-Pro Thinking Level · Extra HighTools | 74.10 | 2026-08-13 | 1600B | Free Commercial |
— | ![]() Hy4 preview Thinking Level · HighTools | 74.10 | 2026-08-28 | 770B | Free Commercial |
— | ![]() Qwen3.8-Flash-Next Thinking Level · Extra HighTools | 73.50 | 2026-08-26 | 125B | Conditional |
— | ![]() Qwen3.8-Max-0902 Thinking Level · Extra HighTools | 73.30 | 2026-09-02 | 2400B | Closed |
— | ![]() GLM-5.3 Thinking Level · MaxTools | 73.00 | 2026-08-14 | 744B | Conditional |
— | ![]() Qwen3.8-Max Thinking Level · Extra HighTools | 72.50 | 2026-08-03 | 2400B | Conditional |
— | ![]() DeepSeek-V4-Flash Thinking Level · MaxTools | 70.30 | 2026-04-24 | 284B | Free Commercial |