AutomationBench, created by Zapier, evaluates agents on realistic cross-application business workflows across sales, marketing, operations, support, finance, and HR. The public score covers 600 tasks across 47 simulated SaaS tools and uses assertion-based final-state grading. Zapier's verified leaderboard uses a separate held-out private set, so public and private scores must be distinguished.
Browse the latest scores, model modes, release dates, and parameter sizes for AutomationBench.
Data sourced primarily from official releases (GitHub, Hugging Face, papers), then benchmark leaderboards, then third-party evaluators. Learn about our data methodology
| Rank | Model | License | |||
|---|---|---|---|---|---|
![]() GLM-5.3-Flash Thinking Level · MaxTools | 48.80 | 2026-08-26 | 320B | Free Commercial | |
![]() GLM-5.3 Thinking Level · MaxTools | 48.20 | 2026-08-14 | 744B | Conditional | |
![]() Hy4 preview Thinking Level · HighTools | 32.10 | 2026-08-28 | 770B | Free Commercial | |
4 | ![]() DeepSeek-V4-Pro Thinking Level · Extra HighTools | 31.80 | 2026-08-13 | 1600B | Free Commercial |
5 | ![]() Kimi K3 Thinking Level · MaxTools | 30.80 | 2026-07-16 | 2800B | Conditional |
6 | ![]() Gemini 3.7 Flash Thinking EnabledTools | 30.40 | 2026-08-13 | Unknown | Closed |
7 | ![]() Qwen3.8-Max Thinking Level · Extra HighTools | 27.30 | 2026-08-03 | 2400B | Conditional |
8 | ![]() Claude Opus 5 Thinking Level · MaxTools | 26.00 | 2026-07-24 | Unknown | Closed |
9 | ![]() DeepSeek-V4-Flash-Vision-Exp Thinking Level · MaxTools | 25.70 | 2026-08-21 | Unknown | — |
10 | ![]() DeepSeek-V4-Flash Thinking Level · MaxTools | 25.10 | 2026-04-24 | 284B | Free Commercial |
DeepSeek reported 25.7 on AutomationBench (Public) for DeepSeek-V4-Flash-Vision-Exp on August 21, 2026.