ZeroBench is a difficult visual-reasoning benchmark with 100 manually curated main questions and 334 subquestions, spanning natural and synthetic images and both single-image and multi-image settings. Main-set results commonly report pass@1, pass@5, and pass^5. Vendor-reported release scores and benchmark-team official leaderboard runs should be interpreted separately with their evaluation settings.
Browse the latest scores, model modes, release dates, and parameter sizes for ZeroBench Main.
Data sourced primarily from official releases (GitHub, Hugging Face, papers), then benchmark leaderboards, then third-party evaluators. Learn about our data methodology
| Rank | Model | License | |||
|---|---|---|---|---|---|
![]() Kimi K3 Thinking Level · MaxTools | 41.00 | 2026-07-16 | 2800B | Conditional | |
![]() DeepSeek-V4-Flash-Vision-Exp Thinking Level · MaxTools | 35.00 | 2026-08-21 | Unknown | — | |
![]() Kimi K3 Thinking Level · Max | 23.00 | 2026-07-16 | 2800B | Conditional |
The official ZeroBench site separates benchmark-team evaluations from externally reported model-card and release scores. Compare provenance, sampling, tools, and inference settings alongside the score.