ExploitBench evaluates how far AI agents progress through a deterministic exploit-capability ladder on 41 real V8 vulnerabilities, from reaching vulnerable code through exploit primitives to arbitrary code execution. The GLM-5.3 release evaluation averages capability coverage over all tasks and three revisions with up to 300 agent-environment interaction rounds.
Browse the latest scores, model modes, release dates, and parameter sizes for ExploitBench.
Data sourced primarily from official releases (GitHub, Hugging Face, papers), then benchmark leaderboards, then third-party evaluators. Learn about our data methodology
| Rank | Model | License | |||
|---|---|---|---|---|---|
![]() GLM-5.3 Thinking Level · MaxTools | 54.40 | 2026-08-14 | 753.3B | Closed |
Z.ai reported an average capability coverage score of 54.4 for GLM-5.3 on 2026-08-14, using Claude Code 2.1.207, max reasoning effort, no web tools, and 128K maximum output.