ExploitGym evaluates whether AI agents can turn real-world vulnerabilities into working exploits across userspace programs, V8, and the Linux kernel. This entry records the number of solved instances under the 6-hour timeout budget used in the GLM-5.3 release evaluation; the current benchmark release contains 869 instances.
Browse the latest scores, model modes, release dates, and parameter sizes for ExploitGym (6h).
Data sourced primarily from official releases (GitHub, Hugging Face, papers), then benchmark leaderboards, then third-party evaluators. Learn about our data methodology
| Rank | Model | License | |||
|---|---|---|---|---|---|
![]() GLM-5.3 Thinking Level · MaxTools | 130.00 | 2026-08-14 | 753.3B | Closed |
Z.ai reported 130 solved instances for GLM-5.3 under the 6-hour budget on 2026-08-14, using Claude Code 2.1.207, max reasoning effort, no web tools, and a TPS-normalized API inference-time budget.