ExploitGym evaluates whether AI agents can turn real-world vulnerabilities into working exploits across userspace programs, V8, and the Linux kernel. This entry records the number of solved instances under the 2-hour timeout budget used in the GLM-5.3 release evaluation; the current benchmark release contains 869 instances.
查看 ExploitGym (2h) 的最新得分、模型模式、发布时间与参数规模,快速了解当前完整榜单表现。
数据优先来自官方发布(GitHub、Hugging Face、论文),其次为评测基准官方结果,最后为第三方评测机构数据。 了解数据收集方法
| 排名 | 模型 | 开源情况 | |||
|---|---|---|---|---|---|
![]() GLM-5.3 思考水平·Max工具 | 105.00 | 2026-08-14 | 7533.3亿 | 闭源 |
Z.ai reported 105 solved instances for GLM-5.3 under the 2-hour budget on 2026-08-14, using Claude Code 2.1.207, max reasoning effort, no web tools, and a TPS-normalized API inference-time budget.