ExploitBench evaluates how far AI agents progress through a deterministic exploit-capability ladder on 41 real V8 vulnerabilities, from reaching vulnerable code through exploit primitives to arbitrary code execution. The GLM-5.3 release evaluation averages capability coverage over all tasks and three revisions with up to 300 agent-environment interaction rounds.
查看 ExploitBench 的最新得分、模型模式、发布时间与参数规模,快速了解当前完整榜单表现。
数据优先来自官方发布(GitHub、Hugging Face、论文),其次为评测基准官方结果,最后为第三方评测机构数据。 了解数据收集方法
| 排名 | 模型 | 开源情况 | |||
|---|---|---|---|---|---|
![]() GLM-5.3 思考水平·Max工具 | 54.40 | 2026-08-14 | 7533.3亿 | 闭源 |
Z.ai reported an average capability coverage score of 54.4 for GLM-5.3 on 2026-08-14, using Claude Code 2.1.207, max reasoning effort, no web tools, and 128K maximum output.