ExploitBench evaluates how far AI agents progress through a deterministic exploit-capability ladder on 41 real V8 vulnerabilities, from reaching vulnerable code through exploit primitives to arbitrary code execution. The GLM-5.3 release evaluation averages capability coverage over all tasks and three revisions with up to 300 agent-environment interaction rounds.
查看 ExploitBench 的最新得分、模型模式、发布时间与参数规模,快速了解当前完整榜单表现。
数据优先来自官方发布(GitHub、Hugging Face、论文),其次为评测基准官方结果,最后为第三方评测机构数据。 了解数据收集方法
| 排名 | 模型 | 开源情况 | |||
|---|---|---|---|---|---|
— | ![]() GPT-6 Astra 思考水平·Max工具 | 100.00 | 2026-09-03 | 未知 | 闭源 |
— | ![]() GPT-5.6 Sol unknown | 78.50 | 2026-06-26 | 未知 | 闭源 |
— | ![]() GLM-5.3 思考水平·Max工具 | 54.40 | 2026-08-14 | 7440亿 | 有条件商用 |
2026-09-05:已核对 Astra / Sol 原始来源;不同版本、子集与工具配置分别保存。 https://openai.com/index/gpt-6-astra/