NL2Repo-Bench evaluates whether coding agents can build a complete, installable Python repository from a natural-language specification and an empty workspace. The 104 tasks are graded against upstream pytest suites.
查看 NL2Repo-Bench 的最新得分、模型模式、发布时间与参数规模,快速了解当前完整榜单表现。
数据优先来自官方发布(GitHub、Hugging Face、论文),其次为评测基准官方结果,最后为第三方评测机构数据。 了解数据收集方法
| 排名 | 模型 | 开源情况 | |||
|---|---|---|---|---|---|
![]() DeepSeek-V4-Pro 思考水平·极高工具 | 61.50 | 2026-08-13 | 16000亿 | 免费商用 | |
![]() GLM-5.3 思考水平·Max工具 | 58.00 | 2026-08-14 | 7533.3亿 | — | |
![]() Qwen3.8-Max 思考水平·极高工具 | 55.90 | 2026-08-03 | 24000亿 | 免费商用 | |
4 | ![]() DeepSeek-V4-Flash 思考水平·Max工具 | 54.20 | 2026-04-24 | 2840亿 | 免费商用 |
5 | ![]() GLM-5.2 开启思考工具 | 48.90 | 2026-06-13 | 7533.3亿 | 免费商用 |
6 | ![]() Qwen3.8-27B 开启思考工具 | 42.30 | 2026-08-14 | 270亿 | 免费商用 |
7 | ![]() MiniMax-M2.7 开启思考工具 | 39.80 | 2026-03-18 | 2290亿 | 非商用 |
DeepSeek-V4-Flash-0731 官方发布成绩:54.2(2026-07-31,DeepSeek Harness minimal,max effort,temperature=1.0,top_p=0.95)。