DeepSeek-V4-ProvsGLM-5.2
Across 11 shared benchmarks, GLM-5.2 leads overall: DeepSeek-V4-Pro wins 4, GLM-5.2 wins 7, with 0 ties and an average score difference of -12.23.
DeepSeek-V4-Pro
DeepSeek-AI · 2026-08-13 · Reasoning model
GLM-5.2
智谱AI · 2026-06-13 · Reasoning model
DeepSeek-V4-Pro4 wins(36%)(64%)7 winsGLM-5.2
Benchmark scores
Grouped by capability, sorted by largest gap within each. 11 shared benchmarks.
Coding and Software Engineer
DeepSeek-V4-Pro 2/3| Benchmark | DeepSeek-V4-Pro | GLM-5.2 | Diff |
|---|---|---|---|
| DeepSWE | 62.7011 / 27极高强度思考(工具) | 4421 / 27Deep Thinking (With Tools) | +18.70 |
| NL2Repo-Bench | 61.501 / 8极高强度思考(工具) | 48.906 / 8Thinking (With Tools) | +12.60 |
| SWE-Bench Pro - Public | 52.1040 / 57Normal (With Tools) | 62.109 / 57Thinking (With Tools) | -10 |
Math and Reasoning
GLM-5.2 3/3| Benchmark | DeepSeek-V4-Pro | GLM-5.2 | Diff |
|---|---|---|---|
| IMO-AnswerBench | 35.3023 / 23Normal (No Tools) | 912 / 23Thinking (No Tools) | -55.70 |
| FrontierMath Tier 4 v2 | 2.4430 / 34最高(无工具) | 29.2716 / 34最高(无工具) | -26.83 |
| FrontierMath v2 | 45.2626 / 34最高(无工具) | 59.2119 / 34最高(无工具) | -13.94 |
General Knowledge
GLM-5.2 2/2| Benchmark | DeepSeek-V4-Pro | GLM-5.2 | Diff |
|---|---|---|---|
| HLE | 7.70165 / 181Normal (No Tools) | 54.7015 / 181Thinking (With Tools) | -47 |
| LiveBench | 73.5823 / 115Normal (No Tools) | 76.249 / 115Normal (No Tools) | -2.66 |
AI Agent - Tool Usage
DeepSeek-V4-Pro 1/1| Benchmark | DeepSeek-V4-Pro | GLM-5.2 | Diff |
|---|---|---|---|
| Terminal-Bench 2.1 | 87.906 / 44极高强度思考(工具) | 8117 / 44Thinking High (With Tools) | +6.90 |
Commonsense Reasoning
DeepSeek-V4-Pro 1/1| Benchmark | DeepSeek-V4-Pro | GLM-5.2 | Diff |
|---|---|---|---|
| SimpleBench | 61.2016 / 67Normal (No Tools) | 58.8020 / 67Normal (No Tools) | +2.40 |
General Evaluation
GLM-5.2 1/1| Benchmark | DeepSeek-V4-Pro | GLM-5.2 | Diff |
|---|---|---|---|
| GPQA Diamond | 72.90147 / 226Normal (No Tools) | 91.8626 / 226最高(无工具) | -18.96 |
Specs
| Field | DeepSeek-V4-Pro | GLM-5.2 |
|---|---|---|
| Publisher | DeepSeek-AI | 智谱AI |
| Release date | 2026-08-13 | 2026-06-13 |
| Model type | Reasoning model | Reasoning model |
| Architecture | MoE | MoE |
| Parameters | 1.6T | 753.33B |
| Context length | 1M | 1M |
| Max output | 384K | 128K |
API pricing
Prices use DataLearner records when available; missing fields are not inferred.
| Item | DeepSeek-V4-Pro | GLM-5.2 |
|---|---|---|
| Text input | $0.435 / 1M tokens | $1.4 / 1M tokens |
| Text output | $0.87 / 1M tokens | $4.4 / 1M tokens |
| Cache read | $0.003625 / 1M tokens | $0.26 / 1M tokens |
Summary
- DeepSeek-V4-Proleads in:Coding and Software Engineer (2/3), AI Agent - Tool Usage (1/1), Commonsense Reasoning (1/1)
- GLM-5.2leads in:Math and Reasoning (3/3), General Knowledge (2/2), General Evaluation (1/1)
On average across the 11 shared benchmarks, GLM-5.2 scores 12.23 higher.
Largest single-benchmark gap: IMO-AnswerBench — DeepSeek-V4-Pro 35.30 vs GLM-5.2 91 (-55.70).
Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.