DataLearner logo

DeepSeek-V4-ProvsGLM-5.2

Across 11 shared benchmarks, GLM-5.2 leads overall: DeepSeek-V4-Pro wins 4, GLM-5.2 wins 7, with 0 ties and an average score difference of -12.23.

DeepSeek-AI
DeepSeek-V4-Pro

DeepSeek-AI · 2026-08-13 · Reasoning model

智谱AI
GLM-5.2

智谱AI · 2026-06-13 · Reasoning model

DeepSeek-V4-Pro4 wins(36%)(64%)7 winsGLM-5.2

Benchmark scores

Grouped by capability, sorted by largest gap within each. 11 shared benchmarks.

Coding and Software Engineer

DeepSeek-V4-Pro 2/3
BenchmarkDeepSeek-V4-ProGLM-5.2Diff
DeepSWE62.7011 / 27极高强度思考(工具)4421 / 27Deep Thinking (With Tools)+18.70
NL2Repo-Bench61.501 / 8极高强度思考(工具)48.906 / 8Thinking (With Tools)+12.60
SWE-Bench Pro - Public52.1040 / 57Normal (With Tools)62.109 / 57Thinking (With Tools)-10

Math and Reasoning

GLM-5.2 3/3
BenchmarkDeepSeek-V4-ProGLM-5.2Diff
IMO-AnswerBench35.3023 / 23Normal (No Tools)912 / 23Thinking (No Tools)-55.70
FrontierMath Tier 4 v22.4430 / 34最高(无工具)29.2716 / 34最高(无工具)-26.83
FrontierMath v245.2626 / 34最高(无工具)59.2119 / 34最高(无工具)-13.94

General Knowledge

GLM-5.2 2/2
BenchmarkDeepSeek-V4-ProGLM-5.2Diff
HLE7.70165 / 181Normal (No Tools)54.7015 / 181Thinking (With Tools)-47
LiveBench73.5823 / 115Normal (No Tools)76.249 / 115Normal (No Tools)-2.66

AI Agent - Tool Usage

DeepSeek-V4-Pro 1/1
BenchmarkDeepSeek-V4-ProGLM-5.2Diff
Terminal-Bench 2.187.906 / 44极高强度思考(工具)8117 / 44Thinking High (With Tools)+6.90

Commonsense Reasoning

DeepSeek-V4-Pro 1/1
BenchmarkDeepSeek-V4-ProGLM-5.2Diff
SimpleBench61.2016 / 67Normal (No Tools)58.8020 / 67Normal (No Tools)+2.40

General Evaluation

GLM-5.2 1/1
BenchmarkDeepSeek-V4-ProGLM-5.2Diff
GPQA Diamond72.90147 / 226Normal (No Tools)91.8626 / 226最高(无工具)-18.96

Specs

FieldDeepSeek-V4-ProGLM-5.2
PublisherDeepSeek-AI智谱AI
Release date2026-08-132026-06-13
Model typeReasoning modelReasoning model
ArchitectureMoEMoE
Parameters1.6T753.33B
Context length1M1M
Max output384K128K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemDeepSeek-V4-ProGLM-5.2
Text input$0.435 / 1M tokens$1.4 / 1M tokens
Text output$0.87 / 1M tokens$4.4 / 1M tokens
Cache read$0.003625 / 1M tokens$0.26 / 1M tokens

Summary

  • DeepSeek-V4-Proleads in:Coding and Software Engineer (2/3), AI Agent - Tool Usage (1/1), Commonsense Reasoning (1/1)
  • GLM-5.2leads in:Math and Reasoning (3/3), General Knowledge (2/2), General Evaluation (1/1)

On average across the 11 shared benchmarks, GLM-5.2 scores 12.23 higher.

Largest single-benchmark gap: IMO-AnswerBench — DeepSeek-V4-Pro 35.30 vs GLM-5.2 91 (-55.70).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.