DataLearner logo

GLM-5.2 Benchmark Details

GLM-5.2 currently shows benchmark results led by τ²-Bench - Telecom (3 / 264, score 99.10), HLE (27 / 563, score 54.70), IMO-AnswerBench (2 / 24, score 91). This page also compares it with 6 competitor models and 3 predecessor or same-series models, including performance and pricing views when available.

Benchmark Results

GLM-5.2

Benchmark Results

Thinking
Tool usage

General Knowledge

7 evaluations
Benchmark / mode
Score
Rank/total
LiveBench
Standard Mode
73.18
21 / 117
HLE
Standard Mode
9.80
392 / 563
HLE
Thinking Enabled
40.50
128 / 563
HLE
Thinking EnabledTools
54.70
27 / 563
HLE
Max
41.10
123 / 563
CritPt
Standard Mode
3.10
110 / 200
20.90
33 / 200

Other

4 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Standard Mode
71.21
310 / 462
GPQA Diamond
Thinking Level · Low
87.88
120 / 462
GPQA Diamond
Thinking Enabled
91.20
63 / 462
91.86
57 / 462

Writing and Creative Capabilities

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
1752.80
22 / 106

Common Sense Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleBench
Standard Mode
58.80
36 / 93

Coding and Software Engineer

12 evaluations
Benchmark / mode
Score
Rank/total
1593.25
5 / 35
FrontierSWE
MaxTools
74.40
3 / 4
WeirdML v2
HighTools
67.31
21 / 52
Program Bench
Thinking EnabledTools
63.70
2 / 11
SWE-Bench Pro - Public
Thinking EnabledTools
62.10
12 / 62
51.20
63 / 130
NL2Repo-Bench
Thinking EnabledTools
48.90
13 / 16
DeepSWE
HighTools
36.28
74 / 85
DeepSWE
MaxTools
43.78
68 / 85
DeepSWE
Deep Thinking ModeTools
44
67 / 85
34.30
5 / 5
SWE-Marathon
MaxTools
13
6 / 6

Agent Level Benchmark

5 evaluations
Benchmark / mode
Score
Rank/total
99.10
3 / 264
50.80
22 / 244
τ³-Banking
Standard ModeTools
16.70
100 / 164
τ³-Banking
MaxTools
34.60
48 / 164
τ³-Banking
Thinking Level · Extra HighTools
37.11
40 / 164

Instruction Following

1 evaluations
Benchmark / mode
Score
Rank/total
73.30
44 / 282

AI Agent - Tool Usage

6 evaluations
Benchmark / mode
Score
Rank/total
Terminal-Bench 2.1
Standard ModeTools
51.70
127 / 192
81
50 / 192
77.90
68 / 192
MCP-Atlas
Thinking EnabledTools
77.80
17 / 43
Tool Decathlon
Thinking EnabledTools
48.20
4 / 10
1
74 / 86

Text Embedding

1 evaluations
Benchmark / mode
Score
Rank/total
72.34
58 / 126

Math and Reasoning

6 evaluations
Benchmark / mode
Score
Rank/total
AIME 2026
Thinking Enabled
99.20
3 / 29
IMO-AnswerBench
Thinking Enabled
91
2 / 24
FrontierMath v2
Standard Mode
42.46
34 / 58
FrontierMath v2
Thinking Level · Low
54.74
28 / 58
59.21
22 / 58
29.27
20 / 42

Long Context

2 evaluations
Benchmark / mode
Score
Rank/total
AA-LCR
Standard Mode
42.30
154 / 170
78.30
63 / 170

Productivity Knowledge

4 evaluations
Benchmark / mode
Score
Rank/total
GDPval-AA v2
Standard ModeTools
1306
61 / 105
GDPval-AA v2
MaxTools
1406
46 / 105
AA-Briefcase
MaxTools
1231
42 / 83
90.97
12 / 43

Multimodal Understanding

1 evaluations
Benchmark / mode
Score
Rank/total
10.40
78 / 118

Claw-style Agent Evaluation

1 evaluations
Benchmark / mode
Score
Rank/total
86.98
5 / 45

Competitor Comparison

Benchmark scores for GLM-5.2 compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkGLM-5.2CurrentKimi K2.6DeepSeek-V4-ProQwen3.7 MaxKimi K2.7 CodeMiniMax M3Claude Opus 4.8
CritPt
Score
综合评估
20.90Thinking Level · High
8.00Thinking Enabled
12.90Thinking Level · High
13.40Thinking Enabled
10.00Thinking Enabled
3.70Thinking Enabled
20.90Thinking Level · High
HLE
Accuracy
综合评估
54.70Thinking Enabled | Tools
54.00Thinking Enabled | Tools
48.20Thinking Level · Extra High | Tools
53.50Thinking Enabled | Tools
35.00Thinking Enabled
39.00Thinking Enabled
57.90Extended Thinking | Tools
LiveBench
Accuracy
综合评估
73.18Standard Mode
70.54Thinking Enabled
71.57Standard Mode
73.14Deep Thinking Mode
68.41Standard Mode
67.26Deep Thinking Mode
78.93Deep Thinking Mode
GPQA Diamond
Accuracy
科学与综合推理
91.86Thinking Level · High
90.50Thinking Enabled
90.50Thinking Level · High
92.40Thinking Level · High
89.60Thinking Enabled
92.90Thinking Enabled
93.60Thinking Level · High
Creative Writing
Elo、大模型评判两两对战
写作和创作
1752.80Standard Mode
1721.30Standard Mode
1552.10Standard Mode
--
--
--
1835.20Standard Mode
SimpleBench
Score (AVG@5)
常识推理
58.80Standard Mode
--
50.90Standard Mode
70.40Standard Mode
57.90Thinking Enabled
45.80Thinking Enabled
64.80Standard Mode
DeepSWE
Pass@1 (DeepSWE v1.1)
编程与软件工程
44.00Deep Thinking Mode | Tools
--
62.83Thinking Level · High | Tools
--
31.00Standard Mode | Tools
--
58.97Deep Thinking Mode | Tools
NL2Repo-Bench
Average test pass rate
编程与软件工程
48.90Thinking Enabled | Tools
--
61.50Thinking Level · Extra High | Tools
--
--
--
--
PostTrain Bench
Score
编程与软件工程
34.30Thinking Level · High | Tools
--
--
--
--
37.00Thinking Enabled | Tools
--
Program Bench
Score
编程与软件工程
63.70Thinking Enabled | Tools
48.30Thinking Enabled | Tools
--
--
53.60Thinking Enabled | Tools
--
--
SciCode
Score
编程与软件工程
51.20Thinking Level · High
51.50Thinking Enabled
50.80Thinking Level · High
53.50Thinking Level · High
47.80Thinking Enabled
45.37Thinking Enabled
54.40Thinking Level · High
SWE-Bench Pro - Public
Accuracy
编程与软件工程
62.10Thinking Enabled | Tools
58.60Thinking Enabled | Tools
55.40Thinking Level · Extra High | Tools
60.60Thinking Enabled | Tools
--
59.00Thinking Enabled | Tools
69.20Extended Thinking | Tools
21 additional benchmarks remain in the chart above.

Standard API Pricing: GLM-5.2 vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier.

These models use different currencies or billing units, so the page falls back to raw price values instead of a shared bar chart.

GLM-5.2
Supplier: 智谱AI
Standard input: $1.4 / 1M tokens
Standard output: $4.4 / 1M tokens
Kimi K2.6
Supplier: Facebook AI研究实验室
Standard input: $0.95 / 1M tokens
Standard output: $4 / 1M tokens
DeepSeek-V4-Pro
Supplier: DeepSeek-AI
Standard input: $0.435 / 1M tokens
Standard output: $0.87 / 1M tokens
Qwen3.7 Max
Supplier: 阿里巴巴
Standard input: ¥12 / 1M tokens
Standard output: ¥36 / 1M tokens
Kimi K2.7 Code
Supplier: Moonshot AI
Standard input: $0.95 / 1M tokens
Standard output: $4 / 1M tokens
MiniMax M3
Supplier: MiniMaxAI
Standard input: ¥2.1 / 1M tokens
Standard output: ¥8.4 / 1M tokens
Claude Opus 4.8
Supplier: Anthropic
Standard input: $5 / 1M tokens
Standard output: $25 / 1M tokens
ModelSupplierStandard inputStandard outputBase price applies to
GLM-5.2
智谱AI$1.4 / 1M tokens$4.4 / 1M tokens
Kimi K2.6
Facebook AI研究实验室$0.95 / 1M tokens$4 / 1M tokens
DeepSeek-V4-Pro
DeepSeek-AI$0.435 / 1M tokens$0.87 / 1M tokens
Qwen3.7 Max
阿里巴巴¥12 / 1M tokens¥36 / 1M tokens
Kimi K2.7 Code
Moonshot AI$0.95 / 1M tokens$4 / 1M tokens
MiniMax M3
MiniMaxAI¥2.1 / 1M tokens¥8.4 / 1M tokens
Claude Opus 4.8
Anthropic$5 / 1M tokens$25 / 1M tokens

Version History

How each version of the GLM-5.2 series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkGLM-5.2CurrentGLM 5.1GLM-5GLM-4.7
CritPt
Score
综合评估
20.90Thinking Level · High
4.60Thinking Enabled
2.00Thinking Enabled
1.70Thinking Enabled
HLE
Accuracy
综合评估
54.70Thinking Enabled | Tools
52.30Thinking Enabled | Tools
50.40Thinking Enabled | Tools
42.80Thinking Enabled | Tools
LiveBench
Accuracy
综合评估
73.18Standard Mode
70.18Standard Mode
68.85Standard Mode
58.09Standard Mode
GPQA Diamond
Accuracy
科学与综合推理
91.86Thinking Level · High
86.20Thinking Enabled
86.00Thinking Enabled
85.70Thinking Enabled
Creative Writing
Elo、大模型评判两两对战
写作和创作
1752.80Standard Mode
1589.20Standard Mode
1597.50Standard Mode
1410.80Standard Mode
SimpleBench
Score (AVG@5)
常识推理
58.80Standard Mode
55.10Standard Mode
53.20Standard Mode
47.70Thinking Enabled
SciCode
Score
编程与软件工程
51.20Thinking Level · High
44.80Thinking Enabled
--
--
SWE-Bench Pro - Public
Accuracy
编程与软件工程
62.10Thinking Enabled | Tools
58.40Thinking Enabled | Tools
--
40.60Thinking Enabled | Tools
Text Arena (Coding)
Arena Score
编程与软件工程
1593.25Thinking Level · High
1534.00Standard Mode
--
--
WeirdML v2
Average accuracy across 17 tasks (%)
编程与软件工程
67.31Thinking Level · High | Tools
57.10Standard Mode | Tools
--
--
Terminal Bench Hard
Accuracy
Agent能力评测
50.80Thinking Level · High | Tools
43.20Thinking Enabled | Tools
43.00Thinking Enabled | Tools
33.30Thinking Enabled | Tools
τ²-Bench - Telecom
Accuracy
Agent能力评测
99.10Thinking Level · High | Tools
97.70Thinking Enabled | Tools
98.00Thinking Enabled | Tools
95.90Thinking Enabled | Tools
12 additional benchmarks remain in the chart above.

Single-Benchmark Version Trend

Viewing: CritPt · 综合评估

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the GLM-5.2 Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier.

These models use different currencies or billing units, so the page falls back to raw price values instead of a shared bar chart.

GLM-5.2
Supplier: 智谱AI
Standard input: $1.4 / 1M tokens
Standard output: $4.4 / 1M tokens
GLM 5.1
Supplier: 智谱AI
Standard input: $1.4 / 1M tokens
Standard output: $4.4 / 1M tokens
GLM-5
Supplier: 智谱AI
Standard input: $1 / 1M tokens
Standard output: $3.2 / 1M tokens
GLM-4.7
Supplier: 智谱AI
Standard input: ¥4 / 1M tokens
Standard output: ¥16 / 1M tokens
ModelSupplierStandard inputStandard outputBase price applies to
GLM-5.2
智谱AI$1.4 / 1M tokens$4.4 / 1M tokens
GLM 5.1
智谱AI$1.4 / 1M tokens$4.4 / 1M tokens
GLM-5
智谱AI$1 / 1M tokens$3.2 / 1M tokens
GLM-4.7
智谱AI¥4 / 1M tokens¥16 / 1M tokens
GLM-5.2 Benchmark Results Analysis & Model Comparisons | DataLearnerAI