DataLearner logo

GLM-5 Benchmark Details

GLM-5 currently shows benchmark results led by τ²-Bench - Telecom (8 / 264, score 98), HLE (51 / 568, score 50.40), τ²-Bench (4 / 44, score 89.70). This page also compares it with 3 competitor models and 4 predecessor or same-series models, including performance and pricing views when available. 2 source links are attached for reference.

Benchmark Results

GLM-5

Benchmark Results

Thinking
Tool usage

Abstract Generalization

2 evaluations
Benchmark / mode
Score
Rank/total
ARC-AGI-1
Thinking Enabled
44.67
134 / 176
ARC-AGI-2
Thinking Enabled
4.86
127 / 164

Knowledge Exams

4 evaluations
Benchmark / mode
Score
Rank/total
HLE
Standard Mode
7.60
433 / 568
HLE
Thinking Enabled
30.50
204 / 568
HLE
Thinking Enabled
29.30
217 / 568
HLE
Thinking EnabledTools
50.40
51 / 568

Scientific Reasoning

3 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Standard Mode
66.60
350 / 461
GPQA Diamond
Thinking Enabled
86
148 / 461
CritPt
Thinking Enabled
2
131 / 204

Repository Engineering

1 evaluations
Benchmark / mode
Score
Rank/total
SWE-bench Verified
Thinking Enabled
77.80
25 / 116

Writing

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
1597.50
39 / 106

Commonsense

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleBench
Standard Mode
53.20
45 / 93

Service Workflows

4 evaluations
Benchmark / mode
Score
Rank/total
τ²-Bench - Telecom
Standard ModeTools
97.40
16 / 264
τ²-Bench - Telecom
Thinking EnabledTools
98
8 / 264
τ²-Bench
Thinking EnabledTools
89.70
4 / 44
τ³-Banking
Thinking EnabledTools
9.79
133 / 167

Mathematics

3 evaluations
Benchmark / mode
Score
Rank/total
AIME 2026
Thinking Enabled
92.70
19 / 30
IMO-AnswerBench
Thinking Enabled
82.50
17 / 24
2.10
56 / 80

Instruction Following

3 evaluations
Benchmark / mode
Score
Rank/total
IF Bench
Standard Mode
55.20
132 / 282
IF Bench
Thinking Enabled
72.30
52 / 282
IF Bench
Thinking EnabledTools
72
54 / 282

Fact Finding

2 evaluations
Benchmark / mode
Score
Rank/total
BrowseComp
Thinking Enabled
62
36 / 58
BrowseComp
Thinking EnabledTools
75.90
27 / 58

Cross-capability Suites

1 evaluations
Benchmark / mode
Score
Rank/total
LiveBench
Standard Mode
68.85
43 / 117

Agentic Development

3 evaluations
Benchmark / mode
Score
Rank/total
Terminal Bench 2.0
Thinking EnabledTools
61.10
18 / 48
Terminal Bench Hard
Standard ModeTools
39.40
53 / 244
Terminal Bench Hard
Thinking EnabledTools
43
44 / 244

Cross-industry Work

1 evaluations
Benchmark / mode
Score
Rank/total
GDPval-AA
Thinking Enabled
46
8 / 15

Long Reasoning

3 evaluations
Benchmark / mode
Score
Rank/total
AA-LCR
Standard Mode
43.70
155 / 174
AA-LCR
Thinking Enabled
75.70
86 / 174
LongBench v2
Standard Mode
60.80
7 / 14

Tool Orchestration

2 evaluations
Benchmark / mode
Score
Rank/total
Claw Bench
Thinking EnabledTools
91.70
5 / 29
Pinch Bench
Thinking EnabledTools
86.40
13 / 38

Honesty & Factuality

1 evaluations
Benchmark / mode
Score
Rank/total
0.27
47 / 47

Code Generation & Editing

1 evaluations
Benchmark / mode
Score
Rank/total
Vibe Code Bench v1.1
Thinking EnabledTools
23.36
44 / 60

Capability Indices

1 evaluations
Benchmark / mode
Score
Rank/total
ECI
unknown
145.84
71 / 167

Competitor Comparison

Benchmark scores for GLM-5 compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkGLM-5CurrentKimi K2.5MiniMax M2.5
ARC-AGI-1
Score (%)
Abstract Generalization
44.67Thinking Enabled
65.33Thinking Enabled
63.67Thinking Enabled
ARC-AGI-2
Score (% solved); cost per task (USD)
Abstract Generalization
4.86Thinking Enabled
11.81Thinking Enabled
4.86Thinking Enabled
HLE
Accuracy
Knowledge Exams
50.40Thinking Enabled | Tools
50.20Thinking Enabled | Tools
20.50Thinking Enabled
CritPt
Score
Scientific Reasoning
2.00Thinking Enabled
3.10Thinking Enabled
1.10Thinking Enabled
GPQA Diamond
Accuracy
Scientific Reasoning
86.00Thinking Enabled
87.60Thinking Enabled
85.20Thinking Enabled
SWE-bench Verified
Accuracy
Repository Engineering
77.80Thinking Enabled
76.80Thinking Enabled | Tools
80.20Thinking Enabled | Tools
Creative Writing
Elo、大模型评判两两对战
Writing
1597.50Standard Mode
1575.80Standard Mode
1358.40Standard Mode
SimpleBench
Score (AVG@5)
Commonsense
53.20Standard Mode
46.80Thinking Enabled
--
τ²-Bench - Telecom
Accuracy
Service Workflows
98.00Thinking Enabled | Tools
95.90Thinking Enabled | Tools
97.80Thinking Enabled | Tools
τ³-Banking
Score
Service Workflows
9.79Thinking Enabled | Tools
14.20Thinking Enabled | Tools
--
AIME 2026
Accuracy
Mathematics
92.70Thinking Enabled
92.50Thinking Enabled
--
FrontierMath - Tier 4
Accuracy
Mathematics
2.10Standard Mode
4.20Standard Mode
--
13 additional benchmarks remain in the chart above.

Standard API Pricing: GLM-5 vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
GLM-5
智谱AI$1 / 1M tokens$3.2 / 1M tokens—
Kimi K2.5
Moonshot AI$0.6 / 1M tokens$3 / 1M tokens—
MiniMax M2.5
MiniMaxAI$0.3 / 1M tokens$1.2 / 1M tokens—

Version History

How each version of the GLM-5 series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkGLM-5CurrentGLM-4.7GLM-4.6GLM-4.5
HLE
Accuracy
Knowledge Exams
50.40Thinking Enabled | Tools
42.80Thinking Enabled | Tools
30.40Thinking Enabled | Tools
14.40Thinking Enabled
CritPt
Score
Scientific Reasoning
2.00Thinking Enabled
1.70Thinking Enabled
1.10Thinking Enabled
--
GPQA Diamond
Accuracy
Scientific Reasoning
86.00Thinking Enabled
85.70Thinking Enabled
82.90Thinking Enabled | Tools
79.10Thinking Enabled
SWE-bench Verified
Accuracy
Repository Engineering
77.80Thinking Enabled
73.80Thinking Enabled | Tools
68.00Standard Mode
64.20Thinking Enabled
Creative Writing
Elo、大模型评判两两对战
Writing
1597.50Standard Mode
1410.80Standard Mode
1408.60Standard Mode
1340.50Standard Mode
SimpleBench
Score (AVG@5)
Commonsense
53.20Standard Mode
47.70Thinking Enabled
--
--
τ²-Bench
Accuracy
Service Workflows
89.70Thinking Enabled | Tools
87.40Thinking Enabled | Tools
75.90Thinking Enabled | Tools
--
τ²-Bench - Telecom
Accuracy
Service Workflows
98.00Thinking Enabled | Tools
95.90Thinking Enabled | Tools
76.90Standard Mode | Tools
43.00Thinking Enabled | Tools
τ³-Banking
Score
Service Workflows
9.79Thinking Enabled | Tools
12.20Thinking Enabled | Tools
13.40Thinking Enabled | Tools
--
AIME 2026
Accuracy
Mathematics
92.70Thinking Enabled
92.90Thinking Enabled
--
--
FrontierMath - Tier 4
Accuracy
Mathematics
2.10Standard Mode
2.10Standard Mode
2.10Standard Mode
--
IF Bench
Accuracy
Instruction Following
72.30Thinking Enabled
67.90Thinking Enabled
43.00Thinking Enabled
44.10Thinking Enabled
6 additional benchmarks remain in the chart above.

Single-Benchmark Version Trend

Viewing: HLE · Knowledge Exams

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the GLM-5 Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier.

These models use different currencies or billing units, so the page falls back to raw price values instead of a shared bar chart.

GLM-5
Supplier: 智谱AI
Standard input: $1 / 1M tokens
Standard output: $3.2 / 1M tokens
GLM-4.7
Supplier: 智谱AI
Standard input: ¥4 / 1M tokens
Standard output: ¥16 / 1M tokens
GLM-4.6
Supplier: 智谱AI
Standard input: ¥5 / 1M tokens
Standard output: ¥5 / 1M tokens
GLM-4.5
Supplier: 智谱AI
Standard input: ¥0.8 / 1M tokens
Standard output: ¥2 / 1M tokens
GLM4
Supplier: 智谱AI
Standard input: ¥5 / 1M tokens
Standard output: ¥5 / 1M tokens
ModelSupplierStandard inputStandard outputBase price applies to
GLM-5
智谱AI$1 / 1M tokens$3.2 / 1M tokens—
GLM-4.7
智谱AI¥4 / 1M tokens¥16 / 1M tokens—
GLM-4.6
智谱AI¥5 / 1M tokens¥5 / 1M tokens—
GLM-4.5
智谱AI¥0.8 / 1M tokens¥2 / 1M tokens—
GLM4
智谱AI¥5 / 1M tokens¥5 / 1M tokens—

Sources