DataLearner logo

GLM 5.1 Benchmark Details

GLM 5.1 currently shows benchmark results led by τ²-Bench - Telecom (13 / 264, score 97.70), HLE (43 / 568, score 52.30), IF Bench (24 / 282, score 76.30). This page also compares it with 3 competitor models and 3 predecessor or same-series models, including performance and pricing views when available. 1 source link is attached for reference.

Benchmark Results

GLM 5.1

Benchmark Results

Thinking
Tool usage
Internet

Knowledge Exams

4 evaluations
Benchmark / mode
Score
Rank/total
HLE
Standard Mode
27.90
232 / 568
HLE
Thinking Mode
31
196 / 568
HLE
Thinking Mode
30.10
208 / 568
HLE
Thinking ModeTools
52.30
43 / 568

Scientific Reasoning

3 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Standard Mode
83.90
181 / 461
GPQA Diamond
Thinking Mode
86.20
146 / 461
CritPt
Thinking Mode
4.60
106 / 204

Writing

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
1589.20
40 / 106

Commonsense

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleBench
Standard Mode
55.10
42 / 93

Repository Engineering

1 evaluations
Benchmark / mode
Score
Rank/total
SWE-Bench Pro - Public
Thinking ModeTools
58.40
20 / 62

Service Workflows

3 evaluations
Benchmark / mode
Score
Rank/total
τ²-Bench - Telecom
Standard ModeTools
97.10
17 / 264
τ²-Bench - Telecom
Thinking ModeTools
97.70
13 / 264
τ³-Banking
Thinking ModeTools
13.60
118 / 167

Instruction Following

2 evaluations
Benchmark / mode
Score
Rank/total
IF Bench
Standard Mode
52
148 / 282
IF Bench
Thinking Mode
76.30
24 / 282

Fact Finding

1 evaluations
Benchmark / mode
Score
Rank/total
BrowseComp
Thinking ModeToolsInternet
79.30
20 / 58

Cross-capability Suites

1 evaluations
Benchmark / mode
Score
Rank/total
LiveBench
Standard Mode
70.18
37 / 117

Agentic Development

7 evaluations
Benchmark / mode
Score
Rank/total
Terminal Bench 2.0
Thinking ModeTools
63.50
13 / 48
Terminal-Bench 2.1
Thinking ModeTools
61.80
112 / 199
Terminal-Bench 2.1
Thinking Level · HighTools
58.70
121 / 199
Terminal-Bench 2.1
Thinking Level · MaxTools
58.70
121 / 199
Terminal Bench Hard
Standard ModeTools
35.60
70 / 244
Terminal Bench Hard
Thinking ModeTools
43.20
41 / 244
Terminal-Bench 4.0
Thinking ModeTools
2
76 / 95

Tool Orchestration

3 evaluations
Benchmark / mode
Score
Rank/total
MCP-Atlas
Standard ModeTools
75.60
24 / 44
59.95
33 / 45
Tool Decathlon
Thinking ModeTools
40.70
6 / 10

Memory & Persistence

2 evaluations
Benchmark / mode
Score
Rank/total
Context Arena
Standard Mode
30.29
116 / 126
Context Arena
Thinking Mode
62.05
74 / 126

Mathematics

2 evaluations
Benchmark / mode
Score
Rank/total
AIME 2026
Thinking Mode
95.30
13 / 30
IMO-AnswerBench
Thinking Mode
83.80
14 / 24

Long Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
AA-LCR
Standard Mode
53.30
138 / 174

Scientific Computing

1 evaluations
Benchmark / mode
Score
Rank/total
SciCode
Thinking Mode
44.80
94 / 134

ML Engineering

1 evaluations
Benchmark / mode
Score
Rank/total
WeirdML v2
Standard ModeTools
57.10
32 / 52

Preference Arenas

1 evaluations
Benchmark / mode
Score
Rank/total
Text Arena (Coding)
Standard Mode
1534
13 / 35

Documents & Charts

1 evaluations
Benchmark / mode
Score
Rank/total
GDP.pdf
Thinking Mode
8.40
90 / 122

Honesty & Factuality

1 evaluations
Benchmark / mode
Score
Rank/total
0.85
44 / 47

Capability Indices

1 evaluations
Benchmark / mode
Score
Rank/total
ECI
unknown
149.86
48 / 167

Competitor Comparison

Benchmark scores for GLM 5.1 compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkGLM 5.1CurrentKimi K2.6MiniMax-M2.7DeepSeek-V4-Pro
HLE
Accuracy
Knowledge Exams
52.30Thinking Enabled | Tools
54.00Thinking Enabled | Tools
29.60Thinking Enabled
48.20Thinking Level · Extra High | Tools
CritPt
Score
Scientific Reasoning
4.60Thinking Enabled
8.00Thinking Enabled
0.60Thinking Enabled
12.90Thinking Level · High
GPQA Diamond
Accuracy
Scientific Reasoning
86.20Thinking Enabled
90.50Thinking Enabled
87.00Thinking Enabled
90.50Thinking Level · High
Creative Writing
Elo、大模型评判两两对战
Writing
1589.20Standard Mode
1721.30Standard Mode
--
1552.10Standard Mode
SimpleBench
Score (AVG@5)
Commonsense
55.10Standard Mode
--
--
50.90Standard Mode
SWE-Bench Pro - Public
Accuracy
Repository Engineering
58.40Thinking Enabled | Tools
58.60Thinking Enabled | Tools
56.20Thinking Enabled | Tools
55.40Thinking Level · Extra High | Tools
τ²-Bench - Telecom
Accuracy
Service Workflows
97.70Thinking Enabled | Tools
95.90Thinking Enabled | Tools
85.00Thinking Enabled | Tools
96.20Thinking Level · High | Tools
τ³-Banking
Score
Service Workflows
13.60Thinking Enabled | Tools
23.30Thinking Enabled | Tools
9.90Thinking Enabled | Tools
30.10Thinking Level · High | Tools
IF Bench
Accuracy
Instruction Following
76.30Thinking Enabled
76.00Thinking Enabled
76.00Thinking Enabled | Tools
76.50Thinking Level · High
BrowseComp
Accuracy
Fact Finding
79.30Thinking Enabled | Tools
83.20Thinking Enabled | Tools
--
83.40Thinking Level · Extra High | Tools
LiveBench
Accuracy
Cross-capability Suites
70.18Standard Mode
70.54Thinking Enabled
63.49Deep Thinking Mode
71.57Standard Mode
Terminal Bench 2.0
Accuracy
Agentic Development
63.50Thinking Enabled | Tools
66.70Thinking Enabled | Tools
57.00Thinking Enabled | Tools
67.90Thinking Level · Extra High | Tools
14 additional benchmarks remain in the chart above.

Standard API Pricing: GLM 5.1 vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
GLM 5.1
智谱AI$1.4 / 1M tokens$4.4 / 1M tokens—
Kimi K2.6
Moonshot AI$0.95 / 1M tokens$4 / 1M tokens—
MiniMax-M2.7
MiniMaxAI$0.3 / 1M tokens$1.2 / 1M tokens—
DeepSeek-V4-Pro
DeepSeek-AI$0.435 / 1M tokens$0.87 / 1M tokens—

Version History

How each version of the GLM 5.1 series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkGLM 5.1CurrentGLM-5GLM-4.7GLM-4.6
HLE
Accuracy
Knowledge Exams
52.30Thinking Enabled | Tools
50.40Thinking Enabled | Tools
42.80Thinking Enabled | Tools
30.40Thinking Enabled | Tools
CritPt
Score
Scientific Reasoning
4.60Thinking Enabled
2.00Thinking Enabled
1.70Thinking Enabled
1.10Thinking Enabled
GPQA Diamond
Accuracy
Scientific Reasoning
86.20Thinking Enabled
86.00Thinking Enabled
85.70Thinking Enabled
82.90Thinking Enabled | Tools
Creative Writing
Elo、大模型评判两两对战
Writing
1589.20Standard Mode
1597.50Standard Mode
1410.80Standard Mode
1408.60Standard Mode
SimpleBench
Score (AVG@5)
Commonsense
55.10Standard Mode
53.20Standard Mode
47.70Thinking Enabled
--
SWE-Bench Pro - Public
Accuracy
Repository Engineering
58.40Thinking Enabled | Tools
--
40.60Thinking Enabled | Tools
--
τ²-Bench - Telecom
Accuracy
Service Workflows
97.70Thinking Enabled | Tools
98.00Thinking Enabled | Tools
95.90Thinking Enabled | Tools
76.90Standard Mode | Tools
τ³-Banking
Score
Service Workflows
13.60Thinking Enabled | Tools
9.79Thinking Enabled | Tools
12.20Thinking Enabled | Tools
13.40Thinking Enabled | Tools
IF Bench
Accuracy
Instruction Following
76.30Thinking Enabled
72.30Thinking Enabled
67.90Thinking Enabled
43.00Thinking Enabled
BrowseComp
Accuracy
Fact Finding
79.30Thinking Enabled | Tools
75.90Thinking Enabled | Tools
52.00Thinking Enabled | Tools
45.10Thinking Enabled | Tools
LiveBench
Accuracy
Cross-capability Suites
70.18Standard Mode
68.85Standard Mode
58.09Standard Mode
55.19Standard Mode
Terminal Bench 2.0
Accuracy
Agentic Development
63.50Thinking Enabled | Tools
61.10Thinking Enabled | Tools
41.00Thinking Enabled | Tools
--
8 additional benchmarks remain in the chart above.

Single-Benchmark Version Trend

Viewing: HLE · Knowledge Exams

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the GLM 5.1 Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier.

These models use different currencies or billing units, so the page falls back to raw price values instead of a shared bar chart.

GLM 5.1
Supplier: 智谱AI
Standard input: $1.4 / 1M tokens
Standard output: $4.4 / 1M tokens
GLM-5
Supplier: 智谱AI
Standard input: $1 / 1M tokens
Standard output: $3.2 / 1M tokens
GLM-4.7
Supplier: 智谱AI
Standard input: ¥4 / 1M tokens
Standard output: ¥16 / 1M tokens
GLM-4.6
Supplier: 智谱AI
Standard input: ¥5 / 1M tokens
Standard output: ¥5 / 1M tokens
ModelSupplierStandard inputStandard outputBase price applies to
GLM 5.1
智谱AI$1.4 / 1M tokens$4.4 / 1M tokens—
GLM-5
智谱AI$1 / 1M tokens$3.2 / 1M tokens—
GLM-4.7
智谱AI¥4 / 1M tokens¥16 / 1M tokens—
GLM-4.6
智谱AI¥5 / 1M tokens¥5 / 1M tokens—

Sources