DataLearner logo

GLM-5 Benchmark Details

GLM-5 currently shows benchmark results led by τ²-Bench (4 / 43, score 89.70), τ²-Bench - Telecom (5 / 35, score 98), HLE (25 / 173, score 50.40). This page also compares it with 3 competitor models and 4 predecessor or same-series models, including performance and pricing views when available. 2 source links are attached for reference.

Benchmark Results

GLM-5

Benchmark Results

Thinking
Tool usage

General Knowledge

6 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Thinking Enabled
86
48 / 188
LiveBench
Standard Mode
68.85
43 / 115
50.40
25 / 173
HLE
Thinking Enabled
30.50
86 / 173
ARC-AGI
Thinking Enabled
44.70
47 / 68
ARC-AGI-2
Thinking Enabled
4.90
47 / 62

Coding and Software Engineer

1 evaluations
Benchmark / mode
Score
Rank/total
SWE-bench Verified
Thinking Enabled
77.80
25 / 113

Common Sense Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
Simple Bench
Standard Mode
53.20
23 / 63

Agent Level Benchmark

3 evaluations
Benchmark / mode
Score
Rank/total

Math and Reasoning

3 evaluations
Benchmark / mode
Score
Rank/total
AIME 2026
Thinking Enabled
92.70
10 / 19
IMO-AnswerBench
Thinking Enabled
82.50
15 / 21
2.10
56 / 80

Instruction Following

1 evaluations
Benchmark / mode
Score
Rank/total
72
12 / 31

AI Agent - Information Search

2 evaluations
Benchmark / mode
Score
Rank/total
75.90
24 / 53
BrowseComp
Thinking Enabled
62
33 / 53

AI Agent - Tool Usage

1 evaluations
Benchmark / mode
Score
Rank/total
61.10
18 / 47

Productivity Knowledge

1 evaluations
Benchmark / mode
Score
Rank/total
GDPval-AA
Thinking Enabled
46
14 / 21

Long Context

2 evaluations
Benchmark / mode
Score
Rank/total
AA-LCR
Thinking Enabled
63
15 / 16
LongBench v2
Standard Mode
60.80
6 / 11

Claw-style Agent Evaluation

2 evaluations
Benchmark / mode
Score
Rank/total
Claw Bench
Thinking EnabledTools
91.70
5 / 29
Pinch Bench
Thinking EnabledTools
86.40
12 / 37

Competitor Comparison

Benchmark scores for GLM-5 compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkGLM-5CurrentKimi K2.5MiniMax M2.5
ARC-AGI
综合评估
44.70Thinking Enabled
65.30Thinking Enabled
63.70Thinking Enabled
ARC-AGI-2
综合评估
4.90Thinking Enabled
11.80Thinking Enabled
4.90Thinking Enabled
GPQA Diamond
综合评估
86.00Thinking Enabled
87.60Thinking Enabled
85.20Thinking Enabled
HLE
综合评估
50.40Thinking Enabled | Tools
50.20Thinking Enabled | Tools
19.40Thinking Enabled
LiveBench
综合评估
68.85Standard Mode
69.07Thinking Enabled
60.14Deep Thinking Mode
SWE-bench Verified
编程与软件工程
77.80Thinking Enabled
76.80Thinking Enabled | Tools
80.20Thinking Enabled | Tools
Simple Bench
常识推理
53.20Standard Mode
46.80Thinking Enabled
--
τ²-Bench - Telecom
Agent能力评测
98.00Thinking Enabled | Tools
--
97.80Thinking Enabled | Tools
AIME 2026
数学推理
92.70Thinking Enabled
92.50Thinking Enabled
--
2.10Standard Mode
4.20Standard Mode
--
IMO-AnswerBench
数学推理
82.50Thinking Enabled
81.80Thinking Enabled
--
IF Bench
指令跟随
72.00Thinking Enabled | Tools
--
70.00Thinking Enabled | Tools
7 additional benchmarks remain in the chart above.

Standard API Pricing: GLM-5 vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
GLM-5
智谱AI$1 / 1M tokens$3.2 / 1M tokens
Kimi K2.5
Moonshot AI$0.6 / 1M tokens$3 / 1M tokens
MiniMax M2.5
MiniMaxAI$0.3 / 1M tokens$2.4 / 1M tokens

Version History

How each version of the GLM-5 series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkGLM-5CurrentGLM-4.7GLM-4.6GLM-4.5
GPQA Diamond
综合评估
86.00Thinking Enabled
85.70Thinking Enabled
82.90Thinking Enabled | Tools
79.10Thinking Enabled
HLE
综合评估
50.40Thinking Enabled | Tools
42.80Thinking Enabled | Tools
30.40Thinking Enabled | Tools
14.40Thinking Enabled
LiveBench
综合评估
68.85Standard Mode
58.09Standard Mode
55.19Standard Mode
--
SWE-bench Verified
编程与软件工程
77.80Thinking Enabled
73.80Thinking Enabled | Tools
68.00Standard Mode
64.20Thinking Enabled
Simple Bench
常识推理
53.20Standard Mode
47.70Thinking Enabled
--
--
Terminal Bench Hard
Agent能力评测
43.00Thinking Enabled | Tools
33.30Thinking Enabled | Tools
--
--
τ²-Bench
Agent能力评测
89.70Thinking Enabled | Tools
87.40Thinking Enabled | Tools
75.90Thinking Enabled | Tools
--
τ²-Bench - Telecom
Agent能力评测
98.00Thinking Enabled | Tools
--
71.00Thinking Enabled | Tools
--
AIME 2026
数学推理
92.70Thinking Enabled
92.90Thinking Enabled
--
--
2.10Standard Mode
2.10Standard Mode
2.10Standard Mode
--
IF Bench
指令跟随
72.00Thinking Enabled | Tools
--
43.00Thinking Enabled
--
BrowseComp
AI Agent - 信息收集
75.90Thinking Enabled | Tools
52.00Thinking Enabled | Tools
45.10Thinking Enabled | Tools
--
1 additional benchmarks remain in the chart above.

Single-Benchmark Version Trend

Viewing: GPQA Diamond · 综合评估

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the GLM-5 Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier.

These models use different currencies or billing units, so the page falls back to raw price values instead of a shared bar chart.

GLM-5
Supplier: 智谱AI
Standard input: $1 / 1M tokens
Standard output: $3.2 / 1M tokens
GLM-4.7
Supplier: 智谱AI
Standard input: ¥4 / 1M tokens
Standard output: ¥16 / 1M tokens
GLM-4.6
Supplier: 智谱AI
Standard input: ¥5 / 1M tokens
Standard output: ¥5 / 1M tokens
GLM-4.5
Supplier: 智谱AI
Standard input: ¥0.8 / 1M tokens
Standard output: ¥2 / 1M tokens
GLM4
Supplier: 智谱AI
Standard input: ¥5 / 1M tokens
Standard output: ¥5 / 1M tokens
ModelSupplierStandard inputStandard outputBase price applies to
GLM-5
智谱AI$1 / 1M tokens$3.2 / 1M tokens
GLM-4.7
智谱AI¥4 / 1M tokens¥16 / 1M tokens
GLM-4.6
智谱AI¥5 / 1M tokens¥5 / 1M tokens
GLM-4.5
智谱AI¥0.8 / 1M tokens¥2 / 1M tokens
GLM4
智谱AI¥5 / 1M tokens¥5 / 1M tokens

Sources