DataLearner logo

GLM 5.1 Benchmark Details

GLM 5.1 currently shows benchmark results led by HLE (19 / 173, score 52.30), AIME 2026 (4 / 19, score 95.30), GPQA Diamond (47 / 188, score 86.20). This page also compares it with 3 competitor models and 3 predecessor or same-series models, including performance and pricing views when available. 1 source link is attached for reference.

Benchmark Results

GLM 5.1

Benchmark Results

Thinking
Tool usage
Internet

General Knowledge

4 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Thinking Mode
86.20
47 / 188
LiveBench
Standard Mode
70.18
37 / 115
HLE
Thinking Mode
31
82 / 173
HLE
Thinking ModeTools
52.30
19 / 173

Coding and Software Engineer

1 evaluations
Benchmark / mode
Score
Rank/total
SWE-Bench Pro - Public
Thinking ModeTools
58.40
15 / 55

AI Agent - Information Search

1 evaluations
Benchmark / mode
Score
Rank/total
BrowseComp
Thinking ModeToolsInternet
79.30
17 / 53

AI Agent - Tool Usage

4 evaluations
Benchmark / mode
Score
Rank/total
MCP-Atlas
Standard ModeTools
75.60
11 / 28
Terminal Bench 2.0
Thinking ModeTools
63.50
13 / 47
TerminalBench 2.1
Thinking Level · HighTools
58.70
25 / 29
Tool Decathlon
Thinking ModeTools
40.70
5 / 9

Math and Reasoning

2 evaluations
Benchmark / mode
Score
Rank/total
AIME 2026
Thinking Mode
95.30
4 / 19
IMO-AnswerBench
Thinking Mode
83.80
12 / 21

Competitor Comparison

Benchmark scores for GLM 5.1 compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

10 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkGLM 5.1CurrentKimi K2.6MiniMax-M2.7DeepSeek-V4-Pro
GPQA Diamond
综合评估
86.20Thinking Enabled
90.50Thinking Enabled
87.00Thinking Enabled
90.10Thinking Level · High
HLE
综合评估
52.30Thinking Enabled | Tools
54.00Thinking Enabled | Tools
28.00Thinking Enabled
48.20Thinking Level · Extra High | Tools
LiveBench
综合评估
70.18Standard Mode
72.17Thinking Enabled
63.49Deep Thinking Mode
73.58Standard Mode
SWE-Bench Pro - Public
编程与软件工程
58.40Thinking Enabled | Tools
58.60Thinking Enabled | Tools
56.20Thinking Enabled | Tools
55.40Thinking Level · Extra High | Tools
BrowseComp
AI Agent - 信息收集
79.30Thinking Enabled | Tools
83.20Thinking Enabled | Tools
--
83.40Thinking Level · Extra High | Tools
Terminal Bench 2.0
AI Agent - 工具使用
63.50Thinking Enabled | Tools
66.70Thinking Enabled | Tools
--
67.90Thinking Level · Extra High | Tools
TerminalBench 2.1
AI Agent - 工具使用
58.70Thinking Level · High | Tools
53.56Thinking Enabled
--
--
Tool Decathlon
AI Agent - 工具使用
40.70Thinking Enabled | Tools
50.00Thinking Enabled | Tools
--
--
AIME 2026
数学推理
95.30Thinking Enabled
96.40Thinking Enabled
--
--
IMO-AnswerBench
数学推理
83.80Thinking Enabled
86.00Thinking Enabled
--
89.80Thinking Level · High

Standard API Pricing: GLM 5.1 vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
GLM 5.1
智谱AI$1.4 / 1M tokens$4.4 / 1M tokens
Kimi K2.6
Facebook AI研究实验室$0.95 / 1M tokens$4 / 1M tokens
MiniMax-M2.7
MiniMaxAI$0.3 / 1M tokens$1.2 / 1M tokens
DeepSeek-V4-Pro
DeepSeek-AI$0.435 / 1M tokens$0.87 / 1M tokens

Version History

How each version of the GLM 5.1 series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

9 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkGLM 5.1CurrentGLM-5GLM-4.7GLM-4.6
GPQA Diamond
综合评估
86.20Thinking Enabled
86.00Thinking Enabled
85.70Thinking Enabled
82.90Thinking Enabled | Tools
HLE
综合评估
52.30Thinking Enabled | Tools
50.40Thinking Enabled | Tools
42.80Thinking Enabled | Tools
30.40Thinking Enabled | Tools
LiveBench
综合评估
70.18Standard Mode
68.85Standard Mode
58.09Standard Mode
55.19Standard Mode
SWE-Bench Pro - Public
编程与软件工程
58.40Thinking Enabled | Tools
--
40.60Thinking Enabled | Tools
--
BrowseComp
AI Agent - 信息收集
79.30Thinking Enabled | Tools
75.90Thinking Enabled | Tools
52.00Thinking Enabled | Tools
45.10Thinking Enabled | Tools
MCP-Atlas
AI Agent - 工具使用
75.60Standard Mode | Tools
--
58.10Standard Mode | Tools
--
Terminal Bench 2.0
AI Agent - 工具使用
63.50Thinking Enabled | Tools
61.10Thinking Enabled | Tools
41.00Thinking Enabled | Tools
--
AIME 2026
数学推理
95.30Thinking Enabled
92.70Thinking Enabled
92.90Thinking Enabled
--
IMO-AnswerBench
数学推理
83.80Thinking Enabled
82.50Thinking Enabled
--
--

Single-Benchmark Version Trend

Viewing: GPQA Diamond · 综合评估

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the GLM 5.1 Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier.

These models use different currencies or billing units, so the page falls back to raw price values instead of a shared bar chart.

GLM 5.1
Supplier: 智谱AI
Standard input: $1.4 / 1M tokens
Standard output: $4.4 / 1M tokens
GLM-5
Supplier: 智谱AI
Standard input: $1 / 1M tokens
Standard output: $3.2 / 1M tokens
GLM-4.7
Supplier: 智谱AI
Standard input: ¥4 / 1M tokens
Standard output: ¥16 / 1M tokens
GLM-4.6
Supplier: 智谱AI
Standard input: ¥5 / 1M tokens
Standard output: ¥5 / 1M tokens
ModelSupplierStandard inputStandard outputBase price applies to
GLM 5.1
智谱AI$1.4 / 1M tokens$4.4 / 1M tokens
GLM-5
智谱AI$1 / 1M tokens$3.2 / 1M tokens
GLM-4.7
智谱AI¥4 / 1M tokens¥16 / 1M tokens
GLM-4.6
智谱AI¥5 / 1M tokens¥5 / 1M tokens

Sources