GLM-5 Benchmark Details
GLM-5 currently shows benchmark results led by τ²-Bench (4 / 43, score 89.70), τ²-Bench - Telecom (5 / 35, score 98), HLE (25 / 173, score 50.40). This page also compares it with 3 competitor models and 4 predecessor or same-series models, including performance and pricing views when available. 2 source links are attached for reference.
Benchmark Results
Benchmark Results
General Knowledge
6 evaluationsCoding and Software Engineer
1 evaluationsAgent Level Benchmark
3 evaluationsMath and Reasoning
3 evaluationsAI Agent - Information Search
2 evaluationsClaw-style Agent Evaluation
2 evaluationsCompetitor Comparison
Benchmark scores for GLM-5 compared against top models in its class
12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.
| Benchmark | GLM-5Current | Kimi K2.5 | MiniMax M2.5 |
|---|---|---|---|
ARC-AGI 综合评估 | 44.70Thinking Enabled | 65.30Thinking Enabled | 63.70Thinking Enabled |
ARC-AGI-2 综合评估 | 4.90Thinking Enabled | 11.80Thinking Enabled | 4.90Thinking Enabled |
GPQA Diamond 综合评估 | 86.00Thinking Enabled | 87.60Thinking Enabled | 85.20Thinking Enabled |
HLE 综合评估 | 50.40Thinking Enabled | Tools | 50.20Thinking Enabled | Tools | 19.40Thinking Enabled |
LiveBench 综合评估 | 68.85Standard Mode | 69.07Thinking Enabled | 60.14Deep Thinking Mode |
SWE-bench Verified 编程与软件工程 | 77.80Thinking Enabled | 76.80Thinking Enabled | Tools | 80.20Thinking Enabled | Tools |
Simple Bench 常识推理 | 53.20Standard Mode | 46.80Thinking Enabled | -- |
τ²-Bench - Telecom Agent能力评测 | 98.00Thinking Enabled | Tools | -- | 97.80Thinking Enabled | Tools |
AIME 2026 数学推理 | 92.70Thinking Enabled | 92.50Thinking Enabled | -- |
2.10Standard Mode | 4.20Standard Mode | -- | |
IMO-AnswerBench 数学推理 | 82.50Thinking Enabled | 81.80Thinking Enabled | -- |
IF Bench 指令跟随 | 72.00Thinking Enabled | Tools | -- | 70.00Thinking Enabled | Tools |
Standard API Pricing: GLM-5 vs. Peer Models
Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.
Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens
| Model | Supplier | Standard input | Standard output | Base price applies to |
|---|---|---|---|---|
GLM-5 | 智谱AI | $1 / 1M tokens | $3.2 / 1M tokens | — |
Kimi K2.5 | Moonshot AI | $0.6 / 1M tokens | $3 / 1M tokens | — |
MiniMax M2.5 | MiniMaxAI | $0.3 / 1M tokens | $2.4 / 1M tokens | — |
Version History
How each version of the GLM-5 series stacks up on benchmark tests
12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.
| Benchmark | GLM-5Current | GLM-4.7 | GLM-4.6 | GLM-4.5 |
|---|---|---|---|---|
GPQA Diamond 综合评估 | 86.00Thinking Enabled | 85.70Thinking Enabled | 82.90Thinking Enabled | Tools | 79.10Thinking Enabled |
HLE 综合评估 | 50.40Thinking Enabled | Tools | 42.80Thinking Enabled | Tools | 30.40Thinking Enabled | Tools | 14.40Thinking Enabled |
LiveBench 综合评估 | 68.85Standard Mode | 58.09Standard Mode | 55.19Standard Mode | -- |
SWE-bench Verified 编程与软件工程 | 77.80Thinking Enabled | 73.80Thinking Enabled | Tools | 68.00Standard Mode | 64.20Thinking Enabled |
Simple Bench 常识推理 | 53.20Standard Mode | 47.70Thinking Enabled | -- | -- |
Terminal Bench Hard Agent能力评测 | 43.00Thinking Enabled | Tools | 33.30Thinking Enabled | Tools | -- | -- |
τ²-Bench Agent能力评测 | 89.70Thinking Enabled | Tools | 87.40Thinking Enabled | Tools | 75.90Thinking Enabled | Tools | -- |
τ²-Bench - Telecom Agent能力评测 | 98.00Thinking Enabled | Tools | -- | 71.00Thinking Enabled | Tools | -- |
AIME 2026 数学推理 | 92.70Thinking Enabled | 92.90Thinking Enabled | -- | -- |
2.10Standard Mode | 2.10Standard Mode | 2.10Standard Mode | -- | |
IF Bench 指令跟随 | 72.00Thinking Enabled | Tools | -- | 43.00Thinking Enabled | -- |
BrowseComp AI Agent - 信息收集 | 75.90Thinking Enabled | Tools | 52.00Thinking Enabled | Tools | 45.10Thinking Enabled | Tools | -- |
Single-Benchmark Version Trend
Viewing: GPQA Diamond · 综合评估
Standard API Pricing Across the GLM-5 Series
Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.
Source: DataLearnerAI. Standard text prices shown here use the default supplier.
These models use different currencies or billing units, so the page falls back to raw price values instead of a shared bar chart.
| Model | Supplier | Standard input | Standard output | Base price applies to |
|---|---|---|---|---|
GLM-5 | 智谱AI | $1 / 1M tokens | $3.2 / 1M tokens | — |
GLM-4.7 | 智谱AI | ¥4 / 1M tokens | ¥16 / 1M tokens | — |
GLM-4.6 | 智谱AI | ¥5 / 1M tokens | ¥5 / 1M tokens | — |
GLM-4.5 | 智谱AI | ¥0.8 / 1M tokens | ¥2 / 1M tokens | — |
GLM4 | 智谱AI | ¥5 / 1M tokens | ¥5 / 1M tokens | — |