DataLearner logo

GLM-5.3 Benchmark Details

GLM-5.3 currently shows benchmark results led by HLE (6 / 233, score 62.50), τ³-Banking (5 / 167, score 50.30), Creative Writing (5 / 106, score 2064.10). This page also compares it with 4 competitor models and 4 predecessor or same-series models, including performance and pricing views when available.

Benchmark Results

GLM-5.3

Benchmark Results

Thinking
Tool usage

Knowledge Exams

1 evaluations
Benchmark / mode
Score
Rank/total
HLE
MaxTools
62.50
6 / 233

Writing

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
2064.10
5 / 106

Commonsense

1 evaluations
Benchmark / mode
Score
Rank/total
66.20
23 / 96

Memory & Persistence

3 evaluations
Benchmark / mode
Score
Rank/total
78.61
43 / 126
88.53
18 / 126
85.61
23 / 126

Repository Engineering

6 evaluations
Benchmark / mode
Score
Rank/total
95.60
5 / 10
84.30
7 / 11
FrontierSWE
MaxTools
78.10
2 / 4
DeepSWE
MaxTools
68.96
23 / 91
58
6 / 16
SWE-Marathon
MaxTools
42.50
3 / 7

Capability Indices

2 evaluations
Benchmark / mode
Score
Rank/total
ECI
unknown
155.56
22 / 167
Vals Index
MaxTools
56.97
23 / 42

Tool Orchestration

2 evaluations
Benchmark / mode
Score
Rank/total
73
12 / 14
28.50
12 / 24

Code Generation & Editing

2 evaluations
Benchmark / mode
Score
Rank/total
78.12
19 / 60
19
10 / 15

ML Engineering

1 evaluations
Benchmark / mode
Score
Rank/total
39.80
1 / 5

Office & Business

1 evaluations
Benchmark / mode
Score
Rank/total
48.80
7 / 23

Scientific Reasoning

2 evaluations
Benchmark / mode
Score
Rank/total
22.97
10 / 13
19.10
43 / 204

Scientific Computing

2 evaluations
Benchmark / mode
Score
Rank/total
59
9 / 134
8.10
9 / 12

Service Workflows

2 evaluations
Benchmark / mode
Score
Rank/total
68.54
5 / 26
τ³-Banking
MaxTools
50.30
5 / 167

Finance

3 evaluations
Benchmark / mode
Score
Rank/total
73.09
3 / 20
56.34
30 / 42
55.84
14 / 42

Legal

2 evaluations
Benchmark / mode
Score
Rank/total
49.04
6 / 42

Vulnerability Analysis

1 evaluations
Benchmark / mode
Score
Rank/total
CyberGym
MaxTools
84.50
5 / 11

Mathematics

3 evaluations
Benchmark / mode
Score
Rank/total
68.77
16 / 58
49
19 / 28
29.27
20 / 42

Agentic Development

2 evaluations
Benchmark / mode
Score
Rank/total
41.82
19 / 96
28.30
6 / 11

Exploitation

4 evaluations
Benchmark / mode
Score
Rank/total
130
1 / 1
105
1 / 1
ExploitBench
MaxTools
54.40
3 / 3

Documents & Charts

1 evaluations
Benchmark / mode
Score
Rank/total
11.20
77 / 122

Coding Indices

1 evaluations
Benchmark / mode
Score
Rank/total
53.60
6 / 11

Maintenance & Optimization

1 evaluations
Benchmark / mode
Score
Rank/total
44.22
17 / 43

Algorithmic Coding

1 evaluations
Benchmark / mode
Score
Rank/total
68.44
8 / 26

Clinical Workflows

2 evaluations
Benchmark / mode
Score
Rank/total
MedScribe
MaxTools
88.81
8 / 66
MedCode
MaxTools
42.86
37 / 64

Cross-capability Suites

1 evaluations
Benchmark / mode
Score
Rank/total
SuperCLUE
unknown
71.29
4 / 13

Competitor Comparison

Benchmark scores for GLM-5.3 compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkGLM-5.3CurrentKimi K3Claude Opus 5GPT-5.6 SolDeepSeek-V4-Pro
HLE
Accuracy
Knowledge Exams
62.50Thinking Level · High | Tools
59.80Thinking Level · High | Tools
63.60Thinking Level · High | Tools
44.50Thinking Level · High
48.20Thinking Level · Extra High | Tools
Creative Writing
Elo、大模型评判两两对战
Writing
2064.10Standard Mode
2070.60Standard Mode
2120.60Standard Mode
1963.40Standard Mode
1552.10Standard Mode
SimpleBench
Score (AVG@5)
Commonsense
66.20Thinking Level · High
60.70Thinking Level · High
80.60Thinking Level · High
64.80Thinking Level · Extra High
50.90Standard Mode
Context Arena
Accuracy (8 needles, 4K-128K context)
Memory & Persistence
88.53Thinking Level · High
71.75Thinking Level · High
97.72Thinking Level · High
97.63Thinking Level · High
76.09Thinking Enabled
DeepSWE
Pass@1 (DeepSWE v1.1)
Repository Engineering
68.96Thinking Level · High | Tools
68.51Thinking Level · High | Tools
73.65Thinking Level · High | Tools
72.70Thinking Level · Extra High | Tools
62.83Thinking Level · High | Tools
FrontierSWE
Dominance score
Repository Engineering
78.10Thinking Level · High | Tools
81.20Thinking Level · High | Tools
--
--
--
NL2Repo-Bench
Average test pass rate
Repository Engineering
58.00Thinking Level · High | Tools
58.00Thinking Level · High | Tools
75.30Thinking Level · High | Tools
56.80Thinking Level · High | Tools
61.50Thinking Level · Extra High | Tools
SWE-Bench Pro V2
Resolve Rate (%, pass@1)
Repository Engineering
95.60Thinking Level · High | Tools
97.70Thinking Level · High | Tools
99.40Thinking Level · Extra High | Tools
95.50Thinking Level · Extra High | Tools
--
SWE-Bench Pro V2 Hard
Resolve Rate (%, pass@1)
Repository Engineering
84.30Thinking Level · High | Tools
88.20Thinking Level · High | Tools
98.00Thinking Level · Extra High | Tools
82.40Thinking Level · Extra High | Tools
--
SWE-Marathon
Resolution rate (Pass@1)
Repository Engineering
42.50Thinking Level · High | Tools
42.00Thinking Level · High | Tools
--
--
--
ECI
ECI score (capability index, higher is better)
Capability Indices
155.56Thinking Level · High
157.68Thinking Level · High
162.67Thinking Level · High
161.99Thinking Level · High
155.39Thinking Level · High
Vals Index
跨行业任务准确率综合指数
Capability Indices
56.97Thinking Level · High | Tools
57.81Thinking Level · High | Tools
67.21Thinking Level · High | Tools
72.63Thinking Level · Extra High
52.37Thinking Level · High | Tools
32 additional benchmarks remain in the chart above.

Standard API Pricing: GLM-5.3 vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier.

These models use different currencies or billing units, so the page falls back to raw price values instead of a shared bar chart.

GLM-5.3
Supplier: 智谱AI
Standard input: $1.4 / 1M tokens
Standard output: $4.4 / 1M tokens
Kimi K3
Supplier: Moonshot AI
Standard input: ¥20 / 1M tokens
Standard output: ¥100 / 1M tokens
Claude Opus 5
Supplier: Anthropic
Standard input: $5 / 1M tokens
Standard output: $25 / 1M tokens
GPT-5.6 Sol
Supplier: OpenAI
Standard input: $4 / 1M tokens
Standard output: $20 / 1M tokens
DeepSeek-V4-Pro
Supplier: DeepSeek-AI
Standard input: $0.435 / 1M tokens
Standard output: $0.87 / 1M tokens
ModelSupplierStandard inputStandard outputBase price applies to
GLM-5.3
智谱AI$1.4 / 1M tokens$4.4 / 1M tokens—
Kimi K3
Moonshot AI¥20 / 1M tokens¥100 / 1M tokens—
Claude Opus 5
Anthropic$5 / 1M tokens$25 / 1M tokens—
GPT-5.6 Sol
OpenAI$4 / 1M tokens$20 / 1M tokens—
DeepSeek-V4-Pro
DeepSeek-AI$0.435 / 1M tokens$0.87 / 1M tokens—

Version History

How each version of the GLM-5.3 series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkGLM-5.3CurrentGLM-5.2GLM 5.1GLM-5GLM-4.7
HLE
Accuracy
Knowledge Exams
62.50Thinking Level · High | Tools
54.70Thinking Enabled | Tools
52.30Thinking Enabled | Tools
50.40Thinking Enabled | Tools
42.80Thinking Enabled | Tools
Creative Writing
Elo、大模型评判两两对战
Writing
2064.10Standard Mode
1752.80Standard Mode
1589.20Standard Mode
1597.50Standard Mode
1410.80Standard Mode
SimpleBench
Score (AVG@5)
Commonsense
66.20Thinking Level · High
58.80Standard Mode
55.10Standard Mode
53.20Standard Mode
47.70Thinking Enabled
Context Arena
Accuracy (8 needles, 4K-128K context)
Memory & Persistence
88.53Thinking Level · High
72.34Thinking Level · High
62.05Thinking Enabled
--
--
DeepSWE
Pass@1 (DeepSWE v1.1)
Repository Engineering
68.96Thinking Level · High | Tools
44.00Deep Thinking Mode | Tools
--
--
--
FrontierSWE
Dominance score
Repository Engineering
78.10Thinking Level · High | Tools
74.40Thinking Level · High | Tools
--
--
--
NL2Repo-Bench
Average test pass rate
Repository Engineering
58.00Thinking Level · High | Tools
48.90Thinking Enabled | Tools
--
--
--
SWE-Marathon
Resolution rate (Pass@1)
Repository Engineering
42.50Thinking Level · High | Tools
13.00Thinking Level · High | Tools
--
--
--
ECI
ECI score (capability index, higher is better)
Capability Indices
155.56Thinking Level · High
151.76Thinking Level · High
149.86Thinking Level · High
145.84Thinking Level · High
143.53Thinking Level · High
Vals Index
跨行业任务准确率综合指数
Capability Indices
56.97Thinking Level · High | Tools
53.12Thinking Level · High | Tools
--
--
--
Program Bench
Score
Code Generation & Editing
19.00Thinking Level · High | Tools
63.70Thinking Enabled | Tools
--
--
--
Vibe Code Bench v1.1
Pass rate (%)
Code Generation & Editing
78.12Thinking Level · High | Tools
63.96Thinking Level · High | Tools
--
23.36Thinking Enabled | Tools
--
16 additional benchmarks remain in the chart above.

Single-Benchmark Version Trend

Viewing: HLE · Knowledge Exams

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the GLM-5.3 Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier.

These models use different currencies or billing units, so the page falls back to raw price values instead of a shared bar chart.

GLM-5.3
Supplier: 智谱AI
Standard input: $1.4 / 1M tokens
Standard output: $4.4 / 1M tokens
GLM-5.2
Supplier: 智谱AI
Standard input: $1.4 / 1M tokens
Standard output: $4.4 / 1M tokens
GLM 5.1
Supplier: 智谱AI
Standard input: $1.4 / 1M tokens
Standard output: $4.4 / 1M tokens
GLM-5
Supplier: 智谱AI
Standard input: $1 / 1M tokens
Standard output: $3.2 / 1M tokens
GLM-4.7
Supplier: 智谱AI
Standard input: ¥4 / 1M tokens
Standard output: ¥16 / 1M tokens
ModelSupplierStandard inputStandard outputBase price applies to
GLM-5.3
智谱AI$1.4 / 1M tokens$4.4 / 1M tokens—
GLM-5.2
智谱AI$1.4 / 1M tokens$4.4 / 1M tokens—
GLM 5.1
智谱AI$1.4 / 1M tokens$4.4 / 1M tokens—
GLM-5
智谱AI$1 / 1M tokens$3.2 / 1M tokens—
GLM-4.7
智谱AI¥4 / 1M tokens¥16 / 1M tokens—