DataLearner logo

GLM-5.3 Benchmark Details

GLM-5.3 currently shows benchmark results led by HLE (4 / 190, score 62.50), Creative Writing (3 / 99, score 2062.40), Terminal-Bench 2.1 (5 / 49, score 88.20). This page also compares it with 4 competitor models and 4 predecessor or same-series models, including performance and pricing views when available.

Benchmark Results

GLM-5.3

Benchmark Results

Thinking
Tool usage

General Knowledge

1 evaluations
Benchmark / mode
Score
Rank/total
HLE
MaxTools
62.50
4 / 190

Other

1 evaluations
Benchmark / mode
Score
Rank/total
90.91
38 / 271

Writing and Creative Capabilities

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
2062.40
3 / 99

Common Sense Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
66.20
18 / 92

Text Embedding

3 evaluations
Benchmark / mode
Score
Rank/total
78.61
43 / 126
88.53
18 / 126
85.61
23 / 126

AI Agent - Tool Usage

10 evaluations
Benchmark / mode
Score
Rank/total
130
1 / 1
105
1 / 1
88.20
5 / 49
CyberGym
MaxTools
84.50
1 / 5
73
8 / 10
ExploitBench
MaxTools
54.40
2 / 2
48.20
4 / 14
41.82
5 / 13
28.30
4 / 7
8.10
8 / 11

Coding and Software Engineer

6 evaluations
Benchmark / mode
Score
Rank/total
FrontierSWE
MaxTools
78.10
2 / 4
DeepSWE
MaxTools
66.90
12 / 35
58
4 / 12
SWE-Marathon
MaxTools
42.50
2 / 6
39.80
1 / 5
19
6 / 7

Agent Level Benchmark

1 evaluations
Benchmark / mode
Score
Rank/total
28.50
5 / 15

Productivity Knowledge

1 evaluations
Benchmark / mode
Score
Rank/total
GDPval-AA v2
MaxTools
1769
4 / 25

Math and Reasoning

2 evaluations
Benchmark / mode
Score
Rank/total
68.77
16 / 58
29.27
19 / 41

Competitor Comparison

Benchmark scores for GLM-5.3 compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkGLM-5.3CurrentKimi K3Claude Opus 5GPT-5.6 SolDeepSeek-V4-Pro
HLE
综合评估
62.50Thinking Level · High | Tools
56.00Thinking Level · High | Tools
64.70Thinking Level · High | Tools
49.50Thinking Level · High
48.20Thinking Level · Extra High | Tools
GPQA Diamond
科学与综合推理
90.91Thinking Level · High
93.50Thinking Level · High
93.88Thinking Level · High
93.50Thinking Level · High
90.10Thinking Level · High
Creative Writing
写作和创作
2062.40Standard Mode
2070.80Standard Mode
2116.10Standard Mode
1964.10Standard Mode
1552.10Standard Mode
SimpleBench
常识推理
66.20Thinking Level · High
60.70Thinking Level · High
80.60Thinking Level · High
64.80Thinking Level · Extra High
50.90Standard Mode
Context Arena
文本向量检索
88.53Thinking Level · High
71.75Thinking Level · High
97.72Thinking Level · High
97.63Thinking Level · High
76.09Thinking Enabled
AutomationBench
AI Agent - 工具使用
48.20Thinking Level · High | Tools
30.80Thinking Level · High | Tools
26.00Thinking Level · High | Tools
--
31.80Thinking Level · Extra High | Tools
CyberGym
AI Agent - 工具使用
84.50Thinking Level · High | Tools
--
--
--
83.30Thinking Level · Extra High | Tools
Terminal-Bench 2.1
AI Agent - 工具使用
88.20Thinking Level · High | Tools
88.30Thinking Level · High | Tools
--
88.80Thinking Level · High
87.90Thinking Level · Extra High | Tools
Terminal-Bench 3.0
AI Agent - 工具使用
28.30Thinking Level · High | Tools
--
--
34.60Thinking Level · High | Tools
--
Terminal-Bench 4.0
AI Agent - 工具使用
41.82Thinking Level · High | Tools
--
51.82Thinking Level · High | Tools
37.27Thinking Level · High | Tools
--
Terminal-Bench-Science 0.1
AI Agent - 工具使用
8.10Thinking Level · High | Tools
7.10Thinking Level · High | Tools
30.00Thinking Level · High | Tools
22.40Thinking Level · High | Tools
--
Toolathlon-Verified
AI Agent - 工具使用
73.00Thinking Level · High | Tools
76.50Thinking Level · High | Tools
--
--
74.10Thinking Level · Extra High | Tools
10 additional benchmarks remain in the chart above.

Standard API Pricing: GLM-5.3 vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier.

These models use different currencies or billing units, so the page falls back to raw price values instead of a shared bar chart.

GLM-5.3
Supplier: 智谱AI
Standard input: $1.4 / 1M tokens
Standard output: $4.4 / 1M tokens
Kimi K3
Supplier: Moonshot AI
Standard input: ¥20 / 1M tokens
Standard output: ¥100 / 1M tokens
Claude Opus 5
Supplier: Anthropic
Standard input: $5 / 1M tokens
Standard output: $25 / 1M tokens
GPT-5.6 Sol
Supplier: OpenAI
Standard input: $4 / 1M tokens
Standard output: $20 / 1M tokens
DeepSeek-V4-Pro
Supplier: DeepSeek-AI
Standard input: $0.435 / 1M tokens
Standard output: $0.87 / 1M tokens
ModelSupplierStandard inputStandard outputBase price applies to
GLM-5.3
智谱AI$1.4 / 1M tokens$4.4 / 1M tokens
Kimi K3
Moonshot AI¥20 / 1M tokens¥100 / 1M tokens
Claude Opus 5
Anthropic$5 / 1M tokens$25 / 1M tokens
GPT-5.6 Sol
OpenAI$4 / 1M tokens$20 / 1M tokens
DeepSeek-V4-Pro
DeepSeek-AI$0.435 / 1M tokens$0.87 / 1M tokens

Version History

How each version of the GLM-5.3 series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkGLM-5.3CurrentGLM-5.2GLM 5.1GLM-5GLM-4.7
HLE
综合评估
62.50Thinking Level · High | Tools
54.70Thinking Enabled | Tools
52.30Thinking Enabled | Tools
50.40Thinking Enabled | Tools
42.80Thinking Enabled | Tools
GPQA Diamond
科学与综合推理
90.91Thinking Level · High
91.86Thinking Level · High
86.20Thinking Enabled
86.00Thinking Enabled
85.70Thinking Enabled
Creative Writing
写作和创作
2062.40Standard Mode
1750.90Standard Mode
1589.20Standard Mode
1597.60Standard Mode
1410.50Standard Mode
SimpleBench
常识推理
66.20Thinking Level · High
58.80Standard Mode
55.10Standard Mode
53.20Standard Mode
47.70Thinking Enabled
Context Arena
文本向量检索
88.53Thinking Level · High
72.34Thinking Level · High
62.05Thinking Enabled
--
--
Terminal-Bench 2.1
AI Agent - 工具使用
88.20Thinking Level · High | Tools
81.00Thinking Level · High | Tools
58.70Thinking Level · High | Tools
--
--
DeepSWE
编程与软件工程
66.90Thinking Level · High | Tools
44.00Deep Thinking Mode | Tools
--
--
--
FrontierSWE
编程与软件工程
78.10Thinking Level · High | Tools
74.40Thinking Level · High | Tools
--
--
--
NL2Repo-Bench
编程与软件工程
58.00Thinking Level · High | Tools
48.90Thinking Enabled | Tools
--
--
--
PostTrain Bench
编程与软件工程
39.80Thinking Level · High | Tools
34.30Thinking Level · High | Tools
--
--
--
Program Bench
编程与软件工程
19.00Thinking Level · High | Tools
63.70Thinking Enabled | Tools
--
--
--
SWE-Marathon
编程与软件工程
42.50Thinking Level · High | Tools
13.00Thinking Level · High | Tools
--
--
--
2 additional benchmarks remain in the chart above.

Single-Benchmark Version Trend

Viewing: HLE · 综合评估

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the GLM-5.3 Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier.

These models use different currencies or billing units, so the page falls back to raw price values instead of a shared bar chart.

GLM-5.3
Supplier: 智谱AI
Standard input: $1.4 / 1M tokens
Standard output: $4.4 / 1M tokens
GLM-5.2
Supplier: 智谱AI
Standard input: $1.4 / 1M tokens
Standard output: $4.4 / 1M tokens
GLM 5.1
Supplier: 智谱AI
Standard input: $1.4 / 1M tokens
Standard output: $4.4 / 1M tokens
GLM-5
Supplier: 智谱AI
Standard input: $1 / 1M tokens
Standard output: $3.2 / 1M tokens
GLM-4.7
Supplier: 智谱AI
Standard input: ¥4 / 1M tokens
Standard output: ¥16 / 1M tokens
ModelSupplierStandard inputStandard outputBase price applies to
GLM-5.3
智谱AI$1.4 / 1M tokens$4.4 / 1M tokens
GLM-5.2
智谱AI$1.4 / 1M tokens$4.4 / 1M tokens
GLM 5.1
智谱AI$1.4 / 1M tokens$4.4 / 1M tokens
GLM-5
智谱AI$1 / 1M tokens$3.2 / 1M tokens
GLM-4.7
智谱AI¥4 / 1M tokens¥16 / 1M tokens