DataLearner logo

GLM-5.3-Flash Benchmark Details

GLM-5.3-Flash currently shows benchmark results led by HLE (15 / 183, score 55.30), AutomationBench (1 / 9, score 48.80), GDPval-AA v2 (2 / 15, score 1773). This page also compares it with 3 competitor models and 2 predecessor or same-series models, including performance and pricing views when available. 3 source links are attached for reference.

Benchmark Results

GLM-5.3-Flash

Benchmark Results

Thinking
Tool usage

General Knowledge

1 evaluations
Benchmark / mode
Score
Rank/total
HLE
MaxTools
55.30
15 / 183

AI Agent - Tool Usage

3 evaluations
Benchmark / mode
Score
Rank/total
84.30
11 / 46
78.40
1 / 7
48.80
1 / 9

Coding and Software Engineer

2 evaluations
Benchmark / mode
Score
Rank/total
DeepSWE
MaxTools
63.40
11 / 30
56.30
4 / 10

Agent Level Benchmark

1 evaluations
Benchmark / mode
Score
Rank/total
26.30
8 / 13

Productivity Knowledge

2 evaluations
Benchmark / mode
Score
Rank/total
GDPval-AA v2
MaxTools
1773
2 / 15
62.40
2 / 2

Multimodal Understanding

4 evaluations
Benchmark / mode
Score
Rank/total
CharXiv RQ
MaxTools
89.40
4 / 18
MMVU
Max
80.50
2 / 2
Chartography
MaxTools
78
1 / 2
53.40
5 / 5

Competitor Comparison

Benchmark scores for GLM-5.3-Flash compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

10 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkGLM-5.3-FlashCurrentQwen3.8-Flash-NextDeepSeek-V4-Flash-Vision-ExpGemini 3.7 Flash
HLE
综合评估
55.30Thinking Level · High | Tools
35.90Thinking Level · Extra High
--
--
AutomationBench
AI Agent - 工具使用
48.80Thinking Level · High | Tools
--
25.70Thinking Level · High | Tools
30.40Thinking Enabled | Tools
Terminal-Bench 2.1
AI Agent - 工具使用
84.30Thinking Level · High | Tools
--
83.90Thinking Level · High | Tools
85.80Thinking Enabled | Tools
Toolathlon-Verified
AI Agent - 工具使用
78.40Thinking Level · High | Tools
73.50Thinking Level · Extra High | Tools
--
--
DeepSWE
编程与软件工程
63.40Thinking Level · High | Tools
58.70Thinking Level · Extra High | Tools
59.30Thinking Level · High | Tools
65.30Thinking Level · High | Tools
NL2Repo-Bench
编程与软件工程
56.30Thinking Level · High | Tools
48.10Thinking Level · Extra High | Tools
57.70Thinking Level · High | Tools
--
Agents' Last Exam
Agent能力评测
26.30Thinking Level · High | Tools
24.30Thinking Level · Extra High | Tools
27.30Thinking Level · High | Tools
26.30Thinking Level · Medium | Tools
GDPval-AA v2
生产力知识
1773.00Thinking Level · High | Tools
--
--
1525.00Thinking Enabled
Chartography
多模态理解
78.00Thinking Level · High | Tools
--
64.30Thinking Level · High | Tools
--
CharXiv RQ
多模态理解
89.40Thinking Level · High | Tools
90.60Thinking Level · Extra High | Tools
--
88.70Thinking Level · Medium | Tools

Standard API Pricing: GLM-5.3-Flash vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
GLM-5.3-Flash
智谱AI$0.075 / 1M tokens$0.25 / 1M tokens
Gemini 3.7 Flash
DeepMind$0.75 / 1M tokens$3.75 / 1M tokens

Version History

How each version of the GLM-5.3-Flash series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

1 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkGLM-5.3-FlashCurrentGLM-4.7-Flash
HLE
综合评估
55.30Thinking Level · High | Tools
14.40Thinking Enabled

Single-Benchmark Version Trend

Viewing: HLE · 综合评估

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the GLM-5.3-Flash Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier.

These models use different currencies or billing units, so the page falls back to raw price values instead of a shared bar chart.

GLM-5.3-Flash
Supplier: 智谱AI
Standard input: $0.075 / 1M tokens
Standard output: $0.25 / 1M tokens
GLM-4.7-Flash
Supplier: 智谱AI
Standard input: ¥0 / 1M tokens
Standard output: ¥0 / 1M tokens
GLM-4.6V-Flash
Supplier: 智谱AI
Standard input: ¥0 / 1M tokens
Standard output: ¥0 / 1M tokens
ModelSupplierStandard inputStandard outputBase price applies to
GLM-5.3-Flash
智谱AI$0.075 / 1M tokens$0.25 / 1M tokens
GLM-4.7-Flash
智谱AI¥0 / 1M tokens¥0 / 1M tokens
GLM-4.6V-Flash
智谱AI¥0 / 1M tokens¥0 / 1M tokens

Sources