DataLearner logo

GLM-5.3 Benchmark Details

GLM-5.3 currently shows benchmark results led by HLE (3 / 176, score 62.50), Terminal-Bench 2.1 (3 / 38, score 88.20), Automation Bench (1 / 4, score 48.20).

Benchmark Results

GLM-5.3

Benchmark Results

Thinking

General Knowledge

1 evaluations
Benchmark / mode
Score
Rank/total
HLE
MaxTools
62.50
3 / 176

AI Agent - Tool Usage

8 evaluations
Benchmark / mode
Score
Rank/total
130
1 / 1
105
1 / 1
88.20
3 / 38
CyberGym
MaxTools
84.50
1 / 2
73
2 / 3
ExploitBench
MaxTools
54.40
1 / 1
48.20
1 / 4
28.30
1 / 2

Coding and Software Engineer

6 evaluations
Benchmark / mode
Score
Rank/total
FrontierSWE
MaxTools
78.10
2 / 2
DeepSWE
MaxTools
66.90
8 / 22
58
1 / 2
SWE-Marathon
MaxTools
42.50
1 / 3
39.80
1 / 2
19
2 / 2

Agent Level Benchmark

1 evaluations
Benchmark / mode
Score
Rank/total
28.50
4 / 6

Productivity Knowledge

1 evaluations
Benchmark / mode
Score
Rank/total
GDPval-AA v2
MaxTools
1769
2 / 8