DataLearner logo

Opus 4.5 Benchmark Details

Opus 4.5 currently shows benchmark results led by MMLU Pro (2 / 133, score 90), Terminal Bench Hard (1 / 13, score 44), SWE-bench Verified (9 / 114, score 80.90). 3 source links are attached for reference.

Benchmark Results

Opus 4.5

Benchmark Results

Thinking
Tool usage

General Knowledge

9 evaluations
Benchmark / mode
Score
Rank/total
MMLU Pro
Extended
90
2 / 133
ARC-AGI
Extended
80
24 / 68
75.96
11 / 115
LiveBench
Thinking Level · Low
55.77
80 / 115
LiveBench
Thinking Level · Medium
59.10
72 / 115
LiveBench
Thinking Level · High
58.59
74 / 115
HLE
Extended
30.80
88 / 181
HLE
ExtendedTools
43.20
52 / 181
ARC-AGI-2
Extended
37.60
29 / 62

Other

1 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Extended
87
70 / 224

Coding and Software Engineer

6 evaluations
Benchmark / mode
Score
Rank/total
Text Arena (Coding)
Standard Mode
1479
20 / 35
1512
17 / 35
LiveCodeBench
ExtendedTools
87
14 / 126
SWE-bench Verified
ExtendedTools
80.90
9 / 114
WeirdML v2
16KTools
63.70
24 / 52
GSO
Standard ModeTools
26.50
8 / 21

Writing and Creative Capabilities

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
1683.20
25 / 99

Multimodal Understanding

3 evaluations
Benchmark / mode
Score
Rank/total
MMMU
Extended
80.70
10 / 28
GeoBench ACW
Standard Mode
75
10 / 20
VPCT
32K
40
15 / 24

Common Sense Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleBench
Extended
62
14 / 67

Agent Level Benchmark

7 evaluations
Benchmark / mode
Score
Rank/total
METR Time Horizons v1.1
Standard ModeTools
293
6 / 22
288.90
7 / 22
90.70
21 / 35
τ²-Bench
ExtendedTools
81.99
13 / 43
Terminal Bench Hard
ExtendedTools
44
1 / 13
BALROG
Standard ModeTools
43.50
5 / 12
BALROG
64KTools
43
7 / 12

Math and Reasoning

8 evaluations
Benchmark / mode
Score
Rank/total
AIME 2026
Extended
93.30
8 / 19
34.39
30 / 34
FrontierMath
Extended
20.70
17 / 60
4.20
40 / 80
2.10
56 / 80
4.20
40 / 80
4.20
40 / 80

Instruction Following

1 evaluations
Benchmark / mode
Score
Rank/total
IF Bench
ExtendedTools
58
24 / 33

AI Agent - Tool Usage

2 evaluations
Benchmark / mode
Score
Rank/total
MCP-Atlas
Thinking Level · HighTools
69.80
27 / 39
Terminal Bench 2.0
ExtendedTools
59.30
20 / 48

Claw-style Agent Evaluation

2 evaluations
Benchmark / mode
Score
Rank/total
Claw Bench
ExtendedTools
91.50
7 / 29
Pinch Bench
ExtendedTools
87.20
9 / 38

Long Context

1 evaluations
Benchmark / mode
Score
Rank/total
LongBench v2
Standard Mode
64.40
2 / 13

Sources