DataLearner logo

Opus 4.5 Benchmark Details

Opus 4.5 currently shows benchmark results led by MMLU Pro (2 / 133, score 90), Terminal Bench Hard (1 / 13, score 44), SWE-bench Verified (9 / 115, score 80.90). 3 source links are attached for reference.

Benchmark Results

Opus 4.5

Benchmark Results

Thinking
Tool usage

General Knowledge

9 evaluations
Benchmark / mode
Score
Rank/total
MMLU Pro
Extended
90
2 / 133
ARC-AGI-1
Extended
80
46 / 92
75.96
11 / 115
LiveBench
Thinking Level · Low
55.77
80 / 115
LiveBench
Thinking Level · Medium
59.10
72 / 115
LiveBench
Thinking Level · High
58.59
74 / 115
HLE
Extended
30.80
95 / 189
HLE
ExtendedTools
43.20
58 / 189
ARC-AGI-2
Extended
37.60
51 / 85

Other

1 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Extended
87
81 / 270

Coding and Software Engineer

6 evaluations
Benchmark / mode
Score
Rank/total
Text Arena (Coding)
Standard Mode
1479
20 / 35
1512
17 / 35
LiveCodeBench
ExtendedTools
87
15 / 127
SWE-bench Verified
ExtendedTools
80.90
9 / 115
WeirdML v2
16KTools
63.70
24 / 52
GSO
Standard ModeTools
26.50
8 / 21

Writing and Creative Capabilities

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
1683.20
25 / 99

Multimodal Understanding

3 evaluations
Benchmark / mode
Score
Rank/total
MMMU
Extended
80.70
10 / 28
GeoBench ACW
Standard Mode
75
10 / 20
VPCT
32K
40
15 / 24

Common Sense Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleBench
Extended
62
24 / 92

Agent Level Benchmark

7 evaluations
Benchmark / mode
Score
Rank/total
METR Time Horizons v1.1
Standard ModeTools
293
6 / 22
288.90
7 / 22
90.70
21 / 35
τ²-Bench
ExtendedTools
81.99
13 / 44
Terminal Bench Hard
ExtendedTools
44
1 / 13
BALROG
Standard ModeTools
43.50
5 / 12
BALROG
64KTools
43
7 / 12

Math and Reasoning

9 evaluations
Benchmark / mode
Score
Rank/total
AIME 2026
Extended
93.30
8 / 20
34.39
39 / 58
IMO-ProofBench Advanced
Thinking Enabled
23.80
7 / 19
FrontierMath
Extended
20.70
17 / 60
4.20
40 / 80
2.10
56 / 80
4.20
40 / 80
4.20
40 / 80

Instruction Following

1 evaluations
Benchmark / mode
Score
Rank/total
IF Bench
ExtendedTools
58
26 / 35

AI Agent - Tool Usage

2 evaluations
Benchmark / mode
Score
Rank/total
MCP-Atlas
Thinking Level · HighTools
69.80
28 / 41
Terminal Bench 2.0
ExtendedTools
59.30
20 / 48

Claw-style Agent Evaluation

2 evaluations
Benchmark / mode
Score
Rank/total
Claw Bench
ExtendedTools
91.50
7 / 29
Pinch Bench
ExtendedTools
87.20
9 / 38

Long Context

1 evaluations
Benchmark / mode
Score
Rank/total
LongBench v2
Standard Mode
64.40
2 / 13

Sources