DataLearner logo

Opus 4.1 Benchmark Details

Opus 4.1 currently shows benchmark results led by MMLU Pro (7 / 133, score 88), Terminal-Bench (5 / 35, score 46.50), SimpleBench (19 / 67, score 60). 1 source link is attached for reference.

Benchmark Results

Opus 4.1

Benchmark Results

Thinking
Tool usage

General Knowledge

3 evaluations
Benchmark / mode
Score
Rank/total
MMLU Pro
Extended
88
7 / 133
LiveBench
Standard Mode
54.45
82 / 115
61.81
60 / 115

Other

1 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Extended
81
135 / 270

Coding and Software Engineer

1 evaluations
Benchmark / mode
Score
Rank/total
SWE-bench Verified
ExtendedTools
74.50
41 / 114

Math and Reasoning

9 evaluations
Benchmark / mode
Score
Rank/total
AIME2025
Extended
78
60 / 106
IMO 2024
Standard Mode
18.70
3 / 10
12.63
56 / 58
IMO 2025
Standard Mode
11.70
4 / 9
FrontierMath
Standard Mode
5.90
35 / 60
FrontierMath
Extended
7.20
33 / 60
4.20
40 / 80
4.20
40 / 80

AI Agent - Tool Usage

2 evaluations
Benchmark / mode
Score
Rank/total
46.50
5 / 35
Terminal-Bench
ExtendedTools
43.30
9 / 35

Common Sense Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleBench
Extended
60
19 / 67

Instruction Following

1 evaluations
Benchmark / mode
Score
Rank/total
IF Bench
ExtendedTools
55
27 / 34

Agent Level Benchmark

1 evaluations
Benchmark / mode
Score
Rank/total
Terminal Bench Hard
ExtendedTools
32
9 / 13

Sources