DataLearner logo

Claude Opus 4.6 Benchmark Details

Claude Opus 4.6 currently shows benchmark results led by IF Bench (1 / 282, score 94), τ²-Bench - Telecom (2 / 264, score 99.25), HumanEval (2 / 140, score 95). This page also compares it with 2 competitor models and 3 predecessor or same-series models, including performance and pricing views when available. 6 source links are attached for reference.

Benchmark Results

Claude Opus 4.6

Benchmark Results

Thinking
Tool usage
Internet

Knowledge Exams

6 evaluations
Benchmark / mode
Score
Rank/total
MMLU
Extended
91.05
7 / 124
HLE
Standard Mode
19.10
305 / 568
HLE
Standard Mode
19
307 / 568
HLE
ExtendedToolsInternet
53
40 / 568
HLE
Thinking Level · Max
39.90
137 / 568
HLE
Thinking Level · Max
34.44
183 / 568

Code Generation & Editing

3 evaluations
Benchmark / mode
Score
Rank/total
HumanEval
Extended
95
2 / 140
Vibe Code Bench v1.1
Standard ModeTools
57.57
32 / 60
Vibe Code Bench v1.1
Thinking Level · MaxTools
53.50
33 / 60

Abstract Generalization

11 evaluations
Benchmark / mode
Score
Rank/total
86
81 / 176
ARC-AGI-1
Medium
92
50 / 176
94
38 / 176
ARC-AGI-1
Extended
92
50 / 176
ARC-AGI-1
Thinking Level · Max
93
43 / 176
64.60
60 / 164
ARC-AGI-2
Medium
66.25
58 / 164
69.17
49 / 164
ARC-AGI-2
Extended
66.30
57 / 164
ARC-AGI-2
Thinking Level · Max
68.75
50 / 164
ARC-AGI-3 (Standard harness)
Thinking Level · Max
0.51
19 / 42

Scientific Reasoning

6 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Standard Mode
84
177 / 461
90.53
72 / 461
GPQA Diamond
Extended
91.31
63 / 461
GPQA Diamond
Thinking Level · Max
88.38
114 / 461
CritPt
Standard Mode
2.80
124 / 204
CritPt
Thinking Level · Max
12.60
66 / 204

Honesty & Factuality

3 evaluations
Benchmark / mode
Score
Rank/total
SimpleQA
Extended
72
7 / 46
AA-Omniscience
Standard Mode
2.37
38 / 47
AA-Omniscience
Thinking Enabled
13.67
26 / 47

Repository Engineering

4 evaluations
Benchmark / mode
Score
Rank/total
SWE-bench Verified
ExtendedTools
80.84
10 / 116
SWE-bench
ExtendedTools
77.83
1 / 2
72
19 / 30
SWE-Bench Pro - Commercial
Thinking EnabledTools
47.10
1 / 3

Mathematics

12 evaluations
Benchmark / mode
Score
Rank/total
AIME2025
Extended
99.79
8 / 215
MATH-500
Extended
97.60
11 / 46
AIME 2026
HighTools
96.67
7 / 30
HMMT Feb 2026
HighTools
96.21
4 / 10
FrontierMath v2
Thinking Level · Max
65.96
18 / 58
FrontierMath
Thinking Level · Max
40.70
7 / 60
34.45
10 / 17
FrontierMath Tier 4 v2
Thinking Level · Max
26.83
23 / 42
20.80
14 / 80
20.80
14 / 80
14.60
23 / 80
FrontierMath - Tier 4
Thinking Level · Max
22.90
12 / 80

Algorithmic Coding

1 evaluations
Benchmark / mode
Score
Rank/total
76
68 / 251

Writing

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
1804.10
19 / 106

Visual Understanding

6 evaluations
Benchmark / mode
Score
Rank/total
MMMU
Extended
73.90
38 / 74
MMMU
ExtendedTools
77.30
26 / 74
MMMU-Pro
Standard Mode
72.50
116 / 229
MMMU-Pro
unknown
73.90
102 / 229
MMMU-Pro
Thinking EnabledTools
77.30
67 / 229
MMMU-Pro
Thinking Level · Max
75.40
86 / 229

Commonsense

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleBench
Standard Mode
67.60
19 / 93

Service Workflows

6 evaluations
Benchmark / mode
Score
Rank/total
τ²-Bench - Telecom
Standard ModeTools
84.80
90 / 264
99.25
2 / 264
τ²-Bench - Telecom
Thinking Level · MaxTools
92.10
54 / 264
τ²-Bench
ExtendedTools
91.89
1 / 44
SAGE
Thinking Level · MaxTools
51.58
8 / 64
τ³-Banking
Thinking Level · MaxTools
27.32
70 / 167

Instruction Following

3 evaluations
Benchmark / mode
Score
Rank/total
IF Bench
Standard Mode
44.60
182 / 282
IF Bench
Extended
94
1 / 282
IF Bench
Thinking Level · Max
53.10
145 / 282

Fact Finding

1 evaluations
Benchmark / mode
Score
Rank/total
BrowseComp
Thinking EnabledToolsInternet
84
14 / 58

Cross-capability Suites

1 evaluations
Benchmark / mode
Score
Rank/total
74.52
17 / 117

Agentic Development

3 evaluations
Benchmark / mode
Score
Rank/total
Terminal Bench 2.0
ExtendedTools
65.40
11 / 48
Terminal Bench Hard
Standard ModeTools
48.50
25 / 244
Terminal Bench Hard
Thinking Level · MaxTools
46.20
29 / 244

Memory & Persistence

5 evaluations
Benchmark / mode
Score
Rank/total
Context Arena
Standard Mode
60.33
76 / 126
84.61
26 / 126
85.19
24 / 126
81.95
33 / 126
EBR-bench
Thinking Level · MaxTools
12.70
16 / 23

Long Reasoning

2 evaluations
Benchmark / mode
Score
Rank/total
AA-LCR
Standard Mode
67
120 / 174
AA-LCR
Thinking Level · Max
78
70 / 174

Desktop Workflows

1 evaluations
Benchmark / mode
Score
Rank/total
OSWorld-Verified
ExtendedTools
72.70
18 / 28

Tool Orchestration

4 evaluations
Benchmark / mode
Score
Rank/total
Pinch Bench
Thinking EnabledTools
87.40
8 / 38
MCP-Atlas
Thinking Level · MaxTools
76.80
19 / 44
MCP-Atlas
Deep Thinking ModeTools
76.80
19 / 44
69.90
26 / 45

Maintenance & Optimization

2 evaluations
Benchmark / mode
Score
Rank/total
GSO
Standard ModeTools
33.33
5 / 21
GSO
HighTools
41.20
3 / 21

ML Engineering

2 evaluations
Benchmark / mode
Score
Rank/total
WeirdML v2
Standard ModeTools
65.90
23 / 52
WeirdML v2
HighTools
77.95
11 / 52

Preference Arenas

1 evaluations
Benchmark / mode
Score
Rank/total
Text Arena (Coding)
Standard Mode
1555.35
8 / 35

Clinical Workflows

4 evaluations
Benchmark / mode
Score
Rank/total
MedScribe
Standard ModeTools
86.74
14 / 66
MedScribe
Thinking Level · MaxTools
86.13
16 / 66
MedCode
Standard ModeTools
48.24
21 / 64
MedCode
Thinking Level · MaxTools
49.13
17 / 64

Capability Indices

1 evaluations
Benchmark / mode
Score
Rank/total
ECI
unknown
155.36
25 / 167

Competitor Comparison

Benchmark scores for Claude Opus 4.6 compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkClaude Opus 4.6CurrentGPT-5.4Gemini 3.1 Pro Preview
HLE
Accuracy
Knowledge Exams
53.00Extended Thinking | Tools
52.10Thinking Level · Extra High | Tools
51.40Thinking Level · High | Tools
MMLU
Accuracy
Knowledge Exams
91.05Extended Thinking
--
92.60Thinking Level · High
Vibe Code Bench v1.1
Pass rate (%)
Code Generation & Editing
57.57Standard Mode | Tools
48.47Thinking Level · Extra High | Tools
32.03Thinking Level · High | Tools
ARC-AGI-1
Score (%)
Abstract Generalization
94.00Thinking Level · High
93.67Thinking Level · Extra High
--
ARC-AGI-2
Score (% solved); cost per task (USD)
Abstract Generalization
69.17Thinking Level · High
77.10Standard Mode
77.10Thinking Level · High
ARC-AGI-3 (Standard harness)
Action efficiency score(以 ARC Prize Standard harness 口径为准)
Abstract Generalization
0.51Thinking Level · High
0.21Thinking Level · High
0.42Thinking Level · High
CritPt
Score
Scientific Reasoning
12.60Thinking Level · High
23.40Thinking Level · Extra High
17.70Thinking Enabled
GPQA Diamond
Accuracy
Scientific Reasoning
91.31Extended Thinking
92.00Thinking Level · Extra High
94.30Thinking Level · High
AA-Omniscience
AA-Omniscience Index; Accuracy; Hallucination Rate
Honesty & Factuality
13.67Thinking Enabled
5.78Thinking Level · High
--
SWE-Bench Pro - Commercial
Resolve Rate
Repository Engineering
47.10Thinking Enabled | Tools
43.40Thinking Level · Extra High | Tools
32.20Thinking Enabled | Tools
SWE-bench Verified
Accuracy
Repository Engineering
80.84Extended Thinking | Tools
--
80.60Thinking Level · High | Tools
AIME 2026
Accuracy
Mathematics
96.67Thinking Level · High | Tools
99.17Thinking Level · Extra High | Tools
--
33 additional benchmarks remain in the chart above.

Standard API Pricing: Claude Opus 4.6 vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

Gemini 3.1 Pro Preview: Base price applies to <= 200K
ModelSupplierStandard inputStandard outputBase price applies to
Claude Opus 4.6
Anthropic$5 / 1M tokens$25 / 1M tokens—
GPT-5.4
OpenAI$2.5 / 1M tokens$15 / 1M tokens—
Gemini 3.1 Pro Preview
Google DeepMind$2 / 1M tokens$12 / 1M tokens<= 200K

Version History

How each version of the Claude Opus 4.6 series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkClaude Opus 4.6CurrentOpus 4.5Opus 4.1Claude Opus 4
HLE
Accuracy
Knowledge Exams
53.00Extended Thinking | Tools
43.20Extended Thinking | Tools
12.50Thinking Enabled
12.30Thinking Enabled
Vibe Code Bench v1.1
Pass rate (%)
Code Generation & Editing
57.57Standard Mode | Tools
20.63Thinking Level · High | Tools
--
--
ARC-AGI-1
Score (%)
Abstract Generalization
94.00Thinking Level · High
80.00Extended Thinking
--
35.67Standard Mode
ARC-AGI-2
Score (% solved); cost per task (USD)
Abstract Generalization
69.17Thinking Level · High
37.64Extended Thinking
--
8.61Standard Mode
CritPt
Score
Scientific Reasoning
12.60Thinking Level · High
4.60Thinking Enabled
--
--
GPQA Diamond
Accuracy
Scientific Reasoning
91.31Extended Thinking
87.00Extended Thinking
81.00Extended Thinking
79.60Standard Mode
AA-Omniscience
AA-Omniscience Index; Accuracy; Hallucination Rate
Honesty & Factuality
13.67Thinking Enabled
14.00Thinking Enabled
--
--
SWE-bench Verified
Accuracy
Repository Engineering
80.84Extended Thinking | Tools
80.90Extended Thinking | Tools
74.50Extended Thinking | Tools
72.50Standard Mode
AIME 2026
Accuracy
Mathematics
96.67Thinking Level · High | Tools
93.30Extended Thinking
--
--
AIME2025
Accuracy
Mathematics
99.79Extended Thinking
91.30Thinking Enabled
80.30Thinking Enabled
75.50Standard Mode
FrontierMath
Accuracy
Mathematics
40.70Thinking Level · High
20.70Extended Thinking
7.20Extended Thinking
4.50Standard Mode
FrontierMath - Tier 4
Accuracy
Mathematics
22.90Thinking Level · High
4.20Standard Mode
4.2032K
4.2032K
26 additional benchmarks remain in the chart above.

Single-Benchmark Version Trend

Viewing: HLE · Knowledge Exams

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the Claude Opus 4.6 Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
Claude Opus 4.6
Anthropic$5 / 1M tokens$25 / 1M tokens—
Opus 4.5
Anthropic$5 / 1M tokens$25 / 1M tokens—
Opus 4.1
Anthropic$15 / 1M tokens$75 / 1M tokens—
Claude Opus 4
Anthropic$15 / 1M tokens$75 / 1M tokens—

Sources