DataLearner logo

Claude Opus 4.6 Benchmark Details

Claude Opus 4.6 currently shows benchmark results led by τ²-Bench (1 / 43, score 91.89), IF Bench (1 / 31, score 94), HumanEval (2 / 39, score 95). This page also compares it with 2 competitor models and 3 predecessor or same-series models, including performance and pricing views when available. 6 source links are attached for reference.

Benchmark Results

Claude Opus 4.6

Benchmark Results

Thinking
Tool usage
Internet

General Knowledge

9 evaluations
Benchmark / mode
Score
Rank/total
86
23 / 68
ARC-AGI
Extended
92
13 / 68
GPQA Diamond
Extended
91.31
15 / 188
MMLU
Extended
91.05
7 / 66
76.33
8 / 115
64.60
18 / 62
ARC-AGI-2
Extended
66.30
17 / 62
HLE
ExtendedToolsInternet
53
18 / 173
ARC-AGI-3
Thinking Level · Max
0
4 / 9

Coding and Software Engineer

5 evaluations
Benchmark / mode
Score
Rank/total
HumanEval
Extended
95
2 / 39
SWE-bench Verified
ExtendedTools
80.84
10 / 113
SWE-bench
ExtendedTools
77.83
1 / 2
76
38 / 123
72
14 / 23

Common Sense

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleQA
Extended
72
7 / 47

Math and Reasoning

7 evaluations
Benchmark / mode
Score
Rank/total
AIME2025
Extended
99.79
7 / 107
MATH-500
Extended
97.60
10 / 44
FrontierMath
Thinking Level · Max
40.70
7 / 60
20.80
14 / 80
20.80
14 / 80
14.60
23 / 80
FrontierMath - Tier 4
Thinking Level · Max
22.90
12 / 80

Multimodal Understanding

2 evaluations
Benchmark / mode
Score
Rank/total
MMMU
Extended
73.90
19 / 29
MMMU
ExtendedTools
77.30
16 / 29

Common Sense Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
Simple Bench
Standard Mode
67.60
8 / 63

Agent Level Benchmark

2 evaluations
Benchmark / mode
Score
Rank/total
99.25
2 / 35
τ²-Bench
ExtendedTools
91.89
1 / 43

Instruction Following

1 evaluations
Benchmark / mode
Score
Rank/total
IF Bench
Extended
94
1 / 31

AI Agent - Information Search

1 evaluations
Benchmark / mode
Score
Rank/total
BrowseComp
Thinking EnabledToolsInternet
84
11 / 53

AI Agent - Tool Usage

3 evaluations
Benchmark / mode
Score
Rank/total
MCP-Atlas
Deep Thinking ModeTools
76.80
10 / 28
OSWorld-Verified
ExtendedTools
72.70
15 / 25
Terminal Bench 2.0
ExtendedTools
65.40
11 / 47

Productivity Knowledge

1 evaluations
Benchmark / mode
Score
Rank/total
GDPval-AA
ExtendedToolsInternet
1606
3 / 21

Claw-style Agent Evaluation

1 evaluations
Benchmark / mode
Score
Rank/total
Pinch Bench
Thinking EnabledTools
87.40
7 / 37

Competitor Comparison

Benchmark scores for Claude Opus 4.6 compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkClaude Opus 4.6CurrentGPT-5.4Gemini 3.1 Pro Preview
ARC-AGI
综合评估
92.00Extended Thinking
93.70Thinking Level · Extra High
--
ARC-AGI-2
综合评估
66.30Extended Thinking
77.10Standard Mode
77.10Thinking Level · High
GPQA Diamond
综合评估
91.31Extended Thinking
92.80Thinking Level · Extra High
94.30Thinking Level · High
HLE
综合评估
53.00Extended Thinking | Tools
52.10Thinking Level · Extra High | Tools
51.40Thinking Level · High | Tools
LiveBench
综合评估
76.33Thinking Level · High
80.28Deep Thinking Mode
79.93Thinking Level · High
MMLU
综合评估
91.05Extended Thinking
--
92.60Thinking Level · High
LiveCodeBench
编程与软件工程
76.00Extended Thinking
--
91.70Thinking Level · High | Tools
SWE-bench Verified
编程与软件工程
80.84Extended Thinking | Tools
--
80.60Thinking Level · High | Tools
FrontierMath
数学推理
40.70Thinking Level · High
47.60Thinking Level · Extra High
36.90Thinking Level · High
22.90Thinking Level · High
27.10Thinking Level · Extra High
16.70Standard Mode
MMMU
多模态理解
77.30Extended Thinking | Tools
--
80.50Thinking Level · High
Simple Bench
常识推理
67.60Standard Mode
--
79.60Standard Mode
7 additional benchmarks remain in the chart above.

Standard API Pricing: Claude Opus 4.6 vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

Claude Opus 4.6: Base price applies to <= 200K
GPT-5.4: Base price applies to <= 272K
Gemini 3.1 Pro Preview: Base price applies to <= 200K
ModelSupplierStandard inputStandard outputBase price applies to
Claude Opus 4.6
Anthropic$5 / 1M tokens$25 / 1M tokens<= 200K
GPT-5.4
OpenAI$2.5 / 1M tokens$15 / 1M tokens<= 272K
Gemini 3.1 Pro Preview
Google Deep Mind$2 / 1M tokens$12 / 1M tokens<= 200K

Version History

How each version of the Claude Opus 4.6 series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkClaude Opus 4.6CurrentOpus 4.5Opus 4.1Claude Opus 4
ARC-AGI
综合评估
92.00Extended Thinking
80.00Extended Thinking
--
35.70Standard Mode
ARC-AGI-2
综合评估
66.30Extended Thinking
37.60Extended Thinking
--
8.60Standard Mode
GPQA Diamond
综合评估
91.31Extended Thinking
87.00Extended Thinking
81.00Extended Thinking
79.60Standard Mode
HLE
综合评估
53.00Extended Thinking | Tools
43.20Extended Thinking | Tools
--
10.70Standard Mode
LiveBench
综合评估
76.33Thinking Level · High
75.9664K
61.8132K
--
LiveCodeBench
编程与软件工程
76.00Extended Thinking
87.00Extended Thinking | Tools
--
56.60Standard Mode
SWE-bench Verified
编程与软件工程
80.84Extended Thinking | Tools
80.90Extended Thinking | Tools
74.50Extended Thinking | Tools
72.50Standard Mode
AIME2025
数学推理
99.79Extended Thinking
--
78.00Extended Thinking
75.50Standard Mode
FrontierMath
数学推理
40.70Thinking Level · High
20.70Extended Thinking
7.20Extended Thinking
4.50Standard Mode
22.90Thinking Level · High
4.20Standard Mode
4.2032K
4.20Thinking Enabled
MATH-500
数学推理
97.60Extended Thinking
--
--
98.20Standard Mode
MMMU
多模态理解
77.30Extended Thinking | Tools
80.70Extended Thinking
--
--
7 additional benchmarks remain in the chart above.

Single-Benchmark Version Trend

Viewing: ARC-AGI · 综合评估

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the Claude Opus 4.6 Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

Claude Opus 4.6: Base price applies to <= 200K
ModelSupplierStandard inputStandard outputBase price applies to
Claude Opus 4.6
Anthropic$5 / 1M tokens$25 / 1M tokens<= 200K
Opus 4.5
Facebook AI研究实验室$5 / 1M tokens$25 / 1M tokens
Opus 4.1
Anthropic$15 / 1M tokens$75 / 1M tokens
Claude Opus 4
Anthropic$15 / 1M tokens$75 / 1M tokens

Sources