DataLearner logo

Claude Opus 4.6 Benchmark Details

Claude Opus 4.6 currently shows benchmark results led by τ²-Bench (1 / 43, score 91.89), IF Bench (1 / 30, score 94), HumanEval (2 / 39, score 95). This page also compares it with 2 competitor models and 3 predecessor or same-series models, including performance and pricing views when available. 6 source links are attached for reference.

Benchmark Results

Claude Opus 4.6

Benchmark Results

Thinking
Tool usage
Internet

General Knowledge

9 evaluations
Benchmark / mode
Score
Rank/total
86
22 / 67
ARC-AGI
Extended
92
12 / 67
GPQA Diamond
Extended
91.31
15 / 187
MMLU
Extended
91.05
7 / 66
76.33
8 / 115
64.60
17 / 61
ARC-AGI-2
Extended
66.30
16 / 61
HLE
ExtendedToolsInternet
53
16 / 170
ARC-AGI-3
Thinking Level · Max
0
3 / 8

Coding and Software Engineer

5 evaluations
Benchmark / mode
Score
Rank/total
HumanEval
Extended
95
2 / 39
SWE-bench Verified
ExtendedTools
80.84
9 / 111
SWE-bench
ExtendedTools
77.83
1 / 2
76
38 / 123
72
13 / 22

Common Sense

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleQA
Extended
72
7 / 47

Math and Reasoning

7 evaluations
Benchmark / mode
Score
Rank/total
AIME2025
Extended
99.79
7 / 107
MATH-500
Extended
97.60
10 / 44
FrontierMath
Thinking Level · Max
40.70
7 / 60
20.80
14 / 80
20.80
14 / 80
14.60
23 / 80
FrontierMath - Tier 4
Thinking Level · Max
22.90
12 / 80

Multimodal Understanding

2 evaluations
Benchmark / mode
Score
Rank/total
MMMU
Extended
73.90
19 / 29
MMMU
ExtendedTools
77.30
16 / 29

Common Sense Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
Simple Bench
Standard Mode
67.60
8 / 63

Agent Level Benchmark

2 evaluations
Benchmark / mode
Score
Rank/total
99.25
2 / 35
τ²-Bench
ExtendedTools
91.89
1 / 43

Instruction Following

1 evaluations
Benchmark / mode
Score
Rank/total
IF Bench
Extended
94
1 / 30

AI Agent - Information Search

1 evaluations
Benchmark / mode
Score
Rank/total
BrowseComp
Thinking EnabledToolsInternet
84
10 / 52

AI Agent - Tool Usage

3 evaluations
Benchmark / mode
Score
Rank/total
MCP-Atlas
Deep Thinking ModeTools
76.80
10 / 27
OSWorld-Verified
ExtendedTools
72.70
11 / 20
Terminal Bench 2.0
ExtendedTools
65.40
11 / 47

Productivity Knowledge

1 evaluations
Benchmark / mode
Score
Rank/total
GDPval-AA
ExtendedToolsInternet
1606
3 / 21

Claw-style Agent Evaluation

1 evaluations
Benchmark / mode
Score
Rank/total
Pinch Bench
Thinking EnabledTools
87.40
7 / 37

Competitor Comparison

Benchmark scores for Claude Opus 4.6 compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkClaude Opus 4.6CurrentGPT-5.4Gemini 3.1 Pro Preview
ARC-AGI
综合评估
92.00Extended Thinking
93.70Standard Mode
--
ARC-AGI-2
综合评估
66.30Extended Thinking
77.10Standard Mode
77.10Thinking Level · High
GPQA Diamond
综合评估
91.31Extended Thinking
--
94.30Thinking Level · High
HLE
综合评估
53.00Extended Thinking | Tools
52.10Thinking Level · Extra High | Tools
51.40Thinking Level · High | Tools
MMLU
综合评估
91.05Extended Thinking
--
92.60Thinking Level · High
LiveCodeBench
编程与软件工程
76.00Extended Thinking
--
91.70Thinking Level · High | Tools
SWE-bench Verified
编程与软件工程
80.84Extended Thinking | Tools
--
80.60Thinking Level · High | Tools
FrontierMath
数学推理
40.70Thinking Level · High
--
36.90Thinking Level · High
22.90Thinking Level · High
27.10Thinking Level · Extra High
16.70Thinking Level · High
MMMU
多模态理解
77.30Extended Thinking | Tools
--
80.50Thinking Level · High
Simple Bench
常识推理
67.60Standard Mode
--
79.60Standard Mode
τ²-Bench
Agent能力评测
91.89Extended Thinking | Tools
--
90.80Thinking Level · High | Tools
6 additional benchmarks remain in the chart above.

Standard API Pricing: Claude Opus 4.6 vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

Claude Opus 4.6: Base price applies to <= 200K
GPT-5.4: Base price applies to <= 272K
Gemini 3.1 Pro Preview: Base price applies to <= 200K
ModelSupplierStandard inputStandard outputBase price applies to
Claude Opus 4.6
Anthropic$5 / 1M tokens$25 / 1M tokens<= 200K
GPT-5.4
OpenAI$2.5 / 1M tokens$15 / 1M tokens<= 272K
Gemini 3.1 Pro Preview
Google Deep Mind$2 / 1M tokens$12 / 1M tokens<= 200K

Version History

How each version of the Claude Opus 4.6 series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkClaude Opus 4.6CurrentOpus 4.5Opus 4.1Claude Opus 4
ARC-AGI
综合评估
92.00Extended Thinking
--
--
35.70Standard Mode
ARC-AGI-2
综合评估
66.30Extended Thinking
--
--
8.60Standard Mode
GPQA Diamond
综合评估
91.31Extended Thinking
--
81.00Extended Thinking
79.60Standard Mode
HLE
综合评估
53.00Extended Thinking | Tools
43.20Extended Thinking | Tools
--
10.70Standard Mode
LiveBench
综合评估
76.33Thinking Level · High
75.9664K
61.8132K
--
LiveCodeBench
编程与软件工程
76.00Extended Thinking
87.00Extended Thinking | Tools
--
56.60Standard Mode
SWE-bench Verified
编程与软件工程
80.84Extended Thinking | Tools
80.90Extended Thinking | Tools
74.50Extended Thinking | Tools
72.50Standard Mode
AIME2025
数学推理
99.79Extended Thinking
--
78.00Extended Thinking
75.50Standard Mode
FrontierMath
数学推理
40.70Thinking Level · High
--
7.20Extended Thinking
4.50Standard Mode
22.90Thinking Level · High
4.20Standard Mode
4.20Extended Thinking
4.20Thinking Enabled
MATH-500
数学推理
97.60Extended Thinking
--
--
98.20Standard Mode
Simple Bench
常识推理
67.60Standard Mode
62.00Extended Thinking
--
58.80Thinking Enabled
6 additional benchmarks remain in the chart above.

Single-Benchmark Version Trend

Viewing: ARC-AGI · 综合评估

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the Claude Opus 4.6 Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

Claude Opus 4.6: Base price applies to <= 200K
ModelSupplierStandard inputStandard outputBase price applies to
Claude Opus 4.6
Anthropic$5 / 1M tokens$25 / 1M tokens<= 200K
Opus 4.5
Facebook AI研究实验室$5 / 1M tokens$25 / 1M tokens
Opus 4.1
Anthropic$15 / 1M tokens$75 / 1M tokens
Claude Opus 4
Anthropic$15 / 1M tokens$75 / 1M tokens

Sources