DataLearner logo

Opus 4.7 Benchmark Details

Opus 4.7 currently shows benchmark results led by GPQA Diamond (6 / 270, score 94.20), GSO (1 / 21, score 44.12), MMLU (6 / 124, score 91.50). This page also compares it with 3 competitor models and 3 predecessor or same-series models, including performance and pricing views when available. 3 source links are attached for reference.

Benchmark Results

Opus 4.7

Benchmark Results

Thinking
Tool usage

General Knowledge

16 evaluations
Benchmark / mode
Score
Rank/total
ARC-AGI-1
Thinking Level · Low
91
29 / 90
ARC-AGI-1
Medium
91
29 / 90
93.50
22 / 90
ARC-AGI-1
Thinking Level · Max
92
26 / 90
MMLU
Standard Mode
91.50
6 / 124
LiveBench
Thinking Level · Low
70.09
39 / 115
LiveBench
Medium
72.31
27 / 115
74.89
18 / 115
LiveBench
Deep Thinking Mode
76.91
7 / 115
ARC-AGI-2
Thinking Level · Low
62.10
37 / 84
ARC-AGI-2
Medium
67.50
30 / 84
68.30
29 / 84
ARC-AGI-2
Thinking Level · Max
75.80
25 / 84
HLE
Extended
46.90
45 / 189
HLE
ExtendedTools
54.70
19 / 189
0
14 / 15

Other

3 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Extended
94.20
6 / 270
GPQA Diamond
Thinking Level · Max
86.36
85 / 270
GPQA Diamond
Extra-High
90.15
46 / 270

Coding and Software Engineer

7 evaluations
Benchmark / mode
Score
Rank/total
Text Arena (Coding)
Standard Mode
1562.39
7 / 35
SWE-bench Verified
ExtendedTools
87.60
6 / 115
WeirdML v2
Standard ModeTools
76.40
13 / 52
WeirdML v2
HighTools
76.40
13 / 52
64.30
9 / 60
GSO
Standard ModeTools
44.12
1 / 21
GSO
HighTools
44.10
2 / 21

Writing and Creative Capabilities

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
1906.40
7 / 99

Common Sense Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleBench
Standard Mode
62.90
12 / 67

Math and Reasoning

4 evaluations
Benchmark / mode
Score
Rank/total
FrontierMath v2
Thinking Level · Max
70.18
15 / 58
FrontierMath
Extra-High
43.80
6 / 60
FrontierMath Tier 4 v2
Thinking Level · Max
31.71
15 / 40
22.90
12 / 80

AI Agent - Information Search

1 evaluations
Benchmark / mode
Score
Rank/total
BrowseComp
ExtendedTools
79.30
17 / 55

AI Agent - Tool Usage

6 evaluations
Benchmark / mode
Score
Rank/total
MCP-Atlas
Thinking Level · MaxTools
79.10
11 / 41
MCP-Atlas
Deep Thinking ModeTools
79.10
11 / 41
OSWorld-Verified
ExtendedTools
78
11 / 26
69.70
37 / 49
Terminal-Bench 2.1
Thinking Level · MaxTools
68.90
38 / 49
Terminal Bench 2.0
ExtendedTools
69.40
6 / 48

Text Embedding

5 evaluations
Benchmark / mode
Score
Rank/total
Context Arena
Standard Mode
22.12
122 / 126
Context Arena
Thinking Level · Low
46.70
92 / 126
46.15
94 / 126
44.90
96 / 126
Context Arena
Extra-High
45.96
95 / 126

Competitor Comparison

Benchmark scores for Opus 4.7 compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkOpus 4.7CurrentGPT-5.4Gemini 3.1 Pro PreviewGPT-5.5
ARC-AGI-1
综合评估
93.50Thinking Level · High
93.70Thinking Level · Extra High
--
95.00Thinking Level · Extra High
ARC-AGI-2
综合评估
75.80Thinking Level · High
77.10Standard Mode
77.10Thinking Level · High
85.00Thinking Level · Extra High
HLE
综合评估
54.70Extended Thinking | Tools
52.10Thinking Level · Extra High | Tools
51.40Thinking Level · High | Tools
52.20Thinking Level · High | Tools
LiveBench
综合评估
76.91Deep Thinking Mode
80.28Deep Thinking Mode
79.93Thinking Level · High
80.71Deep Thinking Mode
MMLU
综合评估
91.50Standard Mode
--
92.60Thinking Level · High
--
GPQA Diamond
科学与综合推理
94.20Extended Thinking
92.80Thinking Level · Extra High
94.30Thinking Level · High
94.00Thinking Level · Extra High
GSO
编程与软件工程
44.12Standard Mode | Tools
31.37Thinking Level · Extra High | Tools
22.55Standard Mode | Tools
40.20Thinking Level · Extra High | Tools
SWE-Bench Pro - Public
编程与软件工程
64.30Extended Thinking | Tools
57.70Thinking Level · Extra High
54.20Thinking Level · High | Tools
58.60Thinking Level · High | Tools
SWE-bench Verified
编程与软件工程
87.60Extended Thinking | Tools
--
80.60Thinking Level · High | Tools
--
Text Arena (Coding)
编程与软件工程
1562.39Standard Mode
1457.23Thinking Level · High
1461.49Standard Mode
1504.74Thinking Level · Extra High
WeirdML v2
编程与软件工程
76.40Standard Mode | Tools
77.70Thinking Level · Extra High | Tools
72.10Standard Mode | Tools
84.91Thinking Level · Extra High | Tools
Creative Writing
写作和创作
1906.40Standard Mode
1835.50Standard Mode
1488.80Standard Mode
1844.00Standard Mode
11 additional benchmarks remain in the chart above.

Standard API Pricing: Opus 4.7 vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

Gemini 3.1 Pro Preview: Base price applies to <= 200K
ModelSupplierStandard inputStandard outputBase price applies to
Opus 4.7
Anthropic$5 / 1M tokens$25 / 1M tokens
GPT-5.4
OpenAI$2.5 / 1M tokens$15 / 1M tokens
Gemini 3.1 Pro Preview
Google Deep Mind$2 / 1M tokens$12 / 1M tokens<= 200K
GPT-5.5
OpenAI$5 / 1M tokens$30 / 1M tokens

Version History

How each version of the Opus 4.7 series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkOpus 4.7CurrentClaude Opus 4.6Opus 4.5Opus 4.1
ARC-AGI-1
综合评估
93.50Thinking Level · High
92.00Extended Thinking
80.00Extended Thinking
--
ARC-AGI-2
综合评估
75.80Thinking Level · High
66.30Extended Thinking
37.60Extended Thinking
--
HLE
综合评估
54.70Extended Thinking | Tools
53.00Extended Thinking | Tools
43.20Extended Thinking | Tools
--
LiveBench
综合评估
76.91Deep Thinking Mode
76.33Thinking Level · High
75.9664K
61.8132K
MMLU
综合评估
91.50Standard Mode
91.05Extended Thinking
--
--
GPQA Diamond
科学与综合推理
94.20Extended Thinking
91.31Extended Thinking
87.00Extended Thinking
81.00Extended Thinking
GSO
编程与软件工程
44.12Standard Mode | Tools
41.20Thinking Level · High | Tools
26.50Standard Mode | Tools
--
SWE-bench Verified
编程与软件工程
87.60Extended Thinking | Tools
80.84Extended Thinking | Tools
80.90Extended Thinking | Tools
74.50Extended Thinking | Tools
Text Arena (Coding)
编程与软件工程
1562.39Standard Mode
1555.35Standard Mode
1512.0032K
--
WeirdML v2
编程与软件工程
76.40Standard Mode | Tools
77.95Thinking Level · High | Tools
63.7016K | Tools
--
Creative Writing
写作和创作
1906.40Standard Mode
1803.80Standard Mode
1683.20Standard Mode
--
SimpleBench
常识推理
62.90Standard Mode
67.60Standard Mode
62.00Extended Thinking
60.00Extended Thinking
9 additional benchmarks remain in the chart above.

Single-Benchmark Version Trend

Viewing: ARC-AGI-1 · 综合评估

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the Opus 4.7 Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
Opus 4.7
Anthropic$5 / 1M tokens$25 / 1M tokens
Claude Opus 4.6
Anthropic$5 / 1M tokens$25 / 1M tokens
Opus 4.5
Facebook AI研究实验室$5 / 1M tokens$25 / 1M tokens
Opus 4.1
Anthropic$15 / 1M tokens$75 / 1M tokens

Sources