DataLearner logo

Opus 4.7 Benchmark Details

Opus 4.7 currently shows benchmark results led by GPQA Diamond (4 / 187, score 94.20), SWE-bench Verified (5 / 111, score 87.60), LiveBench (7 / 115, score 76.91). This page also compares it with 2 competitor models and 3 predecessor or same-series models, including performance and pricing views when available. 3 source links are attached for reference.

Benchmark Results

Opus 4.7

Benchmark Results

Thinking
Tool usage

General Knowledge

17 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Extended
94.20
4 / 187
ARC-AGI
Thinking Level · Low
91
14 / 67
ARC-AGI
Medium
91
14 / 67
93.50
10 / 67
ARC-AGI
Thinking Level · Max
92
12 / 67
MMLU
Standard Mode
91.50
6 / 66
LiveBench
Thinking Level · Low
70.09
39 / 115
LiveBench
Medium
72.31
27 / 115
74.89
18 / 115
LiveBench
Deep Thinking Mode
76.91
7 / 115
ARC-AGI-2
Thinking Level · Low
62.10
18 / 61
ARC-AGI-2
Medium
67.50
15 / 61
68.30
14 / 61
ARC-AGI-2
Thinking Level · Max
75.80
10 / 61
HLE
Extended
46.90
36 / 170
HLE
ExtendedTools
54.70
11 / 170
0
7 / 8

Coding and Software Engineer

2 evaluations
Benchmark / mode
Score
Rank/total
SWE-bench Verified
ExtendedTools
87.60
5 / 111
64.30
6 / 51

Common Sense Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
Simple Bench
Standard Mode
61.70
13 / 63

Math and Reasoning

2 evaluations
Benchmark / mode
Score
Rank/total
FrontierMath
Extra-High
43.80
6 / 60
22.90
12 / 80

AI Agent - Information Search

1 evaluations
Benchmark / mode
Score
Rank/total
BrowseComp
ExtendedTools
79.30
16 / 52

AI Agent - Tool Usage

4 evaluations
Benchmark / mode
Score
Rank/total
MCP-Atlas
Deep Thinking ModeTools
79.10
7 / 27
OSWorld-Verified
ExtendedTools
78
8 / 20
69.70
18 / 25
Terminal Bench 2.0
ExtendedTools
69.40
6 / 47

Competitor Comparison

Benchmark scores for Opus 4.7 compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

11 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkOpus 4.7CurrentGPT-5.4Gemini 3.1 Pro Preview
ARC-AGI
综合评估
92.00Thinking Level · High
93.70Standard Mode
--
ARC-AGI-2
综合评估
75.80Thinking Level · High
77.10Standard Mode
--
HLE
综合评估
54.70Extended Thinking | Tools
52.10Thinking Level · Extra High | Tools
51.40Thinking Level · High | Tools
LiveBench
综合评估
76.91Deep Thinking Mode
80.28Deep Thinking Mode
79.93Thinking Level · High
SWE-Bench Pro - Public
编程与软件工程
64.30Extended Thinking | Tools
--
54.20Thinking Level · High | Tools
SWE-bench Verified
编程与软件工程
87.60Extended Thinking | Tools
--
80.60Thinking Level · High | Tools
22.90Thinking Level · Extra High
27.10Thinking Level · Extra High
--
BrowseComp
AI Agent - 信息收集
79.30Extended Thinking | Tools
82.70Thinking Level · Extra High | Tools
85.90Thinking Level · High | Tools
MCP-Atlas
AI Agent - 工具使用
79.10Deep Thinking Mode | Tools
70.60Thinking Level · Extra High | Tools
--
OSWorld-Verified
AI Agent - 工具使用
78.00Extended Thinking | Tools
75.00Thinking Level · Extra High | Tools
--
Terminal Bench 2.0
AI Agent - 工具使用
69.40Extended Thinking | Tools
75.10Thinking Level · Extra High | Tools
68.50Thinking Level · High | Tools

Standard API Pricing: Opus 4.7 vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

GPT-5.4: Base price applies to <= 272K
Gemini 3.1 Pro Preview: Base price applies to <= 200K
ModelSupplierStandard inputStandard outputBase price applies to
Opus 4.7
Anthropic$5 / 1M tokens$25 / 1M tokens
GPT-5.4
OpenAI$2.5 / 1M tokens$15 / 1M tokens<= 272K
Gemini 3.1 Pro Preview
Google Deep Mind$2 / 1M tokens$12 / 1M tokens<= 200K

Version History

How each version of the Opus 4.7 series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkOpus 4.7CurrentClaude Opus 4.6Opus 4.5Opus 4.1
ARC-AGI
综合评估
92.00Thinking Level · High
92.00Extended Thinking
--
--
ARC-AGI-2
综合评估
75.80Thinking Level · High
66.30Extended Thinking
--
--
GPQA Diamond
综合评估
94.20Extended Thinking
91.31Extended Thinking
--
81.00Extended Thinking
HLE
综合评估
54.70Extended Thinking | Tools
53.00Extended Thinking | Tools
43.20Extended Thinking | Tools
--
LiveBench
综合评估
76.91Deep Thinking Mode
--
75.9664K
61.8132K
SWE-bench Verified
编程与软件工程
87.60Extended Thinking | Tools
80.84Extended Thinking | Tools
80.90Extended Thinking | Tools
74.50Extended Thinking | Tools
Simple Bench
常识推理
61.70Standard Mode
67.60Standard Mode
62.00Extended Thinking
--
FrontierMath
数学推理
43.80Thinking Level · Extra High
40.70Thinking Level · High
--
7.20Extended Thinking
22.90Thinking Level · Extra High
22.90Thinking Level · High
4.2032K
4.20Extended Thinking
BrowseComp
AI Agent - 信息收集
79.30Extended Thinking | Tools
84.00Thinking Enabled | Tools
--
--
MCP-Atlas
AI Agent - 工具使用
79.10Deep Thinking Mode | Tools
76.80Deep Thinking Mode | Tools
69.80Thinking Level · High | Tools
--
OSWorld-Verified
AI Agent - 工具使用
78.00Extended Thinking | Tools
72.70Extended Thinking | Tools
--
--
1 additional benchmarks remain in the chart above.

Single-Benchmark Version Trend

Viewing: ARC-AGI · 综合评估

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the Opus 4.7 Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

Claude Opus 4.6: Base price applies to <= 200K
ModelSupplierStandard inputStandard outputBase price applies to
Opus 4.7
Anthropic$5 / 1M tokens$25 / 1M tokens
Claude Opus 4.6
Anthropic$5 / 1M tokens$25 / 1M tokens<= 200K
Opus 4.5
Facebook AI研究实验室$5 / 1M tokens$25 / 1M tokens
Opus 4.1
Anthropic$15 / 1M tokens$75 / 1M tokens

Sources