DataLearner logo

Claude Opus 5 Benchmark Details

Claude Opus 5 currently shows benchmark results led by HLE (1 / 172, score 64.70), SWE-bench Verified (1 / 112, score 96), ARC-AGI (1 / 68, score 97.50). This page also compares it with 3 competitor models and 4 predecessor or same-series models, including performance and pricing views when available. 1 source link is attached for reference.

Benchmark Results

Claude Opus 5

Benchmark Results

Thinking
Tool usage
Internet

General Knowledge

5 evaluations
Benchmark / mode
Score
Rank/total
ARC-AGI
Extra-High
97.50
1 / 68
90.40
1 / 62
HLE
Max
56.30
11 / 172
HLE
MaxTools
64.70
1 / 172
30.20
1 / 9

Coding and Software Engineer

4 evaluations
Benchmark / mode
Score
Rank/total
96
1 / 112
89.50
1 / 23
79.20
2 / 54
DeepSWE
MaxTools
68.80
4 / 19

AI Agent - Information Search

1 evaluations
Benchmark / mode
Score
Rank/total
BrowseComp
MaxToolsInternet
90.80
2 / 53

AI Agent - Tool Usage

2 evaluations
Benchmark / mode
Score
Rank/total
OSWorld 2.0
MaxTools
70.57
1 / 2
26
2 / 2

Productivity Knowledge

2 evaluations
Benchmark / mode
Score
Rank/total
GDPval-AA v2
MaxTools
1861
1 / 5
AA-Briefcase
MaxTools
1720
1 / 2

Competitor Comparison

Benchmark scores for Claude Opus 5 compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

8 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkClaude Opus 5CurrentGPT-5.6 SolKimi K3GLM-5.2
HLE
综合评估
64.70Thinking Level · High | Tools
--
56.00Thinking Level · High | Tools
54.70Thinking Enabled | Tools
DeepSWE
编程与软件工程
68.80Thinking Level · High | Tools
72.70Thinking Level · Extra High | Tools
67.50Thinking Level · High | Tools
44.00Deep Thinking Mode | Tools
SWE-Bench Pro - Public
编程与软件工程
79.20Thinking Level · High | Tools
64.60Thinking Level · Extra High | Tools
--
62.10Thinking Enabled | Tools
BrowseComp
AI Agent - 信息收集
90.80Thinking Level · High | Tools
--
91.20Thinking Level · High | Tools
--
Automation Bench
AI Agent - 工具使用
26.00Thinking Level · High | Tools
--
30.80Thinking Level · High | Tools
--
OSWorld 2.0
AI Agent - 工具使用
70.57Thinking Level · High | Tools
62.60Thinking Level · Extra High | Tools
--
--
AA-Briefcase
生产力知识
1720.00Thinking Level · High | Tools
--
1548.00Thinking Level · High | Tools
--
GDPval-AA v2
生产力知识
1861.00Thinking Level · High | Tools
--
1668.00Thinking Level · High | Tools
--

Standard API Pricing: Claude Opus 5 vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier.

These models use different currencies or billing units, so the page falls back to raw price values instead of a shared bar chart.

Claude Opus 5
Supplier: Anthropic
Standard input: $5 / 1M tokens
Standard output: $25 / 1M tokens
GPT-5.6 Sol
Supplier: OpenAI
Standard input: $5 / 1M tokens
Standard output: $30 / 1M tokens
Kimi K3
Supplier: Moonshot AI
Standard input: ¥20 / 1M tokens
Standard output: ¥100 / 1M tokens
GLM-5.2
Supplier: 智谱AI
Standard input: $1.4 / 1M tokens
Standard output: $4.4 / 1M tokens
ModelSupplierStandard inputStandard outputBase price applies to
Claude Opus 5
Anthropic$5 / 1M tokens$25 / 1M tokens
GPT-5.6 Sol
OpenAI$5 / 1M tokens$30 / 1M tokens
Kimi K3
Moonshot AI¥20 / 1M tokens¥100 / 1M tokens
GLM-5.2
智谱AI$1.4 / 1M tokens$4.4 / 1M tokens

Version History

How each version of the Claude Opus 5 series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

8 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkClaude Opus 5CurrentClaude Opus 4.8Opus 4.7Claude Opus 4.6Opus 4.5
ARC-AGI
综合评估
97.50Thinking Level · Extra High
--
93.50Thinking Level · High
92.00Extended Thinking
80.00Extended Thinking
ARC-AGI-2
综合评估
90.40Thinking Level · High
--
75.80Thinking Level · High
66.30Extended Thinking
37.60Extended Thinking
HLE
综合评估
64.70Thinking Level · High | Tools
57.90Extended Thinking | Tools
54.70Extended Thinking | Tools
53.00Extended Thinking | Tools
43.20Extended Thinking | Tools
DeepSWE
编程与软件工程
68.80Thinking Level · High | Tools
59.00Deep Thinking Mode | Tools
--
--
--
SWE-bench Multilingual
编程与软件工程
89.50Thinking Level · High | Tools
--
--
72.00Extended Thinking | Tools
--
SWE-Bench Pro - Public
编程与软件工程
79.20Thinking Level · High | Tools
69.20Extended Thinking | Tools
64.30Extended Thinking | Tools
--
--
SWE-bench Verified
编程与软件工程
96.00Thinking Level · High | Tools
88.60Extended Thinking | Tools
87.60Extended Thinking | Tools
80.84Extended Thinking | Tools
80.90Extended Thinking | Tools
BrowseComp
AI Agent - 信息收集
90.80Thinking Level · High | Tools
84.30Thinking Level · High | Tools
79.30Extended Thinking | Tools
84.00Thinking Enabled | Tools
--

Single-Benchmark Version Trend

Viewing: ARC-AGI · 综合评估

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the Claude Opus 5 Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

Claude Opus 4.6: Base price applies to <= 200K
ModelSupplierStandard inputStandard outputBase price applies to
Claude Opus 5
Anthropic$5 / 1M tokens$25 / 1M tokens
Claude Opus 4.8
Anthropic$5 / 1M tokens$25 / 1M tokens
Opus 4.7
Anthropic$5 / 1M tokens$25 / 1M tokens
Claude Opus 4.6
Anthropic$5 / 1M tokens$25 / 1M tokens<= 200K
Opus 4.5
Facebook AI研究实验室$5 / 1M tokens$25 / 1M tokens

Sources