DataLearner logo

Claude Opus 5 Benchmark Details

Claude Opus 5 currently shows benchmark results led by Context Arena (1 / 126, score 97.72), SWE-bench Verified (1 / 116, score 96), HLE (2 / 191, score 64.70). This page also compares it with 3 competitor models and 4 predecessor or same-series models, including performance and pricing views when available. 1 source link is attached for reference.

Benchmark Results

Claude Opus 5

Benchmark Results

Thinking
Tool usage
Internet

General Knowledge

10 evaluations
Benchmark / mode
Score
Rank/total
ARC-AGI-1
Extra-High
97.50
4 / 92
90.40
3 / 85
HLE
Max
56.30
15 / 191
HLE
MaxTools
64.70
2 / 191

Other

2 evaluations
Benchmark / mode
Score
Rank/total
87.88
73 / 273
93.88
13 / 273

Coding and Software Engineer

8 evaluations
Benchmark / mode
Score
Rank/total
1668.82
3 / 35
1711.88
1 / 35
96
1 / 116
WeirdML v2
Standard ModeTools
86.34
4 / 52
WeirdML v2
HighTools
91.59
1 / 52
89.50
1 / 28
79.20
2 / 61
DeepSWE
MaxTools
68.80
8 / 35

Writing and Creative Capabilities

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
2120.60
3 / 106

Common Sense Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
80.60
5 / 92

AI Agent - Information Search

1 evaluations
Benchmark / mode
Score
Rank/total
BrowseComp
MaxToolsInternet
90.80
3 / 57

Text Embedding

1 evaluations
Benchmark / mode
Score
Rank/total
97.72
1 / 126

AI Agent - Tool Usage

4 evaluations
Benchmark / mode
Score
Rank/total
MCP-Atlas
Extra-HighTools
85.80
3 / 41
OSWorld 2.0
MaxTools
70.57
3 / 10
51.82
3 / 13

Productivity Knowledge

6 evaluations
Benchmark / mode
Score
Rank/total
GDPval-AA v2
MaxTools
1861
1 / 26
AA-Briefcase
HighTools
1606.44
3 / 19
AA-Briefcase
MaxTools
1684.96
2 / 19
AA-Briefcase
Extra-HighTools
1686.73
1 / 19
AA-AnalystAgent
MaxToolsInternet
53.75
2 / 12
26
12 / 15

Math and Reasoning

2 evaluations
Benchmark / mode
Score
Rank/total
85.61
5 / 58

Competitor Comparison

Benchmark scores for Claude Opus 5 compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkClaude Opus 5CurrentGPT-5.6 SolKimi K3GLM-5.2
AA Intelligence Index (historical versions)
Historical index (not comparable)
综合评估
63.00Thinking Level · Extra High | Tools
61.00Thinking Level · High | Tools
--
--
ARC-AGI-1
Score (%)
综合评估
97.50Thinking Level · Extra High
97.50Thinking Level · Extra High
--
--
ARC-AGI-2
Accuracy and cost per task
综合评估
90.40Thinking Level · High
92.50Thinking Level · High
--
--
ARC-AGI-3 (Standard harness)
Action efficiency score(以 ARC Prize Standard harness 口径为准)
综合评估
30.20Thinking Level · High
7.80Thinking Level · High
--
--
HLE
Accuracy
综合评估
64.70Thinking Level · High | Tools
49.50Thinking Level · High
56.00Thinking Level · High | Tools
54.70Thinking Enabled | Tools
GPQA Diamond
Accuracy
科学与综合推理
93.88Thinking Level · High
94.60Thinking Level · High
93.50Thinking Level · High
91.86Thinking Level · High
DeepSWE
Pass@1 (DeepSWE v1.1)
编程与软件工程
68.80Thinking Level · High | Tools
72.70Thinking Level · Extra High | Tools
67.50Thinking Level · High | Tools
44.00Deep Thinking Mode | Tools
SWE-Bench Pro - Public
Accuracy
编程与软件工程
79.20Thinking Level · High | Tools
64.60Thinking Level · Extra High | Tools
--
62.10Thinking Enabled | Tools
Text Arena (Coding)
Arena Score
编程与软件工程
1711.88Thinking Level · High
1620.27Thinking Level · Extra High
1681.75Thinking Level · High
1593.25Thinking Level · High
WeirdML v2
Average accuracy across 17 tasks (%)
编程与软件工程
91.59Thinking Level · High | Tools
88.76Thinking Level · High | Tools
--
67.31Thinking Level · High | Tools
Creative Writing
Elo、大模型评判两两对战
写作和创作
2120.60Standard Mode
1963.40Standard Mode
2070.60Standard Mode
1752.80Standard Mode
SimpleBench
Score (AVG@5)
常识推理
80.60Thinking Level · High
64.80Thinking Level · Extra High
60.70Thinking Level · High
58.80Standard Mode
12 additional benchmarks remain in the chart above.

Standard API Pricing: Claude Opus 5 vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier.

These models use different currencies or billing units, so the page falls back to raw price values instead of a shared bar chart.

Claude Opus 5
Supplier: Anthropic
Standard input: $5 / 1M tokens
Standard output: $25 / 1M tokens
GPT-5.6 Sol
Supplier: OpenAI
Standard input: $4 / 1M tokens
Standard output: $20 / 1M tokens
Kimi K3
Supplier: Moonshot AI
Standard input: ¥20 / 1M tokens
Standard output: ¥100 / 1M tokens
GLM-5.2
Supplier: 智谱AI
Standard input: $1.4 / 1M tokens
Standard output: $4.4 / 1M tokens
ModelSupplierStandard inputStandard outputBase price applies to
Claude Opus 5
Anthropic$5 / 1M tokens$25 / 1M tokens
GPT-5.6 Sol
OpenAI$4 / 1M tokens$20 / 1M tokens
Kimi K3
Moonshot AI¥20 / 1M tokens¥100 / 1M tokens
GLM-5.2
智谱AI$1.4 / 1M tokens$4.4 / 1M tokens

Version History

How each version of the Claude Opus 5 series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkClaude Opus 5CurrentClaude Opus 4.8Opus 4.7Claude Opus 4.6Opus 4.5
ARC-AGI-1
Score (%)
综合评估
97.50Thinking Level · Extra High
--
93.50Thinking Level · High
92.00Extended Thinking
80.00Extended Thinking
ARC-AGI-2
Accuracy and cost per task
综合评估
90.40Thinking Level · High
--
75.80Thinking Level · High
66.30Extended Thinking
37.60Extended Thinking
HLE
Accuracy
综合评估
64.70Thinking Level · High | Tools
57.90Extended Thinking | Tools
54.70Extended Thinking | Tools
53.00Extended Thinking | Tools
43.20Extended Thinking | Tools
GPQA Diamond
Accuracy
科学与综合推理
93.88Thinking Level · High
93.60Thinking Level · High
94.20Extended Thinking
91.31Extended Thinking
87.00Extended Thinking
DeepSWE
Pass@1 (DeepSWE v1.1)
编程与软件工程
68.80Thinking Level · High | Tools
59.00Deep Thinking Mode | Tools
--
--
--
SWE-bench Multilingual
Accuracy
编程与软件工程
89.50Thinking Level · High | Tools
--
--
72.00Extended Thinking | Tools
--
SWE-Bench Pro - Public
Accuracy
编程与软件工程
79.20Thinking Level · High | Tools
69.20Extended Thinking | Tools
64.30Extended Thinking | Tools
--
--
SWE-bench Verified
Accuracy
编程与软件工程
96.00Thinking Level · High | Tools
88.60Extended Thinking | Tools
87.60Extended Thinking | Tools
80.84Extended Thinking | Tools
80.90Extended Thinking | Tools
Text Arena (Coding)
Arena Score
编程与软件工程
1711.88Thinking Level · High
1545.05Standard Mode
1562.39Standard Mode
1555.35Standard Mode
1512.0032K
WeirdML v2
Average accuracy across 17 tasks (%)
编程与软件工程
91.59Thinking Level · High | Tools
82.89Thinking Level · Extra High | Tools
76.40Standard Mode | Tools
77.95Thinking Level · High | Tools
63.7016K | Tools
Creative Writing
Elo、大模型评判两两对战
写作和创作
2120.60Standard Mode
1835.20Standard Mode
1907.10Standard Mode
1804.10Standard Mode
1683.20Standard Mode
SimpleBench
Score (AVG@5)
常识推理
80.60Thinking Level · High
64.80Standard Mode
61.70Standard Mode
67.60Standard Mode
62.00Extended Thinking
7 additional benchmarks remain in the chart above.

Single-Benchmark Version Trend

Viewing: ARC-AGI-1 · 综合评估

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the Claude Opus 5 Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
Claude Opus 5
Anthropic$5 / 1M tokens$25 / 1M tokens
Claude Opus 4.8
Anthropic$5 / 1M tokens$25 / 1M tokens
Opus 4.7
Anthropic$5 / 1M tokens$25 / 1M tokens
Claude Opus 4.6
Anthropic$5 / 1M tokens$25 / 1M tokens
Opus 4.5
Facebook AI研究实验室$5 / 1M tokens$25 / 1M tokens

Sources