DataLearner logo

Claude Opus 4.8 Benchmark Details

Claude Opus 4.8 currently shows benchmark results led by LiveBench (4 / 115, score 78.79), HLE (8 / 185, score 57.90), SWE-bench Verified (5 / 114, score 88.60). This page also compares it with 4 competitor models and 3 predecessor or same-series models, including performance and pricing views when available. 1 source link is attached for reference.

Benchmark Results

Claude Opus 4.8

Benchmark Results

Thinking
Tool usage
Internet

General Knowledge

6 evaluations
Benchmark / mode
Score
Rank/total
71.42
33 / 115
LiveBench
Medium
75.47
12 / 115
77.16
6 / 115
LiveBench
Deep Thinking Mode
78.79
4 / 115
HLE
Extended Thinking
49.80
34 / 185
HLE
Extended ThinkingTools
57.90
8 / 185

Other

3 evaluations
Benchmark / mode
Score
Rank/total
88.38
56 / 226
93.60
11 / 226
91.04
32 / 226

Coding and Software Engineer

7 evaluations
Benchmark / mode
Score
Rank/total
Text Arena (Coding)
Standard Mode
1545.05
9 / 35
SWE-bench Verified
Extended ThinkingTools
88.60
5 / 114
WeirdML v2
Standard ModeTools
70.45
18 / 52
WeirdML v2
MediumTools
76.04
15 / 52
WeirdML v2
Extra-HighTools
82.89
7 / 52
SWE-Bench Pro - Public
Extended ThinkingTools
69.20
4 / 59
DeepSWE
Deep Thinking ModeTools
59
16 / 31

Writing and Creative Capabilities

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
1835.30
13 / 99

Common Sense Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleBench
Standard Mode
64.80
10 / 67

AI Agent - Information Search

1 evaluations
Benchmark / mode
Score
Rank/total
BrowseComp
HighToolsInternet
84.30
9 / 54

Productivity Knowledge

1 evaluations
Benchmark / mode
Score
Rank/total
GDPval-AA
Extended ThinkingTools
1890
1 / 21

AI Agent - Tool Usage

4 evaluations
Benchmark / mode
Score
Rank/total
OSWorld-Verified
Extended ThinkingTools
83.40
4 / 26
MCP-Atlas
MaxTools
82.20
8 / 40
MCP-Atlas
Deep Thinking ModeTools
82.20
8 / 40
78.90
23 / 47

Math and Reasoning

2 evaluations
Benchmark / mode
Score
Rank/total
80
9 / 34

Competitor Comparison

Benchmark scores for Claude Opus 4.8 compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkClaude Opus 4.8CurrentGPT-5.5 ProGemini 3.1 Pro PreviewKimi K2.6GPT-5.5
HLE
综合评估
57.90Extended Thinking | Tools
57.20Thinking Level · Extra High | Tools
51.40Thinking Level · High | Tools
54.00Thinking Enabled | Tools
52.20Thinking Level · High | Tools
LiveBench
综合评估
78.79Deep Thinking Mode
--
79.93Thinking Level · High
72.17Thinking Enabled
80.71Deep Thinking Mode
GPQA Diamond
科学与综合推理
93.60Thinking Level · High
93.92Thinking Level · Extra High
94.30Thinking Level · High
90.50Thinking Enabled
94.00Thinking Level · Extra High
DeepSWE
编程与软件工程
59.00Deep Thinking Mode | Tools
--
12.00Thinking Level · High | Tools
--
67.00Thinking Level · Extra High | Tools
SWE-Bench Pro - Public
编程与软件工程
69.20Extended Thinking | Tools
--
54.20Thinking Level · High | Tools
58.60Thinking Enabled | Tools
58.60Thinking Level · High | Tools
SWE-bench Verified
编程与软件工程
88.60Extended Thinking | Tools
--
80.60Thinking Level · High | Tools
80.20Thinking Enabled | Tools
--
Text Arena (Coding)
编程与软件工程
1545.05Standard Mode
--
1461.49Standard Mode
--
1504.74Thinking Level · Extra High
WeirdML v2
编程与软件工程
82.89Thinking Level · Extra High | Tools
--
72.10Standard Mode | Tools
--
84.91Thinking Level · Extra High | Tools
Creative Writing
写作和创作
1835.30Standard Mode
--
1488.80Standard Mode
1725.10Standard Mode
1844.00Standard Mode
SimpleBench
常识推理
64.80Standard Mode
76.90Standard Mode
79.60Standard Mode
--
69.00Standard Mode
BrowseComp
AI Agent - 信息收集
84.30Thinking Level · High | Tools
90.10Deep Thinking Mode | Tools
85.90Thinking Level · High | Tools
83.20Thinking Enabled | Tools
84.40Thinking Level · High | Tools
GDPval-AA
生产力知识
1890.00Extended Thinking | Tools
82.30Thinking Level · Extra High
--
--
1769.00Thinking Level · High
5 additional benchmarks remain in the chart above.

Standard API Pricing: Claude Opus 4.8 vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

Gemini 3.1 Pro Preview: Base price applies to <= 200K
ModelSupplierStandard inputStandard outputBase price applies to
Claude Opus 4.8
Anthropic$5 / 1M tokens$25 / 1M tokens
GPT-5.5 Pro
OpenAI$30 / 1M tokens$180 / 1M tokens
Gemini 3.1 Pro Preview
Google Deep Mind$2 / 1M tokens$12 / 1M tokens<= 200K
Kimi K2.6
Facebook AI研究实验室$0.95 / 1M tokens$4 / 1M tokens
GPT-5.5
OpenAI$5 / 1M tokens$30 / 1M tokens

Version History

How each version of the Claude Opus 4.8 series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkClaude Opus 4.8CurrentOpus 4.7Claude Opus 4.6Opus 4.5
HLE
综合评估
57.90Extended Thinking | Tools
54.70Extended Thinking | Tools
53.00Extended Thinking | Tools
43.20Extended Thinking | Tools
LiveBench
综合评估
78.79Deep Thinking Mode
76.91Deep Thinking Mode
76.33Thinking Level · High
75.9664K
GPQA Diamond
科学与综合推理
93.60Thinking Level · High
94.20Extended Thinking
91.31Extended Thinking
87.00Extended Thinking
SWE-Bench Pro - Public
编程与软件工程
69.20Extended Thinking | Tools
64.30Extended Thinking | Tools
--
--
SWE-bench Verified
编程与软件工程
88.60Extended Thinking | Tools
87.60Extended Thinking | Tools
80.84Extended Thinking | Tools
80.90Extended Thinking | Tools
Text Arena (Coding)
编程与软件工程
1545.05Standard Mode
1562.39Standard Mode
1555.35Standard Mode
1512.0032K
WeirdML v2
编程与软件工程
82.89Thinking Level · Extra High | Tools
76.40Standard Mode | Tools
77.95Thinking Level · High | Tools
63.7016K | Tools
Creative Writing
写作和创作
1835.30Standard Mode
1906.40Standard Mode
1803.80Standard Mode
1683.20Standard Mode
SimpleBench
常识推理
64.80Standard Mode
62.90Standard Mode
67.60Standard Mode
62.00Extended Thinking
BrowseComp
AI Agent - 信息收集
84.30Thinking Level · High | Tools
79.30Extended Thinking | Tools
84.00Thinking Enabled | Tools
--
GDPval-AA
生产力知识
1890.00Extended Thinking | Tools
--
1606.00Extended Thinking | Tools
--
MCP-Atlas
AI Agent - 工具使用
82.20Thinking Level · High | Tools
79.10Thinking Level · High | Tools
76.80Thinking Level · High | Tools
69.80Thinking Level · High | Tools
4 additional benchmarks remain in the chart above.

Single-Benchmark Version Trend

Viewing: HLE · 综合评估

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the Claude Opus 4.8 Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

Claude Opus 4.6: Base price applies to <= 200K
ModelSupplierStandard inputStandard outputBase price applies to
Claude Opus 4.8
Anthropic$5 / 1M tokens$25 / 1M tokens
Opus 4.7
Anthropic$5 / 1M tokens$25 / 1M tokens
Claude Opus 4.6
Anthropic$5 / 1M tokens$25 / 1M tokens<= 200K
Opus 4.5
Facebook AI研究实验室$5 / 1M tokens$25 / 1M tokens

Sources