DataLearner logo

Claude Mythos Preview Benchmark Details

Claude Mythos Preview currently shows benchmark results led by GPQA Diamond (1 / 187, score 94.60), HLE (1 / 170, score 64.70), SWE-bench Verified (3 / 111, score 93.90). This page also compares it with 1 competitor models and 1 predecessor or same-series models, including performance and pricing views when available. 1 source link is attached for reference.

Benchmark Results

Claude Mythos Preview

Benchmark Results

Thinking
Tool usage

General Knowledge

3 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Extended Thinking
94.60
1 / 187
HLE
Extended Thinking
56.80
9 / 170
HLE
Extended ThinkingTools
64.70
1 / 170

Coding and Software Engineer

3 evaluations
Benchmark / mode
Score
Rank/total
SWE-bench Verified
Extended ThinkingTools
93.90
3 / 111
SWE-bench Multilingual
Extended ThinkingTools
87.30
1 / 22
SWE-Bench Pro - Public
Extended ThinkingTools
77.80
2 / 51

AI Agent - Information Search

1 evaluations
Benchmark / mode
Score
Rank/total
BrowseComp
Extended ThinkingTools
84.90
5 / 52

AI Agent - Tool Usage

2 evaluations
Benchmark / mode
Score
Rank/total
Terminal Bench 2.0
Extended ThinkingTools
82
2 / 47
OSWorld-Verified
Extended ThinkingTools
79.60
5 / 20

Competitor Comparison

Benchmark scores for Claude Mythos Preview compared against top models in its class

Claude Mythos PreviewGPT-5.4 Pro
Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

3 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkClaude Mythos PreviewCurrentGPT-5.4 Pro
GPQA Diamond
综合评估
94.60Extended Thinking
94.40Thinking Level · High
HLE
综合评估
64.70Extended Thinking | Tools
58.70Thinking Level · High | Tools
BrowseComp
AI Agent - 信息收集
84.90Extended Thinking | Tools
89.30Thinking Level · High | Tools

Standard API Pricing: Claude Mythos Preview vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

GPT-5.4 Pro: Base price applies to <= 272K
ModelSupplierStandard inputStandard outputBase price applies to
Claude Mythos Preview
Anthropic$25 / 1M tokens$125 / 1M tokens
GPT-5.4 Pro
OpenAI$30 / 1M tokens$180 / 1M tokens<= 272K

Version History

How each version of the Claude Mythos Preview series stacks up on benchmark tests

Claude Mythos PreviewClaude Opus 4.6
Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

7 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkClaude Mythos PreviewCurrentClaude Opus 4.6
GPQA Diamond
综合评估
94.60Extended Thinking
91.31Extended Thinking
HLE
综合评估
64.70Extended Thinking | Tools
53.00Extended Thinking | Tools
SWE-bench Multilingual
编程与软件工程
87.30Extended Thinking | Tools
72.00Extended Thinking | Tools
SWE-bench Verified
编程与软件工程
93.90Extended Thinking | Tools
80.84Extended Thinking | Tools
BrowseComp
AI Agent - 信息收集
84.90Extended Thinking | Tools
84.00Thinking Enabled | Tools
OSWorld-Verified
AI Agent - 工具使用
79.60Extended Thinking | Tools
72.70Extended Thinking | Tools
Terminal Bench 2.0
AI Agent - 工具使用
82.00Extended Thinking | Tools
65.40Extended Thinking | Tools

Single-Benchmark Version Trend

Viewing: GPQA Diamond · 综合评估

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the Claude Mythos Preview Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

Claude Opus 4.6: Base price applies to <= 200K
ModelSupplierStandard inputStandard outputBase price applies to
Claude Mythos Preview
Anthropic$25 / 1M tokens$125 / 1M tokens
Claude Opus 4.6
Anthropic$5 / 1M tokens$25 / 1M tokens<= 200K

Sources