DataLearner logo

Claude Fable 5 Benchmark Details

Claude Fable 5 currently shows benchmark results led by ARC-AGI-1 (1 / 92, score 98.50), LiveBench (2 / 117, score 79.53), SWE-bench Verified (2 / 116, score 95). This page also compares it with 4 competitor models and 2 predecessor or same-series models, including performance and pricing views when available.

Benchmark Results

Claude Fable 5

Benchmark Results

Thinking
Tool usage
Internet

General Knowledge

14 evaluations
Benchmark / mode
Score
Rank/total
ARC-AGI-1
Thinking Level · Low
90.50
33 / 92
ARC-AGI-1
Thinking Level · Medium
92.50
25 / 92
ARC-AGI-1
Thinking Level · High
95.50
15 / 92
ARC-AGI-1
Thinking Level · Max
98.50
1 / 92
ARC-AGI-1
Thinking Level · Extra High
98.50
1 / 92
ARC-AGI-2
Thinking Level · Low
76.80
25 / 85
ARC-AGI-2
Thinking Level · Medium
82.50
21 / 85
ARC-AGI-2
Thinking Level · High
87.50
10 / 85
ARC-AGI-2
Thinking Level · Max
89.20
7 / 85
ARC-AGI-2
Thinking Level · Extra High
88.30
9 / 85
LiveBench
Thinking Level · High
75.47
9 / 117
LiveBench
Deep Thinking Mode
79.53
2 / 117
62
5 / 28
HLE
Deep Thinking Mode
59
9 / 197

Other

3 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Thinking Level · Low
78.79
155 / 274
GPQA Diamond
Thinking Level · High
83.33
117 / 274
GPQA Diamond
Thinking Level · Max
85.86
95 / 274

Coding and Software Engineer

10 evaluations
Benchmark / mode
Score
Rank/total
SWE-bench Verified
Thinking Level · HighTools
95
2 / 116
SWE-bench Verified
Deep Thinking ModeTools
95
2 / 116
WeirdML v2
Thinking Level · HighTools
87.85
3 / 52
SWE-Bench Pro - Public
Deep Thinking ModeTools
80.30
2 / 62
CursorBench 3.2
Thinking Level · MaxTools
70.50
2 / 5
DeepSWE
Deep Thinking ModeTools
70
8 / 38
FrontierCode 1.1 Extended
Thinking Level · Extra HighTools
64.90
1 / 7
SciCode
Thinking Level · Max
60.19
1 / 16
APEX-SWE
Thinking Level · MaxTools
58.80
1 / 3
FrontierCode 1.1 Main
Thinking Level · Extra HighTools
53.50
1 / 4

Writing and Creative Capabilities

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
1934.60
7 / 106

Common Sense Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleBench
Standard Mode
81.90
5 / 94

Long Context

1 evaluations
Benchmark / mode
Score
Rank/total
AA-LCR
Thinking Level · Max
76.67
5 / 29

AI Agent - Tool Usage

8 evaluations
Benchmark / mode
Score
Rank/total
Terminal-Bench 2.1
Thinking Level · HighTools
88
9 / 53
Terminal-Bench 2.1
Thinking Level · Extra HighTools
83.80
19 / 53
Terminal-Bench 2.1
Deep Thinking ModeTools
88
9 / 53
OSWorld-Verified
Thinking Level · HighTools
85
1 / 26
MCP-Atlas
Standard ModeTools
83.30
7 / 41
Terminal-Bench 4.0
Thinking Level · MaxTools
44.55
8 / 20
Terminal-Bench 3.0
Thinking Level · MaxTools
34.10
3 / 11
Terminal-Bench-Science 0.1
Thinking Level · MaxTools
21.40
5 / 11

Productivity Knowledge

4 evaluations
Benchmark / mode
Score
Rank/total
GDPval-AA v2
Thinking Level · MaxTools
1741
8 / 27
AA-Briefcase
Thinking Level · MaxTools
1571.79
6 / 20
AA-AnalystAgent
Thinking Level · MaxToolsInternet
48.75
4 / 12
Harvey Lab-AA
Thinking Level · MaxTools
11.30
5 / 6

Agent Level Benchmark

2 evaluations
Benchmark / mode
Score
Rank/total
APEX-Agents
Thinking Level · MaxTools
59.20
1 / 6
τ³-Banking
Thinking Level · MaxTools
38.14
3 / 13

Math and Reasoning

2 evaluations
Benchmark / mode
Score
Rank/total
FrontierMath Tier 4 v2
Thinking Level · Max
87.80
2 / 42
FrontierMath v2
Thinking Level · Max
87.02
3 / 58

Multimodal Understanding

1 evaluations
Benchmark / mode
Score
Rank/total
GeoBench ACW
Standard Mode
89
1 / 20

Competitor Comparison

Benchmark scores for Claude Fable 5 compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkClaude Fable 5CurrentGPT-5.5Gemini 3.1 Pro PreviewGLM-5.2DeepSeek-V4-Pro
ARC-AGI-1
Score (%)
综合评估
98.50Thinking Level · Extra High
95.00Thinking Level · Extra High
--
--
--
ARC-AGI-2
Accuracy and cost per task
综合评估
89.20Thinking Level · High
85.00Thinking Level · Extra High
77.10Thinking Level · High
--
--
HLE
Accuracy
综合评估
59.00Deep Thinking Mode
52.20Thinking Level · High | Tools
51.40Thinking Level · High | Tools
54.70Thinking Enabled | Tools
48.20Thinking Level · Extra High | Tools
LiveBench
Accuracy
综合评估
79.53Deep Thinking Mode
79.91Deep Thinking Mode
77.13Thinking Level · High
73.18Standard Mode
71.57Standard Mode
GPQA Diamond
Accuracy
科学与综合推理
85.86Thinking Level · High
94.00Thinking Level · Extra High
94.30Thinking Level · High
91.86Thinking Level · High
90.10Thinking Level · High
DeepSWE
Pass@1 (DeepSWE v1.1)
编程与软件工程
70.00Deep Thinking Mode | Tools
67.00Thinking Level · Extra High | Tools
12.00Thinking Level · High | Tools
44.00Deep Thinking Mode | Tools
62.70Thinking Level · Extra High | Tools
SWE-Bench Pro - Public
Accuracy
编程与软件工程
80.30Deep Thinking Mode | Tools
58.60Thinking Level · High | Tools
54.20Thinking Level · High | Tools
62.10Thinking Enabled | Tools
55.40Thinking Level · Extra High | Tools
SWE-bench Verified
Accuracy
编程与软件工程
95.00Thinking Level · High | Tools
--
80.60Thinking Level · High | Tools
--
80.60Thinking Level · Extra High | Tools
WeirdML v2
Average accuracy across 17 tasks (%)
编程与软件工程
87.85Thinking Level · High | Tools
84.91Thinking Level · Extra High | Tools
72.10Standard Mode | Tools
67.31Thinking Level · High | Tools
--
Creative Writing
Elo、大模型评判两两对战
写作和创作
1934.60Standard Mode
1843.50Standard Mode
1488.90Standard Mode
1752.80Standard Mode
1552.10Standard Mode
SimpleBench
Score (AVG@5)
常识推理
81.90Standard Mode
69.00Standard Mode
79.60Standard Mode
58.80Standard Mode
50.90Standard Mode
MCP-Atlas
Pass rate / claim coverage
AI Agent - 工具使用
83.30Standard Mode | Tools
75.30Thinking Level · Extra High | Tools
78.20Thinking Level · High | Tools
76.80Thinking Enabled | Tools
--
7 additional benchmarks remain in the chart above.

Standard API Pricing: Claude Fable 5 vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

Gemini 3.1 Pro Preview: Base price applies to <= 200K
ModelSupplierStandard inputStandard outputBase price applies to
Claude Fable 5
Anthropic$10 / 1M tokens$50 / 1M tokens
GPT-5.5
OpenAI$5 / 1M tokens$30 / 1M tokens
Gemini 3.1 Pro Preview
Google Deep Mind$2 / 1M tokens$12 / 1M tokens<= 200K
GLM-5.2
智谱AI$1.4 / 1M tokens$4.4 / 1M tokens
DeepSeek-V4-Pro
DeepSeek-AI$0.435 / 1M tokens$0.87 / 1M tokens

Version History

How each version of the Claude Fable 5 series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkClaude Fable 5CurrentClaude Mythos PreviewClaude Opus 4.8
HLE
Accuracy
综合评估
59.00Deep Thinking Mode
64.70Extended Thinking | Tools
57.90Extended Thinking | Tools
LiveBench
Accuracy
综合评估
79.53Deep Thinking Mode
--
78.93Deep Thinking Mode
GPQA Diamond
Accuracy
科学与综合推理
85.86Thinking Level · High
94.60Extended Thinking
93.60Thinking Level · High
DeepSWE
Pass@1 (DeepSWE v1.1)
编程与软件工程
70.00Deep Thinking Mode | Tools
--
59.00Deep Thinking Mode | Tools
SWE-Bench Pro - Public
Accuracy
编程与软件工程
80.30Deep Thinking Mode | Tools
77.80Extended Thinking | Tools
69.20Extended Thinking | Tools
SWE-bench Verified
Accuracy
编程与软件工程
95.00Thinking Level · High | Tools
93.90Extended Thinking | Tools
88.60Extended Thinking | Tools
WeirdML v2
Average accuracy across 17 tasks (%)
编程与软件工程
87.85Thinking Level · High | Tools
--
82.89Thinking Level · Extra High | Tools
Creative Writing
Elo、大模型评判两两对战
写作和创作
1934.60Standard Mode
--
1835.20Standard Mode
SimpleBench
Score (AVG@5)
常识推理
81.90Standard Mode
--
64.80Standard Mode
MCP-Atlas
Pass rate / claim coverage
AI Agent - 工具使用
83.30Standard Mode | Tools
--
82.20Thinking Level · High | Tools
OSWorld-Verified
Accuracy
AI Agent - 工具使用
85.00Thinking Level · High | Tools
79.60Extended Thinking | Tools
83.40Extended Thinking | Tools
Terminal-Bench 2.1
Accuracy
AI Agent - 工具使用
88.00Thinking Level · High | Tools
--
78.90Thinking Level · High | Tools
4 additional benchmarks remain in the chart above.

Single-Benchmark Version Trend

Viewing: HLE · 综合评估

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the Claude Fable 5 Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
Claude Fable 5
Anthropic$10 / 1M tokens$50 / 1M tokens
Claude Mythos Preview
Anthropic$25 / 1M tokens$125 / 1M tokens
Claude Opus 4.8
Anthropic$5 / 1M tokens$25 / 1M tokens