DataLearner logo
OP

Opus 4.5

Reasoning modelOpusClaude 4.5

Claude Opus 4.5

Release date: 2025-11-25Updated: 2026-07-17 23:18:28.642Knowledge cutoff: 2025-032,682
Live demoGitHubHugging FaceCompare
Parameters
Not disclosed
Context length
200K
Chinese support
Supported
Reasoning ability

Opus 4.5 is a reasoning model from Anthropic, released on 2025-11-25. It accepts text and image input and returns text output. The recorded context window is 200K, and the recorded maximum output is 64K. Cataloged capabilities include Reasoning model and Multilingual. The model weights are proprietary and are not published for download. Use the linked references to confirm current access, licensing, and provider-specific limits.

Data sourced primarily from official releases (GitHub, Hugging Face, papers), then benchmark leaderboards, then third-party evaluators. Learn about our data methodology

Opus 4.5

Model basics

Reasoning traces
Supported
Thinking modes
Thinking Level · Extended (Default)Standard Mode
Context length
200K tokens
Max output length
64K tokens
Model type
Reasoning model
Modality (in / out)
Text, Image → Text
Release date
2025-11-25
Model file size
No data
MoE architecture
No
Total params / Active params
No data / N/A
Knowledge cutoff
2025-03
Opus 4.5

Open source & experience

Code license
Proprietary
Weights license
Proprietary
GitHub repo
GitHub link unavailable
Hugging Face
Hugging Face link unavailable
Opus 4.5

Official resources

Paper
DataLearnerAI blog
No blog post yet
Opus 4.5

API details

API speed
3/5
💡Default unit: $/1M tokens. If vendors use other units, follow their published pricing.
Standard
TypeConditionInputOutput
Text-$5.00/ 1M$25.00/ 1M
Batch
TypeConditionInputOutput
Text-$2.50/ 1M$12.50/ 1M
Cache PricingPrompt Cache
TypeTTLWriteRead
Text5m$6.25/ 1M$0.500/ 1M
Opus 4.5

Benchmark Results

Opus 4.5 currently shows benchmark results led by MMLU Pro (2 / 133, score 90), Terminal Bench Hard (1 / 13, score 44), SWE-bench Verified (9 / 114, score 80.90). This page also consolidates core specs, context limits, and API pricing so you can evaluate the model from benchmark results and deployment constraints together.

Thinking
Tool usage

General Knowledge

9 evaluations
Benchmark / mode
Score
Rank/total
MMLU Pro
Extended
90
2 / 133
ARC-AGI
Extended
80
24 / 68
75.96
11 / 115
55.77
80 / 115
LiveBench
Medium
59.10
72 / 115
58.59
74 / 115
HLE
Extended
30.80
88 / 181
HLE
ExtendedTools
43.20
52 / 181
ARC-AGI-2
Extended
37.60
29 / 62

General Evaluation

1 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Extended
87
70 / 224

Coding and Software Engineer

6 evaluations
Benchmark / mode
Score
Rank/total
Text Arena (Coding)
Standard Mode
1479
20 / 35
1512
17 / 35
LiveCodeBench
ExtendedTools
87
14 / 126
SWE-bench Verified
ExtendedTools
80.90
9 / 114
WeirdML v2
16KTools
63.70
24 / 52
GSO
Standard ModeTools
26.50
8 / 21

Writing and Creative Capabilities

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
1683.20
25 / 99

Multimodal Understanding

3 evaluations
Benchmark / mode
Score
Rank/total
MMMU
Extended
80.70
10 / 28
GeoBench ACW
Standard Mode
75
10 / 20
VPCT
32K
40
15 / 24

Common Sense Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleBench
Extended
62
14 / 67

Agent Level Benchmark

7 evaluations
Benchmark / mode
Score
Rank/total
METR Time Horizons v1.1
Standard ModeTools
293
6 / 22
288.90
7 / 22
90.70
21 / 35
τ²-Bench
ExtendedTools
81.99
13 / 43
Terminal Bench Hard
ExtendedTools
44
1 / 13
BALROG
Standard ModeTools
43.50
5 / 12
BALROG
64KTools
43
7 / 12

Math and Reasoning

8 evaluations
Benchmark / mode
Score
Rank/total
AIME 2026
Extended
93.30
8 / 19
34.39
30 / 34
FrontierMath
Extended
20.70
17 / 60
4.20
40 / 80
2.10
56 / 80
4.20
40 / 80
4.20
40 / 80

Instruction Following

1 evaluations
Benchmark / mode
Score
Rank/total
IF Bench
ExtendedTools
58
24 / 33

AI Agent - Tool Usage

2 evaluations
Benchmark / mode
Score
Rank/total
MCP-Atlas
HighTools
69.80
27 / 39
Terminal Bench 2.0
ExtendedTools
59.30
20 / 48

Claw-style Agent Evaluation

2 evaluations
Benchmark / mode
Score
Rank/total
Claw Bench
ExtendedTools
91.50
7 / 29
Pinch Bench
ExtendedTools
87.20
9 / 38

Long Context

1 evaluations
Benchmark / mode
Score
Rank/total
LongBench v2
Standard Mode
64.40
2 / 13

Compare with other models

No curated comparisons for this model yet.

Want a custom combination? Open the compare tool

Opus 4.5

Publisher

Claude Opus 4.5

Model Overview

Opus 4.5 is a reasoning model from Anthropic, released on 2025-11-25.

It accepts text and image input and produces text output. Its cataloged capabilities include Reasoning model and Multilingual. The recorded context window is 200K, and the recorded maximum output is 64K.

The model weights are proprietary and are not published for download. The page records 6 API pricing rules from the listed provider; current provider pricing and conditions should be checked before deployment. The evaluation section contains 29 cataloged benchmark results with their recorded modes and scores. The page links 2 release, model-card, repository, or provider references for checking the underlying claims. Specifications, availability, and prices can change; undisclosed values are intentionally left unstated.

Opus 4.5

FAQ

What is Opus 4.5?

Opus 4.5 is a reasoning model from Anthropic, released on 2025-11-25. It accepts text and image input and returns text output. The recorded context window is 200K, and the recorded maximum output is 64K. Cataloged capabilities include Reasoning model and Multilingual. The model weights are proprietary and are not published for download. Use the linked references to confirm current access, licensing, and provider-specific limits.

What input and output modalities does Opus 4.5 support?

The current model record lists text and image as input and text as output.

What are the main recorded specifications for Opus 4.5?

The recorded context window is 200K, and the recorded maximum output is 64K. Fields without a source-backed value remain undisclosed.

Does Opus 4.5 have API pricing?

The page records 6 API pricing rules from the listed provider; current provider pricing and conditions should be checked before deployment.

Are benchmark results available for Opus 4.5?

The evaluation section contains 29 cataloged benchmark results with their recorded modes and scores. Compare only results that use the same benchmark version and evaluation mode.

Is Opus 4.5 open source?

The model weights are proprietary and are not published for download. Review the linked license text before commercial or derivative use.

DataLearner on WeChat

Follow DataLearner on WeChat for AI model updates and research notes.

DataLearner WeChat QR code