Claude Opus 5 Benchmark Details
Claude Opus 5 currently shows benchmark results led by Context Arena (1 / 126, score 97.72), SWE-bench Verified (1 / 116, score 96), HLE (2 / 191, score 64.70). This page also compares it with 3 competitor models and 4 predecessor or same-series models, including performance and pricing views when available. 1 source link is attached for reference.
Benchmark Results
Benchmark Results
General Knowledge
10 evaluationsCoding and Software Engineer
8 evaluationsWriting and Creative Capabilities
1 evaluationsAI Agent - Tool Usage
4 evaluationsProductivity Knowledge
6 evaluationsMath and Reasoning
2 evaluationsCompetitor Comparison
Benchmark scores for Claude Opus 5 compared against top models in its class
12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.
| Benchmark | Claude Opus 5Current | GPT-5.6 Sol | Kimi K3 | GLM-5.2 |
|---|---|---|---|---|
63.00Thinking Level · Extra High | Tools | 61.00Thinking Level · High | Tools | -- | -- | |
97.50Thinking Level · Extra High | 97.50Thinking Level · Extra High | -- | -- | |
90.40Thinking Level · High | 92.50Thinking Level · High | -- | -- | |
30.20Thinking Level · High | 7.80Thinking Level · High | -- | -- | |
64.70Thinking Level · High | Tools | 49.50Thinking Level · High | 56.00Thinking Level · High | Tools | 54.70Thinking Enabled | Tools | |
93.88Thinking Level · High | 94.60Thinking Level · High | 93.50Thinking Level · High | 91.86Thinking Level · High | |
68.80Thinking Level · High | Tools | 72.70Thinking Level · Extra High | Tools | 67.50Thinking Level · High | Tools | 44.00Deep Thinking Mode | Tools | |
79.20Thinking Level · High | Tools | 64.60Thinking Level · Extra High | Tools | -- | 62.10Thinking Enabled | Tools | |
1711.88Thinking Level · High | 1620.27Thinking Level · Extra High | 1681.75Thinking Level · High | 1593.25Thinking Level · High | |
91.59Thinking Level · High | Tools | 88.76Thinking Level · High | Tools | -- | 67.31Thinking Level · High | Tools | |
2120.60Standard Mode | 1963.40Standard Mode | 2070.60Standard Mode | 1752.80Standard Mode | |
80.60Thinking Level · High | 64.80Thinking Level · Extra High | 60.70Thinking Level · High | 58.80Standard Mode |
Standard API Pricing: Claude Opus 5 vs. Peer Models
Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.
Source: DataLearnerAI. Standard text prices shown here use the default supplier.
These models use different currencies or billing units, so the page falls back to raw price values instead of a shared bar chart.
| Model | Supplier | Standard input | Standard output | Base price applies to |
|---|---|---|---|---|
Claude Opus 5 | Anthropic | $5 / 1M tokens | $25 / 1M tokens | — |
GPT-5.6 Sol | OpenAI | $4 / 1M tokens | $20 / 1M tokens | — |
Kimi K3 | Moonshot AI | ¥20 / 1M tokens | ¥100 / 1M tokens | — |
GLM-5.2 | 智谱AI | $1.4 / 1M tokens | $4.4 / 1M tokens | — |
Version History
How each version of the Claude Opus 5 series stacks up on benchmark tests
12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.
| Benchmark | Claude Opus 5Current | Claude Opus 4.8 | Opus 4.7 | Claude Opus 4.6 | Opus 4.5 |
|---|---|---|---|---|---|
97.50Thinking Level · Extra High | -- | 93.50Thinking Level · High | 92.00Extended Thinking | 80.00Extended Thinking | |
90.40Thinking Level · High | -- | 75.80Thinking Level · High | 66.30Extended Thinking | 37.60Extended Thinking | |
64.70Thinking Level · High | Tools | 57.90Extended Thinking | Tools | 54.70Extended Thinking | Tools | 53.00Extended Thinking | Tools | 43.20Extended Thinking | Tools | |
93.88Thinking Level · High | 93.60Thinking Level · High | 94.20Extended Thinking | 91.31Extended Thinking | 87.00Extended Thinking | |
68.80Thinking Level · High | Tools | 59.00Deep Thinking Mode | Tools | -- | -- | -- | |
89.50Thinking Level · High | Tools | -- | -- | 72.00Extended Thinking | Tools | -- | |
79.20Thinking Level · High | Tools | 69.20Extended Thinking | Tools | 64.30Extended Thinking | Tools | -- | -- | |
96.00Thinking Level · High | Tools | 88.60Extended Thinking | Tools | 87.60Extended Thinking | Tools | 80.84Extended Thinking | Tools | 80.90Extended Thinking | Tools | |
1711.88Thinking Level · High | 1545.05Standard Mode | 1562.39Standard Mode | 1555.35Standard Mode | 1512.0032K | |
91.59Thinking Level · High | Tools | 82.89Thinking Level · Extra High | Tools | 76.40Standard Mode | Tools | 77.95Thinking Level · High | Tools | 63.7016K | Tools | |
2120.60Standard Mode | 1835.20Standard Mode | 1907.10Standard Mode | 1804.10Standard Mode | 1683.20Standard Mode | |
80.60Thinking Level · High | 64.80Standard Mode | 61.70Standard Mode | 67.60Standard Mode | 62.00Extended Thinking |
Single-Benchmark Version Trend
Viewing: ARC-AGI-1 · 综合评估
Standard API Pricing Across the Claude Opus 5 Series
Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.
Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens
| Model | Supplier | Standard input | Standard output | Base price applies to |
|---|---|---|---|---|
Claude Opus 5 | Anthropic | $5 / 1M tokens | $25 / 1M tokens | — |
Claude Opus 4.8 | Anthropic | $5 / 1M tokens | $25 / 1M tokens | — |
Opus 4.7 | Anthropic | $5 / 1M tokens | $25 / 1M tokens | — |
Claude Opus 4.6 | Anthropic | $5 / 1M tokens | $25 / 1M tokens | — |
Opus 4.5 | Facebook AI研究实验室 | $5 / 1M tokens | $25 / 1M tokens | — |