Claude Opus 4.8vsGPT-5.5 Pro
Across 7 shared benchmarks, GPT-5.5 Pro leads overall: Claude Opus 4.8 wins 2, GPT-5.5 Pro wins 5, with 0 ties and an average score difference of +251.50.
Claude Opus 4.8
Anthropic · 2026-05-28 · Reasoning model
GPT-5.5 Pro
OpenAI · 2026-04-23 · Reasoning model
Claude Opus 4.82 wins(29%)(71%)5 winsGPT-5.5 Pro
Benchmark scores
Grouped by capability, sorted by largest gap within each. 7 shared benchmarks.
Math and Reasoning
GPT-5.5 Pro 2/2| Benchmark | Claude Opus 4.8 | GPT-5.5 Pro | Diff |
|---|---|---|---|
| FrontierMath Tier 4 v2 | 56.109 / 34最高(无工具) | 78.053 / 34极高强度思考(无工具) | -21.95 |
| FrontierMath v2 | 809 / 34最高(无工具) | 87.722 / 34极高强度思考(无工具) | -7.72 |
AI Agent - Information Search
GPT-5.5 Pro 1/1| Benchmark | Claude Opus 4.8 | GPT-5.5 Pro | Diff |
|---|---|---|---|
| BrowseComp | 84.309 / 54Thinking High (With Tools + Internet) | 90.103 / 54Deep Thinking (With Tools + Internet) | -5.80 |
Commonsense Reasoning
GPT-5.5 Pro 1/1| Benchmark | Claude Opus 4.8 | GPT-5.5 Pro | Diff |
|---|---|---|---|
| SimpleBench | 64.8010 / 67Normal (No Tools) | 76.903 / 67Normal (No Tools) | -12.10 |
General Evaluation
GPT-5.5 Pro 1/1| Benchmark | Claude Opus 4.8 | GPT-5.5 Pro | Diff |
|---|---|---|---|
| GPQA Diamond | 93.6011 / 226Thinking High (No Tools) | 93.928 / 226极高强度思考(无工具) | -0.32 |
General Knowledge
Claude Opus 4.8 1/1| Benchmark | Claude Opus 4.8 | GPT-5.5 Pro | Diff |
|---|---|---|---|
| HLE | 57.908 / 181Extended (with tools) | 57.2010 / 181极高强度思考(工具) | +0.70 |
Productivity Knowledge
Claude Opus 4.8 1/1| Benchmark | Claude Opus 4.8 | GPT-5.5 Pro | Diff |
|---|---|---|---|
| GDPval-AA | 1,8901 / 21Extended (with tools) | 82.307 / 21极高强度思考(无工具) | +1,808 |
Specs
| Field | Claude Opus 4.8 | GPT-5.5 Pro |
|---|---|---|
| Publisher | Anthropic | OpenAI |
| Release date | 2026-05-28 | 2026-04-23 |
| Model type | Reasoning model | Reasoning model |
| Architecture | Dense | Dense |
| Parameters | Not available | Not available |
| Context length | 1M | 1000K |
| Max output | 125K | 128K |
API pricing
Prices use DataLearner records when available; missing fields are not inferred.
| Item | Claude Opus 4.8 | GPT-5.5 Pro |
|---|---|---|
| Text input | $5 / 1M tokens | $30 / 1M tokens |
| Text output | $25 / 1M tokens | $180 / 1M tokens |
| Cache read | $0.5 / 1M tokens | Not public |
| Cache write | $6.25 / 1M tokens | Not public |
Summary
- Claude Opus 4.8leads in:General Knowledge (1/1), Productivity Knowledge (1/1)
- GPT-5.5 Proleads in:Math and Reasoning (2/2), AI Agent - Information Search (1/1), Commonsense Reasoning (1/1), General Evaluation (1/1)
On average across the 7 shared benchmarks, Claude Opus 4.8 scores 251.50 higher.
Largest single-benchmark gap: GDPval-AA — Claude Opus 4.8 1,890 vs GPT-5.5 Pro 82.30 (+1,808).
Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.