GPT-5.5 ProvsClaude Mythos Preview
Across 3 shared benchmarks, Claude Mythos Preview leads overall: GPT-5.5 Pro wins 1, Claude Mythos Preview wins 2, with 0 ties and an average score difference of -0.99.
GPT-5.5 Pro
OpenAI · 2026-04-23 · Reasoning model
Claude Mythos Preview
Anthropic · 2026-04-07 · Chat model
GPT-5.5 Pro1 win(33%)(67%)2 winsClaude Mythos Preview
Benchmark scores
Grouped by capability, sorted by largest gap within each. 3 shared benchmarks.
AI Agent - Information Search
GPT-5.5 Pro 1/1| Benchmark | GPT-5.5 Pro | Claude Mythos Preview | Diff |
|---|---|---|---|
| BrowseComp | 90.103 / 53Deep Thinking (With Tools + Internet) | 84.906 / 53Extended (with tools) | +5.20 |
General Evaluation
Claude Mythos Preview 1/1| Benchmark | GPT-5.5 Pro | Claude Mythos Preview | Diff |
|---|---|---|---|
| GPQA Diamond | 93.928 / 225极高强度思考(无工具) | 94.601 / 225Extended (no tools) | -0.68 |
General Knowledge
Claude Mythos Preview 1/1| Benchmark | GPT-5.5 Pro | Claude Mythos Preview | Diff |
|---|---|---|---|
| HLE | 57.209 / 175极高强度思考(工具) | 64.701 / 175Extended (with tools) | -7.50 |
Specs
| Field | GPT-5.5 Pro | Claude Mythos Preview |
|---|---|---|
| Publisher | OpenAI | Anthropic |
| Release date | 2026-04-23 | 2026-04-07 |
| Model type | Reasoning model | Chat model |
| Architecture | Dense | Dense |
| Parameters | Not available | Not available |
| Context length | 1000K | Not available |
| Max output | 128K | 8K |
API pricing
Prices use DataLearner records when available; missing fields are not inferred.
| Item | GPT-5.5 Pro | Claude Mythos Preview |
|---|---|---|
| Text input | $30 / 1M tokens | $25 / 1M tokens |
| Text output | $180 / 1M tokens | $125 / 1M tokens |
Summary
- GPT-5.5 Proleads in:AI Agent - Information Search (1/1)
- Claude Mythos Previewleads in:General Evaluation (1/1), General Knowledge (1/1)
On average across the 3 shared benchmarks, Claude Mythos Preview scores 0.99 higher.
Largest single-benchmark gap: HLE — GPT-5.5 Pro 57.20 vs Claude Mythos Preview 64.70 (-7.50).
Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.