GPT-5.5vsOpus 4.7
Across 17 shared benchmarks, Opus 4.7 leads overall: GPT-5.5 wins 7, Opus 4.7 wins 10, with 0 ties and an average score difference of -3.14.
GPT-5.57 wins(41%)(59%)10 winsOpus 4.7
Benchmark scores
Grouped by capability, sorted by largest gap within each. 17 shared benchmarks.
General Knowledge
Opus 4.7 2/3| Benchmark | GPT-5.5 | Opus 4.7 | Diff |
|---|---|---|---|
| HLE | 13.70351 / 563Normal (No Tools) · Text only | 33.30184 / 563Normal (No Tools) · Text only | -19.60 |
| CritPt | 1.40135 / 200Normal (No Tools) | 5.1094 / 200Normal (No Tools) | -3.70 |
| LiveBench | 79.911 / 117Deep Thinking (No Tools) | 76.538 / 117Deep Thinking (No Tools) | +3.38 |
Math and Reasoning
GPT-5.5 3/3| Benchmark | GPT-5.5 | Opus 4.7 | Diff |
|---|---|---|---|
| MathArena Apex | 80.212 / 17Extra-High (With Tools) | 40.629 / 17Extra-High (With Tools) | +39.59 |
| HMMT Feb 2026 | 98.481 / 10Extra-High (With Tools) | 93.946 / 10Extra-High (With Tools) | +4.54 |
| AIME 2026 | 1001 / 29Extra-High (With Tools) | 95.8310 / 29Extra-High (With Tools) | +4.17 |
Agent Level Benchmark
Opus 4.7 2/2| Benchmark | GPT-5.5 | Opus 4.7 | Diff |
|---|---|---|---|
| Terminal Bench Hard | 49.2024 / 244Normal (With Tools) | 54.5015 / 244Normal (With Tools) | -5.30 |
| τ²-Bench - Telecom | 69.30139 / 264Normal (With Tools) | 74128 / 264Normal (With Tools) | -4.70 |
Claw-style Agent Evaluation
Opus 4.7 1/1| Benchmark | GPT-5.5 | Opus 4.7 | Diff |
|---|---|---|---|
| PinchBench v2 | 75.5018 / 45Reported best (effort unspecified) | 76.0115 / 45Reported best (effort unspecified) | -0.51 |
Coding and Software Engineer
Opus 4.7 1/1| Benchmark | GPT-5.5 | Opus 4.7 | Diff |
|---|---|---|---|
| WeirdML v2 | 67.1522 / 52Normal (With Tools) | 76.4013 / 52Normal (With Tools) | -9.25 |
Commonsense Reasoning
GPT-5.5 1/1| Benchmark | GPT-5.5 | Opus 4.7 | Diff |
|---|---|---|---|
| SimpleBench | 6918 / 93Normal (No Tools) | 61.7027 / 93Normal (No Tools) | +7.30 |
General Evaluation
Opus 4.7 1/1| Benchmark | GPT-5.5 | Opus 4.7 | Diff |
|---|---|---|---|
| GPQA Diamond | 77.27264 / 462Normal (No Tools) | 88.50110 / 462Normal (No Tools) | -11.23 |
Instruction Following
GPT-5.5 1/1| Benchmark | GPT-5.5 | Opus 4.7 | Diff |
|---|---|---|---|
| IF Bench | 46.10174 / 282Normal (No Tools) | 43.60189 / 282Normal (No Tools) | +2.50 |
Long Context
Opus 4.7 1/1| Benchmark | GPT-5.5 | Opus 4.7 | Diff |
|---|---|---|---|
| AA-LCR | 64122 / 170Normal (No Tools) | 75.7082 / 170Normal (No Tools) | -11.70 |
Multimodal Understanding
Opus 4.7 1/1| Benchmark | GPT-5.5 | Opus 4.7 | Diff |
|---|---|---|---|
| MMMU-Pro | 71.40119 / 227Normal (No Tools) | 76.4073 / 227Normal (No Tools) | -5 |
Text Embedding
GPT-5.5 1/1| Benchmark | GPT-5.5 | Opus 4.7 | Diff |
|---|---|---|---|
| Context Arena | 41.9398 / 126Normal (No Tools) | 22.12122 / 126Normal (No Tools) | +19.81 |
Writing and Creative Capabilities
Opus 4.7 1/1| Benchmark | GPT-5.5 | Opus 4.7 | Diff |
|---|---|---|---|
| Creative Writing | 1,84412 / 106Normal (No Tools) | 1,9079 / 106Normal (No Tools) | -63.60 |
Specs
| Field | GPT-5.5 | Opus 4.7 |
|---|---|---|
| Publisher | OpenAI | Anthropic |
| Release date | 2026-04-23 | 2026-04-16 |
| Model type | Reasoning model | Reasoning model |
| Architecture | Dense | Dense |
| Parameters | Not available | Not available |
| Context length | 1000K | 1000K |
| Max output | 128K | 128K |
API pricing
Prices use DataLearner records when available; missing fields are not inferred.
| Item | GPT-5.5 | Opus 4.7 |
|---|---|---|
| Text input | $5 / 1M tokens | $5 / 1M tokens |
| Text output | $30 / 1M tokens | $25 / 1M tokens |
| Cache read | $0.5 / 1M tokens | $0.5 / 1M tokens |
| Cache write | $6.25 / 1M tokens | $6.25 / 1M tokens |
Summary
- GPT-5.5leads in:Math and Reasoning (3/3), Commonsense Reasoning (1/1), Instruction Following (1/1), Text Embedding (1/1)
- Opus 4.7leads in:General Knowledge (2/3), Agent Level Benchmark (2/2), Claw-style Agent Evaluation (1/1), Coding and Software Engineer (1/1), General Evaluation (1/1), Long Context (1/1), Multimodal Understanding (1/1), Writing and Creative Capabilities (1/1)
On average across the 17 shared benchmarks, Opus 4.7 scores 3.14 higher.
Largest single-benchmark gap: Creative Writing — GPT-5.5 1,844 vs Opus 4.7 1,907 (-63.60).
Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.