GPT-5.4 Benchmark Details
GPT-5.4 currently shows benchmark results led by Pinch Bench (1 / 38, score 90.50), LiveBench (5 / 117, score 77.97), Terminal Bench Hard (11 / 244, score 57.60). This page also compares it with 2 competitor models and 2 predecessor or same-series models, including performance and pricing views when available. 2 source links are attached for reference.
Benchmark Results
Benchmark Results
Abstract Generalization
11 evaluationsKnowledge Exams
5 evaluationsScientific Reasoning
8 evaluationsMathematics
7 evaluationsRepository Engineering
3 evaluationsService Workflows
5 evaluationsInstruction Following
3 evaluationsCross-capability Suites
2 evaluationsAgentic Development
5 evaluationsMemory & Persistence
6 evaluationsLong Reasoning
3 evaluationsTool Orchestration
4 evaluationsVisual Understanding
5 evaluationsMaintenance & Optimization
3 evaluationsML Engineering
2 evaluationsPreference Arenas
2 evaluationsCapability Frontier Metrics
1 evaluationsCode Generation & Editing
1 evaluationsClinical Workflows
2 evaluationsCompetitor Comparison
Benchmark scores for GPT-5.4 compared against top models in its class
12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.
| Benchmark | GPT-5.4Current | Gemini 3.1 Pro Preview | Claude Opus 4.6 |
|---|---|---|---|
93.67Thinking Level · Extra High | -- | 94.00Thinking Level · High | |
77.10Standard Mode | 77.10Thinking Level · High | 69.17Thinking Level · High | |
ARC-AGI-3 (Standard harness) Action efficiency score(以 ARC Prize Standard harness 口径为准) Abstract Generalization | 0.21Thinking Level · High | 0.42Thinking Level · High | 0.51Thinking Level · High |
52.10Thinking Level · Extra High | Tools | 51.40Thinking Level · High | Tools | 53.00Extended Thinking | Tools | |
23.40Thinking Level · Extra High | 17.70Thinking Enabled | 12.60Thinking Level · High | |
92.00Thinking Level · Extra High | 94.30Thinking Level · High | 91.31Extended Thinking | |
1835.60Standard Mode | 1488.90Standard Mode | 1804.10Standard Mode | |
99.17Thinking Level · Extra High | Tools | -- | 96.67Thinking Level · High | Tools | |
47.60Thinking Level · Extra High | 36.90Thinking Level · High | 40.70Thinking Level · High | |
27.10Thinking Level · Extra High | 16.70Standard Mode | 22.90Thinking Level · High | |
49.00Thinking Level · Extra High | -- | 26.83Thinking Level · High | |
78.60Thinking Level · Extra High | -- | 65.96Thinking Level · High |
Standard API Pricing: GPT-5.4 vs. Peer Models
Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.
Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens
When a context threshold exists, the charted base price only applies within these limits:
| Model | Supplier | Standard input | Standard output | Base price applies to |
|---|---|---|---|---|
GPT-5.4 | OpenAI | $2.5 / 1M tokens | $15 / 1M tokens | — |
Gemini 3.1 Pro Preview | Google DeepMind | $2 / 1M tokens | $12 / 1M tokens | <= 200K |
Claude Opus 4.6 | Anthropic | $5 / 1M tokens | $25 / 1M tokens | — |
Version History
How each version of the GPT-5.4 series stacks up on benchmark tests
12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.
| Benchmark | GPT-5.4Current | GPT-5.2 | GPT-5.1 |
|---|---|---|---|
93.67Thinking Level · Extra High | 90.50Deep Thinking Mode | 72.83Thinking Level · High | |
77.10Standard Mode | 54.20Deep Thinking Mode | 17.64Thinking Level · High | |
52.10Thinking Level · Extra High | Tools | 45.50Deep Thinking Mode | Tools | 42.70Thinking Level · High | Tools | |
23.40Thinking Level · Extra High | 11.60Thinking Level · Extra High | 4.90Thinking Level · High | |
92.00Thinking Level · Extra High | 93.20Deep Thinking Mode | 88.10Thinking Enabled | |
1835.60Standard Mode | 1699.80Standard Mode | -- | |
99.17Thinking Level · Extra High | Tools | 98.33Thinking Level · High | Tools | -- | |
47.60Thinking Level · Extra High | 40.30Thinking Level · Extra High | Tools | 26.70Thinking Level · High | Tools | |
27.10Thinking Level · Extra High | 18.80Thinking Level · Extra High | 12.50Thinking Level · High | Tools | |
49.00Thinking Level · Extra High | 31.70Thinking Level · Extra High | -- | |
78.60Thinking Level · Extra High | 67.40Thinking Level · Extra High | -- | |
97.73Thinking Level · Extra High | Tools | 96.97Thinking Level · High | Tools | -- |
Single-Benchmark Version Trend
Viewing: ARC-AGI-1 · Abstract Generalization
Standard API Pricing Across the GPT-5.4 Series
Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.
Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens
| Model | Supplier | Standard input | Standard output | Base price applies to |
|---|---|---|---|---|
GPT-5.4 | OpenAI | $2.5 / 1M tokens | $15 / 1M tokens | — |
GPT-5.2 | OpenAI | $1.75 / 1M tokens | $14 / 1M tokens | — |
GPT-5.1 | OpenAI | $1.25 / 1M tokens | $10 / 1M tokens | — |

