DataLearner logo

GPT-5.5 Instant Benchmark Details

GPT-5.5 Instant currently shows benchmark results led by Terminal Bench Hard (45 / 244, score 42.40), GDP.pdf (39 / 122, score 20.20), SciCode (57 / 134, score 52.50). This page also compares it with 1 competitor models and 1 predecessor or same-series models, including performance and pricing views when available.

Benchmark Results

GPT-5.5 Instant

Benchmark Results

Thinking
Tool usage

Agentic Development

2 evaluations
Benchmark / mode
Score
Rank/total
Terminal Bench Hard
Thinking EnabledTools
42.40
45 / 244
Terminal-Bench 4.0
Thinking EnabledTools
12.60
52 / 98

Scientific Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
CritPt
Thinking Enabled
2.60
125 / 204

Scientific Computing

1 evaluations
Benchmark / mode
Score
Rank/total
SciCode
Thinking Enabled
52.50
57 / 134

Service Workflows

1 evaluations
Benchmark / mode
Score
Rank/total
τ³-Banking
Thinking EnabledTools
11.80
128 / 167

Documents & Charts

1 evaluations
Benchmark / mode
Score
Rank/total
GDP.pdf
Thinking Enabled
20.20
39 / 122

Capability Indices

1 evaluations
Benchmark / mode
Score
Rank/total
ECI
unknown
142.48
87 / 167

Competitor Comparison

Benchmark scores for GPT-5.5 Instant compared against top models in its class

GPT-5.5 InstantKimi K2.5 Instant
No benchmark data matches the selected filters.

Standard API Pricing: GPT-5.5 Instant vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier.

Comparable standard text pricing is not available for these models.

Version History

How each version of the GPT-5.5 Instant series stacks up on benchmark tests

GPT-5.5 InstantGPT-5.4
Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

4 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkGPT-5.5 InstantCurrentGPT-5.4
Terminal Bench Hard
Accuracy
Agentic Development
42.40Thinking Enabled | Tools
57.60Thinking Level · Extra High | Tools
CritPt
Score
Scientific Reasoning
2.60Thinking Enabled
23.40Thinking Level · Extra High
τ³-Banking
Score
Service Workflows
11.80Thinking Enabled | Tools
39.43Thinking Level · Extra High | Tools
ECI
ECI score (capability index, higher is better)
Capability Indices
142.48Thinking Level · High
156.92Thinking Level · High

Single-Benchmark Version Trend

Viewing: Terminal Bench Hard · Agentic Development

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the GPT-5.5 Instant Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
GPT-5.4
OpenAI$2.5 / 1M tokens$15 / 1M tokens—