DataLearner logo

GPT-5.4 minivsHaiku 4.5

Across 10 shared benchmarks, Haiku 4.5 leads overall: GPT-5.4 mini wins 4, Haiku 4.5 wins 6, with 0 ties and an average score difference of -3.03.

OpenAI
GPT-5.4 mini

OpenAI · 2026-03-17 · Reasoning model

Anthropic
Haiku 4.5

Anthropic · 2025-10-15 · Multimodal model

GPT-5.4 mini4 wins(40%)(60%)6 winsHaiku 4.5

Benchmark scores

Grouped by capability, sorted by largest gap within each. 10 shared benchmarks.

Agent Level Benchmark

Haiku 4.5 2/2
BenchmarkGPT-5.4 miniHaiku 4.5Diff
τ²-Bench - Telecom23.40234 / 264Normal (With Tools)32.50204 / 264Normal (With Tools)-9.10
Terminal Bench Hard18.20154 / 244Normal (With Tools)27.30120 / 244Normal (With Tools)-9.10

Claw-style Agent Evaluation

Even 2/2
BenchmarkGPT-5.4 miniHaiku 4.5Diff
Claw Bench75.3025 / 29Thinking (With Tools)89.4011 / 29Thinking (With Tools)-14.10
PinchBench v279.2313 / 45Reported best (effort unspecified)67.7028 / 45Reported best (effort unspecified)+11.53

General Evaluation

GPT-5.4 mini 1/1
BenchmarkGPT-5.4 miniHaiku 4.5Diff
GPQA Diamond64.14362 / 462Normal (No Tools)60.50375 / 462Normal (No Tools)+3.64

General Knowledge

Haiku 4.5 1/1
BenchmarkGPT-5.4 miniHaiku 4.5Diff
LiveBench36.95114 / 117Normal (No Tools)45.33105 / 117Normal (No Tools)-8.38

Instruction Following

Haiku 4.5 1/1
BenchmarkGPT-5.4 miniHaiku 4.5Diff
IF Bench38.80225 / 282Normal (No Tools)42202 / 282Normal (No Tools)-3.20

Long Context

Haiku 4.5 1/1
BenchmarkGPT-5.4 miniHaiku 4.5Diff
AA-LCR37158 / 170Normal (No Tools)49.70139 / 170Normal (No Tools)-12.70

Multimodal Understanding

GPT-5.4 mini 1/1
BenchmarkGPT-5.4 miniHaiku 4.5Diff
MMMU-Pro60.50172 / 227Normal (No Tools)55.10188 / 227Normal (No Tools)+5.40

Text Embedding

GPT-5.4 mini 1/1
BenchmarkGPT-5.4 miniHaiku 4.5Diff
Context Arena23.34119 / 126Normal (No Tools)17.68124 / 126Normal (No Tools)+5.66

Specs

FieldGPT-5.4 miniHaiku 4.5
PublisherOpenAIAnthropic
Release date2026-03-172025-10-15
Model typeReasoning modelMultimodal model
ArchitectureDenseDense
ParametersNot availableNot available
Context length400K200K
Max output128K64K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemGPT-5.4 miniHaiku 4.5
Text input$0.75 / 1M tokens$1 / 1M tokens
Text output$4.5 / 1M tokens$5 / 1M tokens
Cache read$0.075 / 1M tokens$0.1 / 1M tokens
Cache writeNot public$1.25 / 1M tokens

Summary

  • GPT-5.4 minileads in:General Evaluation (1/1), Multimodal Understanding (1/1), Text Embedding (1/1)
  • Haiku 4.5leads in:Agent Level Benchmark (2/2), General Knowledge (1/1), Instruction Following (1/1), Long Context (1/1)
  • Tied in:Claw-style Agent Evaluation

On average across the 10 shared benchmarks, Haiku 4.5 scores 3.03 higher.

Largest single-benchmark gap: Claw Bench — GPT-5.4 mini 75.30 vs Haiku 4.5 89.40 (-14.10).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.