DataLearner logo

GPT-5.4vsGemini 3.1 Pro Preview

GPT-5.4 and Gemini 3.1 Pro Preview are tied across 4 shared benchmarks: GPT-5.4 leads on 2, Gemini 3.1 Pro Preview leads on 2, with 0 ties and an average score difference of +82.63.

OpenAI
GPT-5.4

OpenAI · 2026-03-05 · Multimodal model

Google Deep Mind
Gemini 3.1 Pro Preview

Google Deep Mind · 2026-02-20 · Multimodal model

GPT-5.42 wins(50%)(50%)2 winsGemini 3.1 Pro Preview

Benchmark scores

Grouped by capability, sorted by largest gap within each. 4 shared benchmarks.

Claw-style Agent Evaluation

Even 2/2
BenchmarkGPT-5.4Gemini 3.1 Pro PreviewDiff
PinchBench v275.7017 / 45Reported best (effort unspecified)81.019 / 45Reported best (effort unspecified)-5.31
Pinch Bench90.501 / 38Thinking (With Tools)86.7011 / 38Thinking (With Tools)+3.80

Coding and Software Engineer

Gemini 3.1 Pro Preview 1/1
BenchmarkGPT-5.4Gemini 3.1 Pro PreviewDiff
WeirdML v257.4431 / 52Normal (With Tools)72.1017 / 52Normal (With Tools)-14.66

Writing and Creative Capabilities

GPT-5.4 1/1
BenchmarkGPT-5.4Gemini 3.1 Pro PreviewDiff
Creative Writing1,83614 / 106Normal (No Tools)1,48952 / 106Normal (No Tools)+346.70

Specs

FieldGPT-5.4Gemini 3.1 Pro Preview
PublisherOpenAIGoogle Deep Mind
Release date2026-03-052026-02-20
Model typeMultimodal modelMultimodal model
ArchitectureDenseDense
ParametersNot availableNot available
Context length1M1M
Max output125K64K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemGPT-5.4Gemini 3.1 Pro Preview
Text input$2.5 / 1M tokens$2 / 1M tokens
Text output$15 / 1M tokens$12 / 1M tokens
Cache read$0.25 / 1M tokens$0.2 / 1M tokens

Summary

  • GPT-5.4leads in:Writing and Creative Capabilities (1/1)
  • Gemini 3.1 Pro Previewleads in:Coding and Software Engineer (1/1)
  • Tied in:Claw-style Agent Evaluation

On average across the 4 shared benchmarks, GPT-5.4 scores 82.63 higher.

Largest single-benchmark gap: Creative Writing — GPT-5.4 1,836 vs Gemini 3.1 Pro Preview 1,489 (+346.70).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.