DataLearner logo

GPT-5.4 minivsGemini 3.0 Flash

Across 9 shared benchmarks, Gemini 3.0 Flash leads overall: GPT-5.4 mini wins 1, Gemini 3.0 Flash wins 8, with 0 ties and an average score difference of -12.97.

OpenAI
GPT-5.4 mini

OpenAI · 2026-03-17 · Reasoning model

Google Deep Mind
Gemini 3.0 Flash

Google Deep Mind · 2025-12-17 · Chat model

GPT-5.4 mini1 win(11%)(89%)8 winsGemini 3.0 Flash

Benchmark scores

Grouped by capability, sorted by largest gap within each. 9 shared benchmarks.

Agent Level Benchmark

Gemini 3.0 Flash 2/2
BenchmarkGPT-5.4 miniGemini 3.0 FlashDiff
τ²-Bench - Telecom23.40234 / 264Normal (With Tools)43.30186 / 264Normal (With Tools)-19.90
Terminal Bench Hard18.20154 / 244Normal (With Tools)31.80100 / 244Normal (With Tools)-13.60

Claw-style Agent Evaluation

Even 2/2
BenchmarkGPT-5.4 miniGemini 3.0 FlashDiff
Claw Bench75.3025 / 29Thinking (With Tools)85.7015 / 29Thinking (With Tools)-10.40
PinchBench v279.2313 / 45Reported best (effort unspecified)72.0624 / 45Reported best (effort unspecified)+7.17

General Knowledge

Gemini 3.0 Flash 2/2
BenchmarkGPT-5.4 miniGemini 3.0 FlashDiff
LiveBench36.95114 / 117Normal (No Tools)56.3581 / 117Normal (No Tools)-19.40
HLE5.90462 / 563Normal (No Tools) · Text only15334 / 563Normal (No Tools) · Text only-9.10

General Evaluation

Gemini 3.0 Flash 1/1
BenchmarkGPT-5.4 miniGemini 3.0 FlashDiff
GPQA Diamond64.14362 / 462Normal (No Tools)81.20215 / 462Normal (No Tools)-17.06

Instruction Following

Gemini 3.0 Flash 1/1
BenchmarkGPT-5.4 miniGemini 3.0 FlashDiff
IF Bench38.80225 / 282Normal (No Tools)55.10133 / 282Normal (No Tools)-16.30

Multimodal Understanding

Gemini 3.0 Flash 1/1
BenchmarkGPT-5.4 miniGemini 3.0 FlashDiff
MMMU-Pro60.50172 / 227Normal (No Tools)78.6055 / 227Normal (No Tools)-18.10

Specs

FieldGPT-5.4 miniGemini 3.0 Flash
PublisherOpenAIGoogle Deep Mind
Release date2026-03-172025-12-17
Model typeReasoning modelChat model
ArchitectureDenseDense
ParametersNot availableNot available
Context length400K2000K
Max output128K64K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemGPT-5.4 miniGemini 3.0 Flash
Text input$0.75 / 1M tokens$0.5 / 1M tokens
Text output$4.5 / 1M tokens$3 / 1M tokens
Cache read$0.075 / 1M tokensNot public

Summary

  • Gemini 3.0 Flashleads in:Agent Level Benchmark (2/2), General Knowledge (2/2), General Evaluation (1/1), Instruction Following (1/1), Multimodal Understanding (1/1)
  • Tied in:Claw-style Agent Evaluation

On average across the 9 shared benchmarks, Gemini 3.0 Flash scores 12.97 higher.

Largest single-benchmark gap: τ²-Bench - Telecom — GPT-5.4 mini 23.40 vs Gemini 3.0 Flash 43.30 (-19.90).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.