DataLearner logo

GPT-5.1vsGemini 2.5-Pro

Across 15 shared benchmarks, GPT-5.1 leads overall: GPT-5.1 wins 13, Gemini 2.5-Pro wins 2, with 0 ties and an average score difference of +12.83.

OpenAI
GPT-5.1

OpenAI · 2025-11-12 · Reasoning model

Google Deep Mind
Gemini 2.5-Pro

Google Deep Mind · 2025-06-05 · Reasoning model

GPT-5.113 wins(87%)(13%)2 winsGemini 2.5-Pro

Benchmark scores

Grouped by capability, sorted by largest gap within each. 15 shared benchmarks.

General Knowledge

GPT-5.1 4/5
BenchmarkGPT-5.1Gemini 2.5-ProDiff
ARC-AGI72.8028 / 683750 / 68+35.80
LiveBench42.65106 / 115Normal (No Tools)58.3376 / 115Thinking High (No Tools)-15.68
ARC-AGI-217.6036 / 624.9047 / 62+12.70
HLE26.5097 / 17221.60112 / 172+4.90
GPQA Diamond88.1031 / 18786.4045 / 187+1.70

Math and Reasoning

GPT-5.1 3/3
BenchmarkGPT-5.1Gemini 2.5-ProDiff
FrontierMath26.7013 / 60Thinking High (With Tools)1123 / 60+15.70
FrontierMath - Tier 412.5029 / 80Thinking High (With Tools)2.1056 / 80Normal (No Tools)+10.40
AIME20259428 / 1078844 / 107+6

Agent Level Benchmark

GPT-5.1 2/2
BenchmarkGPT-5.1Gemini 2.5-ProDiff
τ²-Bench - Telecom95.6014 / 35Thinking High (With Tools)5432 / 35+41.60
Terminal Bench Hard432 / 13Thinking High (With Tools)2512 / 13+18

AI Agent - Information Search

GPT-5.1 1/1
BenchmarkGPT-5.1Gemini 2.5-ProDiff
BrowseComp50.8043 / 53Thinking High (No Tools)7.8052 / 53+43

AI Agent - Tool Usage

GPT-5.1 1/1
BenchmarkGPT-5.1Gemini 2.5-ProDiff
Terminal Bench 2.047.6038 / 47Thinking High (With Tools)32.6047 / 47+15

Coding and Software Engineer

GPT-5.1 1/1
BenchmarkGPT-5.1Gemini 2.5-ProDiff
SWE-bench Verified76.3034 / 11267.2072 / 112+9.10

Commonsense Reasoning

Gemini 2.5-Pro 1/1
BenchmarkGPT-5.1Gemini 2.5-ProDiff
Simple Bench53.2023 / 63Thinking High (No Tools)62.4011 / 63Thinking (No Tools)-9.20

Multimodal Understanding

GPT-5.1 1/1
BenchmarkGPT-5.1Gemini 2.5-ProDiff
MMMU85.402 / 298210 / 29+3.40

Specs

FieldGPT-5.1Gemini 2.5-Pro
PublisherOpenAIGoogle Deep Mind
Release date2025-11-122025-06-05
Model typeReasoning modelReasoning model
ArchitectureDenseDense
ParametersNot availableNot available
Context length400K1000K
Max output128K64K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemGPT-5.1Gemini 2.5-Pro
Text input$1.25 / 1M tokens$1.25 / 1M tokens
Text output$10 / 1M tokens$10 / 1M tokens
Cache read$0.125 / 1M tokensNot public
Cache write$0 / 1M tokensNot public

Summary

  • GPT-5.1leads in:General Knowledge (4/5), Math and Reasoning (3/3), Agent Level Benchmark (2/2), AI Agent - Information Search (1/1), AI Agent - Tool Usage (1/1), Coding and Software Engineer (1/1), Multimodal Understanding (1/1)
  • Gemini 2.5-Proleads in:Commonsense Reasoning (1/1)

On average across the 15 shared benchmarks, GPT-5.1 scores 12.83 higher.

Largest single-benchmark gap: BrowseComp — GPT-5.1 50.80 vs Gemini 2.5-Pro 7.80 (+43).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.