DataLearner logo

GPT-5.1vsClaude Sonnet 4.5

Across 17 shared benchmarks, GPT-5.1 leads overall: GPT-5.1 wins 10, Claude Sonnet 4.5 wins 7, with 0 ties and an average score difference of +3.72.

OpenAI
GPT-5.1

OpenAI · 2025-11-12 · Reasoning model

Anthropic
Claude Sonnet 4.5

Anthropic · 2025-09-30 · Chat model

GPT-5.110 wins(59%)(41%)7 winsClaude Sonnet 4.5

Benchmark scores

Grouped by capability, sorted by largest gap within each. 17 shared benchmarks.

General Knowledge

GPT-5.1 3/5
BenchmarkGPT-5.1Claude Sonnet 4.5Diff
LiveBench42.65106 / 115Normal (No Tools)53.6983 / 115Normal (No Tools)-11.04
ARC-AGI72.8028 / 6863.7035 / 68+9.10
HLE26.5097 / 17233.6080 / 172-7.10
GPQA Diamond88.1031 / 18783.4063 / 187+4.70
ARC-AGI-217.6036 / 6213.6038 / 62+4

Math and Reasoning

GPT-5.1 2/3
BenchmarkGPT-5.1Claude Sonnet 4.5Diff
FrontierMath26.7013 / 60Thinking High (With Tools)5.2038 / 60+21.50
FrontierMath - Tier 412.5029 / 80Thinking High (With Tools)2.1056 / 80Normal (No Tools)+10.40
AIME20259428 / 1071001 / 107-6

Agent Level Benchmark

Even 2/2
BenchmarkGPT-5.1Claude Sonnet 4.5Diff
Terminal Bench Hard432 / 13Thinking High (With Tools)338 / 13+10
τ²-Bench - Telecom95.6014 / 35Thinking High (With Tools)985 / 35-2.40

AI Agent - Tool Usage

Even 2/2
BenchmarkGPT-5.1Claude Sonnet 4.5Diff
MCP-Atlas50.1025 / 27Thinking High (With Tools)59.5021 / 27Thinking (With Tools)-9.40
Terminal Bench 2.047.6038 / 47Thinking High (With Tools)42.8042 / 47+4.80

Coding and Software Engineer

Even 2/2
BenchmarkGPT-5.1Claude Sonnet 4.5Diff
SWE-Bench Pro - Public50.8040 / 54Thinking High (No Tools)43.6047 / 54+7.20
SWE-bench Verified76.3034 / 112828 / 112-5.70

AI Agent - Information Search

GPT-5.1 1/1
BenchmarkGPT-5.1Claude Sonnet 4.5Diff
BrowseComp50.8043 / 53Thinking High (No Tools)24.1051 / 53+26.70

Commonsense Reasoning

Claude Sonnet 4.5 1/1
BenchmarkGPT-5.1Claude Sonnet 4.5Diff
Simple Bench53.2023 / 63Thinking High (No Tools)54.3022 / 63Normal (No Tools)-1.10

Multimodal Understanding

GPT-5.1 1/1
BenchmarkGPT-5.1Claude Sonnet 4.5Diff
MMMU85.402 / 2977.8015 / 29+7.60

Specs

FieldGPT-5.1Claude Sonnet 4.5
PublisherOpenAIAnthropic
Release date2025-11-122025-09-30
Model typeReasoning modelChat model
ArchitectureDenseDense
ParametersNot availableNot available
Context length400K1000K
Max output128K64K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemGPT-5.1Claude Sonnet 4.5
Text input$1.25 / 1M tokens$3 / 1M tokens
Text output$10 / 1M tokens$15 / 1M tokens
Cache read$0.125 / 1M tokens$0.3 / 1M tokens
Cache write$0 / 1M tokens$3.75 / 1M tokens

Summary

  • GPT-5.1leads in:General Knowledge (3/5), Math and Reasoning (2/3), AI Agent - Information Search (1/1), Multimodal Understanding (1/1)
  • Claude Sonnet 4.5leads in:Commonsense Reasoning (1/1)
  • Tied in:Agent Level Benchmark, AI Agent - Tool Usage, Coding and Software Engineer

On average across the 17 shared benchmarks, GPT-5.1 scores 3.72 higher.

Largest single-benchmark gap: BrowseComp — GPT-5.1 50.80 vs Claude Sonnet 4.5 24.10 (+26.70).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.