DataLearner logo

Gemini 3.1 Pro PreviewvsClaude Opus 4.6

Across 9 shared benchmarks, Claude Opus 4.6 leads overall: Gemini 3.1 Pro Preview wins 4, Claude Opus 4.6 wins 5, with 0 ties and an average score difference of -44.84.

Google Deep Mind
Gemini 3.1 Pro Preview

Google Deep Mind · 2026-02-20 · Multimodal model

Anthropic
Claude Opus 4.6

Anthropic · 2026-02-05 · Reasoning model

Gemini 3.1 Pro Preview4 wins(44%)(56%)5 winsClaude Opus 4.6

Benchmark scores

Grouped by capability, sorted by largest gap within each. 9 shared benchmarks.

Coding and Software Engineer

Claude Opus 4.6 3/4
BenchmarkGemini 3.1 Pro PreviewClaude Opus 4.6Diff
Text Arena (Coding)1,46123 / 35Normal (No Tools)1,5558 / 35Normal (No Tools)-93.86
SWE-Bench Pro - Commercial32.203 / 3Thinking (With Tools)47.101 / 3Thinking (With Tools)-14.90
GSO22.5510 / 21Normal (With Tools)33.335 / 21Normal (With Tools)-10.78
WeirdML v272.1017 / 52Normal (With Tools)65.9023 / 52Normal (With Tools)+6.20

Claw-style Agent Evaluation

Even 2/2
BenchmarkGemini 3.1 Pro PreviewClaude Opus 4.6Diff
PinchBench v281.019 / 45Reported best (effort unspecified)69.9026 / 45Reported best (effort unspecified)+11.11
Pinch Bench86.7011 / 38Thinking (With Tools)87.408 / 38Thinking (With Tools)-0.70

Commonsense Reasoning

Gemini 3.1 Pro Preview 1/1
BenchmarkGemini 3.1 Pro PreviewClaude Opus 4.6Diff
SimpleBench79.608 / 93Normal (No Tools)67.6019 / 93Normal (No Tools)+12

General Knowledge

Gemini 3.1 Pro Preview 1/1
BenchmarkGemini 3.1 Pro PreviewClaude Opus 4.6Diff
LiveBench77.137 / 117Thinking High (No Tools)74.5217 / 117Thinking High (No Tools)+2.61

Writing and Creative Capabilities

Claude Opus 4.6 1/1
BenchmarkGemini 3.1 Pro PreviewClaude Opus 4.6Diff
Creative Writing1,48952 / 106Normal (No Tools)1,80419 / 106Normal (No Tools)-315.20

Specs

FieldGemini 3.1 Pro PreviewClaude Opus 4.6
PublisherGoogle Deep MindAnthropic
Release date2026-02-202026-02-05
Model typeMultimodal modelReasoning model
ArchitectureDenseDense
ParametersNot availableNot available
Context length1M1000K
Max output64K64K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemGemini 3.1 Pro PreviewClaude Opus 4.6
Text input$2 / 1M tokens$5 / 1M tokens
Text output$12 / 1M tokens$25 / 1M tokens
Cache read$0.2 / 1M tokens$0.5 / 1M tokens
Cache writeNot public$6.25 / 1M tokens

Summary

  • Gemini 3.1 Pro Previewleads in:Commonsense Reasoning (1/1), General Knowledge (1/1)
  • Claude Opus 4.6leads in:Coding and Software Engineer (3/4), Writing and Creative Capabilities (1/1)
  • Tied in:Claw-style Agent Evaluation

On average across the 9 shared benchmarks, Claude Opus 4.6 scores 44.84 higher.

Largest single-benchmark gap: Creative Writing — Gemini 3.1 Pro Preview 1,489 vs Claude Opus 4.6 1,804 (-315.20).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.