DataLearner logo

Opus 4.7vsGemini 3.1 Pro Preview

Across 6 shared benchmarks, Opus 4.7 leads overall: Opus 4.7 wins 4, Gemini 3.1 Pro Preview wins 2, with 0 ties and an average score difference of +87.01.

Anthropic
Opus 4.7

Anthropic · 2026-04-16 · Reasoning model

Google Deep Mind
Gemini 3.1 Pro Preview

Google Deep Mind · 2026-02-20 · Multimodal model

Opus 4.74 wins(67%)(33%)2 winsGemini 3.1 Pro Preview

Benchmark scores

Grouped by capability, sorted by largest gap within each. 6 shared benchmarks.

Coding and Software Engineer

Opus 4.7 3/3
BenchmarkOpus 4.7Gemini 3.1 Pro PreviewDiff
Text Arena (Coding)1,5627 / 35Normal (No Tools)1,46123 / 35Normal (No Tools)+100.90
GSO44.121 / 21Normal (With Tools)22.5510 / 21Normal (With Tools)+21.57
WeirdML v276.4013 / 52Normal (With Tools)72.1017 / 52Normal (With Tools)+4.30

Claw-style Agent Evaluation

Gemini 3.1 Pro Preview 1/1
BenchmarkOpus 4.7Gemini 3.1 Pro PreviewDiff
PinchBench v276.0115 / 45Reported best (effort unspecified)81.019 / 45Reported best (effort unspecified)-5

Commonsense Reasoning

Gemini 3.1 Pro Preview 1/1
BenchmarkOpus 4.7Gemini 3.1 Pro PreviewDiff
SimpleBench61.7027 / 93Normal (No Tools)79.608 / 93Normal (No Tools)-17.90

Writing and Creative Capabilities

Opus 4.7 1/1
BenchmarkOpus 4.7Gemini 3.1 Pro PreviewDiff
Creative Writing1,9079 / 106Normal (No Tools)1,48952 / 106Normal (No Tools)+418.20

Specs

FieldOpus 4.7Gemini 3.1 Pro Preview
PublisherAnthropicGoogle Deep Mind
Release date2026-04-162026-02-20
Model typeReasoning modelMultimodal model
ArchitectureDenseDense
ParametersNot availableNot available
Context length1000K1M
Max output128K64K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemOpus 4.7Gemini 3.1 Pro Preview
Text input$5 / 1M tokens$2 / 1M tokens
Text output$25 / 1M tokens$12 / 1M tokens
Cache read$0.5 / 1M tokens$0.2 / 1M tokens
Cache write$6.25 / 1M tokensNot public

Summary

  • Opus 4.7leads in:Coding and Software Engineer (3/3), Writing and Creative Capabilities (1/1)
  • Gemini 3.1 Pro Previewleads in:Claw-style Agent Evaluation (1/1), Commonsense Reasoning (1/1)

On average across the 6 shared benchmarks, Opus 4.7 scores 87.01 higher.

Largest single-benchmark gap: Creative Writing — Opus 4.7 1,907 vs Gemini 3.1 Pro Preview 1,489 (+418.20).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.