DataLearner logo

Claude Opus 4.8vsGemini 3.1 Pro Preview

Across 13 shared benchmarks, Claude Opus 4.8 leads overall: Claude Opus 4.8 wins 8, Gemini 3.1 Pro Preview wins 5, with 0 ties and an average score difference of +12.04.

Anthropic
Claude Opus 4.8

Anthropic · 2026-05-28 · Reasoning model

Google Deep Mind
Gemini 3.1 Pro Preview

Google Deep Mind · 2026-02-20 · Multimodal model

Claude Opus 4.88 wins(62%)(38%)5 winsGemini 3.1 Pro Preview

Benchmark scores

Grouped by capability, sorted by largest gap within each. 13 shared benchmarks.

Coding and Software Engineer

Claude Opus 4.8 4/5
BenchmarkClaude Opus 4.8Gemini 3.1 Pro PreviewDiff
Text Arena (Coding)1,5459 / 35Normal (No Tools)1,46123 / 35Normal (No Tools)+83.56
DeepSWE5913 / 27Deep Thinking (With Tools)1227 / 27Thinking High (With Tools)+47
SWE-Bench Pro - Public69.204 / 57Extended (with tools)54.2034 / 57Thinking High (With Tools)+15
SWE-bench Verified88.605 / 114Extended (with tools)80.6011 / 114Thinking High (With Tools)+8
WeirdML v270.4518 / 52Normal (With Tools)72.1017 / 52Normal (With Tools)-1.65

AI Agent - Tool Usage

Claude Opus 4.8 3/3
BenchmarkClaude Opus 4.8Gemini 3.1 Pro PreviewDiff
OSWorld-Verified83.404 / 26Extended (with tools)76.2012 / 26Thinking (With Tools)+7.20
Terminal-Bench 2.178.9020 / 44Thinking High (With Tools)73.8028 / 44Thinking High (With Tools)+5.10
MCP-Atlas82.206 / 38Deep Thinking (With Tools)78.2012 / 38Thinking High (With Tools)+4

General Knowledge

Even 2/2
BenchmarkClaude Opus 4.8Gemini 3.1 Pro PreviewDiff
HLE57.908 / 181Extended (with tools)51.4024 / 181Thinking High (With Tools)+6.50
LiveBench78.794 / 115Deep Thinking (No Tools)79.933 / 115Thinking High (No Tools)-1.14

AI Agent - Information Search

Gemini 3.1 Pro Preview 1/1
BenchmarkClaude Opus 4.8Gemini 3.1 Pro PreviewDiff
BrowseComp84.309 / 54Thinking High (With Tools + Internet)85.905 / 54Thinking High (With Tools + Internet)-1.60

Commonsense Reasoning

Gemini 3.1 Pro Preview 1/1
BenchmarkClaude Opus 4.8Gemini 3.1 Pro PreviewDiff
SimpleBench64.8010 / 67Normal (No Tools)79.602 / 67Normal (No Tools)-14.80

General Evaluation

Gemini 3.1 Pro Preview 1/1
BenchmarkClaude Opus 4.8Gemini 3.1 Pro PreviewDiff
GPQA Diamond93.6011 / 226Thinking High (No Tools)94.304 / 226Thinking High (No Tools)-0.70

Specs

FieldClaude Opus 4.8Gemini 3.1 Pro Preview
PublisherAnthropicGoogle Deep Mind
Release date2026-05-282026-02-20
Model typeReasoning modelMultimodal model
ArchitectureDenseDense
ParametersNot availableNot available
Context length1M1M
Max output125K64K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemClaude Opus 4.8Gemini 3.1 Pro Preview
Text input$5 / 1M tokens$2 / 1M tokens
Text output$25 / 1M tokens$12 / 1M tokens
Cache read$0.5 / 1M tokensNot public
Cache write$6.25 / 1M tokensNot public

Summary

  • Claude Opus 4.8leads in:Coding and Software Engineer (4/5), AI Agent - Tool Usage (3/3)
  • Gemini 3.1 Pro Previewleads in:AI Agent - Information Search (1/1), Commonsense Reasoning (1/1), General Evaluation (1/1)
  • Tied in:General Knowledge

On average across the 13 shared benchmarks, Claude Opus 4.8 scores 12.04 higher.

Largest single-benchmark gap: Text Arena (Coding) — Claude Opus 4.8 1,545 vs Gemini 3.1 Pro Preview 1,461 (+83.56).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.