DataLearner logo

Claude Mythos PreviewvsGemini 3.1 Pro Preview

Across 8 shared benchmarks, Claude Mythos Preview leads overall: Claude Mythos Preview wins 7, Gemini 3.1 Pro Preview wins 1, with 0 ties and an average score difference of +90.88.

Anthropic
Claude Mythos Preview

Anthropic · 2026-04-07 · Chat model

Google Deep Mind
Gemini 3.1 Pro Preview

Google Deep Mind · 2026-02-20 · Multimodal model

Claude Mythos Preview7 wins(88%)(13%)1 winGemini 3.1 Pro Preview

Benchmark scores

Grouped by capability, sorted by largest gap within each. 8 shared benchmarks.

AI Agent - Tool Usage

Claude Mythos Preview 2/2
BenchmarkClaude Mythos PreviewGemini 3.1 Pro PreviewDiff
Terminal Bench 2.0822 / 48Extended (with tools)68.508 / 48Thinking High (With Tools)+13.50
OSWorld-Verified79.608 / 26Extended (with tools)76.2012 / 26Thinking (With Tools)+3.40

Coding and Software Engineer

Claude Mythos Preview 2/2
BenchmarkClaude Mythos PreviewGemini 3.1 Pro PreviewDiff
SWE-Bench Pro - Public77.803 / 57Extended (with tools)54.2034 / 57Thinking High (With Tools)+23.60
SWE-bench Verified93.904 / 114Extended (with tools)80.6011 / 114Thinking High (With Tools)+13.30

Agent Level Benchmark

Claude Mythos Preview 1/1
BenchmarkClaude Mythos PreviewGemini 3.1 Pro PreviewDiff
METR Time Horizons v1.11,0451 / 22Normal (With Tools)384.152 / 22Normal (With Tools)+660.63

AI Agent - Information Search

Gemini 3.1 Pro Preview 1/1
BenchmarkClaude Mythos PreviewGemini 3.1 Pro PreviewDiff
BrowseComp84.906 / 54Extended (with tools)85.905 / 54Thinking High (With Tools + Internet)-1

General Evaluation

Claude Mythos Preview 1/1
BenchmarkClaude Mythos PreviewGemini 3.1 Pro PreviewDiff
GPQA Diamond94.601 / 226Extended (no tools)94.304 / 226Thinking High (No Tools)+0.30

General Knowledge

Claude Mythos Preview 1/1
BenchmarkClaude Mythos PreviewGemini 3.1 Pro PreviewDiff
HLE64.701 / 181Extended (with tools)51.4024 / 181Thinking High (With Tools)+13.30

Specs

FieldClaude Mythos PreviewGemini 3.1 Pro Preview
PublisherAnthropicGoogle Deep Mind
Release date2026-04-072026-02-20
Model typeChat modelMultimodal model
ArchitectureDenseDense
ParametersNot availableNot available
Context lengthNot available1M
Max output8K64K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemClaude Mythos PreviewGemini 3.1 Pro Preview
Text input$25 / 1M tokens$2 / 1M tokens
Text output$125 / 1M tokens$12 / 1M tokens

Summary

  • Claude Mythos Previewleads in:AI Agent - Tool Usage (2/2), Coding and Software Engineer (2/2), Agent Level Benchmark (1/1), General Evaluation (1/1), General Knowledge (1/1)
  • Gemini 3.1 Pro Previewleads in:AI Agent - Information Search (1/1)

On average across the 8 shared benchmarks, Claude Mythos Preview scores 90.88 higher.

Largest single-benchmark gap: METR Time Horizons v1.1 — Claude Mythos Preview 1,045 vs Gemini 3.1 Pro Preview 384.15 (+660.63).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.