DataLearner logo

Grok 4.6vsClaude Opus 5

Across 3 shared benchmarks, Claude Opus 5 leads overall: Grok 4.6 wins 0, Claude Opus 5 wins 3, with 0 ties and an average score difference of -84.63.

xAI
Grok 4.6

xAI · 2026-08-12 · Coding model

Anthropic
Claude Opus 5

Anthropic · 2026-07-24 · Reasoning model

Grok 4.60 wins(0%)(100%)3 winsClaude Opus 5

Benchmark scores

Grouped by capability, sorted by largest gap within each. 3 shared benchmarks.

Productivity Knowledge

Claude Opus 5 2/2
BenchmarkGrok 4.6Claude Opus 5Diff
AA-Briefcase1,5772 / 6Thinking High (With Tools)1,7201 / 6Max (With Tools)-143
GDPval-AA v21,7533 / 13Thinking High (With Tools)1,8611 / 13Max (With Tools)-108

Coding and Software Engineer

Claude Opus 5 1/1
BenchmarkGrok 4.6Claude Opus 5Diff
DeepSWE65.909 / 27Thinking High (With Tools)68.804 / 27Max (With Tools)-2.90

Specs

FieldGrok 4.6Claude Opus 5
PublisherxAIAnthropic
Release date2026-08-122026-07-24
Model typeCoding modelReasoning model
ArchitectureDenseDense
ParametersNot availableNot available
Context length500K1M
Max outputNot available128K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemGrok 4.6Claude Opus 5
Text input$2 / 1M tokens$5 / 1M tokens
Text output$6 / 1M tokens$25 / 1M tokens
Cache read$0.5 / 1M tokens$0.5 / 1M tokens
Cache writeNot public$6.25 / 1M tokens

Summary

  • Claude Opus 5leads in:Productivity Knowledge (2/2), Coding and Software Engineer (1/1)

On average across the 3 shared benchmarks, Claude Opus 5 scores 84.63 higher.

Largest single-benchmark gap: AA-Briefcase — Grok 4.6 1,577 vs Claude Opus 5 1,720 (-143).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.