DataLearner logo

Grok 4.6vsGrok 4.5

Across 10 shared benchmarks, Grok 4.6 leads overall: Grok 4.6 wins 10, Grok 4.5 wins 0, with 0 ties and an average score difference of +54.32.

xAI
Grok 4.6

xAI · 2026-08-12 · Coding model

xAI
Grok 4.5

xAI · 2026-07-08 · Coding model

Grok 4.610 wins(100%)(0%)0 winsGrok 4.5

Benchmark scores

Grouped by capability, sorted by largest gap within each. 10 shared benchmarks.

Coding and Software Engineer

Grok 4.6 4/4
BenchmarkGrok 4.6Grok 4.5Diff
DeepSWE65.909 / 26Thinking High (With Tools)5317 / 26Thinking High (With Tools)+12.90
FrontierCode 1.161.302 / 5Thinking High (With Tools)56.604 / 5Thinking High (With Tools)+4.70
CursorBench 3.269.902 / 4Thinking High (With Tools)66.704 / 4Thinking High (With Tools)+3.20
APEX-SWE56.402 / 3Thinking High (With Tools)53.603 / 3Thinking High (With Tools)+2.80

Productivity Knowledge

Grok 4.6 3/3
BenchmarkGrok 4.6Grok 4.5Diff
AA-Briefcase1,5772 / 6Thinking High (With Tools)1,3136 / 6Thinking High (With Tools)+264
GDPval-AA v21,7533 / 13Thinking High (With Tools)1,5267 / 13Thinking High (With Tools)+227
Harvey Lab-AA15.803 / 6Thinking High (With Tools)12.904 / 6Thinking High (With Tools)+2.90

Agent Level Benchmark

Grok 4.6 1/1
BenchmarkGrok 4.6Grok 4.5Diff
APEX-Agents57.502 / 5Thinking High (With Tools)47.104 / 5Thinking High (With Tools)+10.40

AI Agent - Tool Usage

Grok 4.6 1/1
BenchmarkGrok 4.6Grok 4.5Diff
Terminal-Bench 3.0264 / 6Thinking High (With Tools)15.705 / 6Thinking High (With Tools)+10.30

General Knowledge

Grok 4.6 1/1
BenchmarkGrok 4.6Grok 4.5Diff
AA Intelligence Index612 / 8Thinking High (With Tools)564 / 8Thinking High (With Tools)+5

Specs

FieldGrok 4.6Grok 4.5
PublisherxAIxAI
Release date2026-08-122026-07-08
Model typeCoding modelCoding model
ArchitectureDenseDense
ParametersNot availableNot available
Context length500K500K
Max outputNot availableNot available

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemGrok 4.6Grok 4.5
Text input$2 / 1M tokens$2 / 1M tokens
Text output$6 / 1M tokens$6 / 1M tokens
Cache read$0.5 / 1M tokens$0.5 / 1M tokens

Summary

  • Grok 4.6leads in:Coding and Software Engineer (4/4), Productivity Knowledge (3/3), Agent Level Benchmark (1/1), AI Agent - Tool Usage (1/1), General Knowledge (1/1)

On average across the 10 shared benchmarks, Grok 4.6 scores 54.32 higher.

Largest single-benchmark gap: AA-Briefcase — Grok 4.6 1,577 vs Grok 4.5 1,313 (+264).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.