DataLearner logo

Grok 4.7vsGrok 4.5

Across 5 shared benchmarks, Grok 4.7 leads overall: Grok 4.7 wins 4, Grok 4.5 wins 1, with 0 ties and an average score difference of +121.80.

xAI
Grok 4.7

xAI · 2026-09-21 · Coding model

xAI
Grok 4.5

xAI · 2026-07-08 · Coding model

Grok 4.74 wins(80%)(20%)1 winGrok 4.5

Benchmark scores

Grouped by capability, sorted by largest gap within each. 5 shared benchmarks.

Productivity Knowledge

Grok 4.7 2/3
BenchmarkGrok 4.7Grok 4.5Diff
AA-Briefcase1,6572 / 85Thinking High (With Tools)1,28337 / 85Thinking High (With Tools)+374
GDPval-AA v21,6956 / 107Thinking High (With Tools)1,43042 / 107Thinking High (With Tools)+265
Harvey Lab-AA19.6042 / 44Thinking High (With Tools)92.429 / 44Thinking High (With Tools)-72.82

AI Agent - Tool Usage

Grok 4.7 1/1
BenchmarkGrok 4.7Grok 4.5Diff
Terminal-Bench 4.03816 / 91Thinking High (With Tools)12.4248 / 91Thinking High (With Tools)+25.58

Coding and Software Engineer

Grok 4.7 1/1
BenchmarkGrok 4.7Grok 4.5Diff
DeepSWE7116 / 89Thinking High (With Tools)53.7657 / 89Thinking High (With Tools)+17.24

Specs

FieldGrok 4.7Grok 4.5
PublisherxAIxAI
Release date2026-09-212026-07-08
Model typeCoding modelCoding model
ArchitectureDenseDense
ParametersNot availableNot available
Context length500K500K
Max outputNot availableNot available

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemGrok 4.7Grok 4.5
Text input$2 / 1M tokens$2 / 1M tokens
Text output$6 / 1M tokens$6 / 1M tokens
Cache read$0.5 / 1M tokens$0.5 / 1M tokens

Summary

  • Grok 4.7leads in:Productivity Knowledge (2/3), AI Agent - Tool Usage (1/1), Coding and Software Engineer (1/1)

On average across the 5 shared benchmarks, Grok 4.7 scores 121.80 higher.

Largest single-benchmark gap: AA-Briefcase — Grok 4.7 1,657 vs Grok 4.5 1,283 (+374).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.