DataLearner logo

Grok 4.7 Benchmark Details

Grok 4.7 currently shows benchmark results led by AA-Briefcase (2 / 85, score 1657), GDPval-AA v2 (6 / 107, score 1695), CursorBench 4.0 (6 / 44, score 46.30). This page also compares it with 2 competitor models and 3 predecessor or same-series models, including performance and pricing views when available.

Benchmark Results

Grok 4.7

Benchmark Results

Thinking

Coding and Software Engineer

2 evaluations
Benchmark / mode
Score
Rank/total
DeepSWE
HighTools
71
16 / 89
46.30
6 / 44

Productivity Knowledge

3 evaluations
Benchmark / mode
Score
Rank/total
GDPval-AA v2
HighTools
1695
6 / 107
AA-Briefcase
HighTools
1657
2 / 85
Harvey Lab-AA
HighTools
19.60
42 / 44

AI Agent - Tool Usage

1 evaluations
Benchmark / mode
Score
Rank/total
38
16 / 91

Other

1 evaluations
Benchmark / mode
Score
Rank/total
56.70
2 / 2

Competitor Comparison

Benchmark scores for Grok 4.7 compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

6 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkGrok 4.7CurrentClaude Opus 5GPT-5.6 Sol
CursorBench 4.0
Task score (%)
编程与软件工程
46.30Thinking Level · High | Tools
46.60Thinking Level · High | Tools
41.70Thinking Level · High | Tools
DeepSWE
Pass@1 (DeepSWE v1.1)
编程与软件工程
71.00Thinking Level · High | Tools
73.65Thinking Level · High | Tools
72.70Thinking Level · Extra High | Tools
AA-Briefcase
AA-Briefcase Elo; Rubric Pass Rate; Analytical Quality Elo; Presentation Elo
生产力知识
1657.00Thinking Level · High | Tools
1645.00Thinking Level · High | Tools
1475.00Thinking Level · High | Tools
GDPval-AA v2
Elo score
生产力知识
1695.00Thinking Level · High | Tools
1735.00Thinking Level · High | Tools
1624.00Thinking Level · High | Tools
Harvey Lab-AA
Criterion pass rate (some entries report all-pass rate)
生产力知识
19.60Thinking Level · High | Tools
93.46Thinking Level · High | Tools
87.18Thinking Level · High | Tools
Terminal-Bench 4.0
Resolution rate (%)
AI Agent - 工具使用
38.00Thinking Level · High | Tools
51.82Thinking Level · High | Tools
37.27Thinking Level · High | Tools

Standard API Pricing: Grok 4.7 vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

Grok 4.7: Base price applies to <= 200000
ModelSupplierStandard inputStandard outputBase price applies to
Grok 4.7
xAI$2 / 1M tokens$6 / 1M tokens<= 200000
Claude Opus 5
Anthropic$5 / 1M tokens$25 / 1M tokens
GPT-5.6 Sol
OpenAI$4 / 1M tokens$20 / 1M tokens

Version History

How each version of the Grok 4.7 series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

6 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkGrok 4.7CurrentGrok 4.6Grok 4.5
CursorBench 4.0
Task score (%)
编程与软件工程
46.30Thinking Level · High | Tools
41.40Thinking Level · Extra High | Tools
--
DeepSWE
Pass@1 (DeepSWE v1.1)
编程与软件工程
71.00Thinking Level · High | Tools
67.48Thinking Level · Medium | Tools
53.76Thinking Level · High | Tools
AA-Briefcase
AA-Briefcase Elo; Rubric Pass Rate; Analytical Quality Elo; Presentation Elo
生产力知识
1657.00Thinking Level · High | Tools
1545.00Thinking Level · Extra High | Tools
1283.00Thinking Level · High | Tools
GDPval-AA v2
Elo score
生产力知识
1695.00Thinking Level · High | Tools
1663.00Thinking Level · Extra High | Tools
1430.00Thinking Level · High | Tools
Harvey Lab-AA
Criterion pass rate (some entries report all-pass rate)
生产力知识
19.60Thinking Level · High | Tools
15.80Thinking Level · High | Tools
92.42Thinking Level · High | Tools
Terminal-Bench 4.0
Resolution rate (%)
AI Agent - 工具使用
38.00Thinking Level · High | Tools
20.30Thinking Level · High | Tools
12.42Thinking Level · High | Tools

Single-Benchmark Version Trend

Viewing: CursorBench 4.0 · 编程与软件工程

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the Grok 4.7 Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

Grok 4.7: Base price applies to <= 200000
Grok 4.6: Base price applies to <= 200000
Grok 4.20: Base price applies to <= 200000
ModelSupplierStandard inputStandard outputBase price applies to
Grok 4.7
xAI$2 / 1M tokens$6 / 1M tokens<= 200000
Grok 4.6
xAI$2 / 1M tokens$6 / 1M tokens<= 200000
Grok 4.5
xAI$2 / 1M tokens$6 / 1M tokens
Grok 4.20
xAI$1.25 / 1M tokens$2.5 / 1M tokens<= 200000