DataLearner logo

Grok 4.6 Benchmark Details

Grok 4.6 currently shows benchmark results led by τ³-Banking (3 / 167, score 50.70), GPQA Diamond (9 / 253, score 94), ECI (18 / 167, score 156.48). This page also compares it with 3 competitor models and 3 predecessor or same-series models, including performance and pricing views when available.

Benchmark Results

Grok 4.6

Benchmark Results

Thinking
Tool usage
Internet

Abstract Generalization

10 evaluations
Benchmark / mode
Score
Rank/total
74.83
97 / 176
ARC-AGI-1
Medium
87.50
70 / 176
87
73 / 176
ARC-AGI-1
HighTools
87
73 / 176
ARC-AGI-1
Extra-High
87
73 / 176
27.64
100 / 164
ARC-AGI-2
Medium
61.25
64 / 164
65.14
59 / 164
ARC-AGI-2
Extra-High
67.08
54 / 164
2.11
14 / 42

Scientific Reasoning

7 evaluations
Benchmark / mode
Score
Rank/total
94
9 / 253
GPQA Diamond
Extra-High
93.18
19 / 253
30.63
8 / 13
5.70
92 / 204
CritPt
Medium
17.70
46 / 204
CritPt
High
17.10
50 / 204
CritPt
Extra-High
19.70
41 / 204

Commonsense

1 evaluations
Benchmark / mode
Score
Rank/total
75.90
13 / 96

Memory & Persistence

5 evaluations
Benchmark / mode
Score
Rank/total
74.16
52 / 126
81.36
37 / 126
79.03
42 / 126
Context Arena
Extra-High
80.94
38 / 126
EBR-bench
Extra-HighTools
30.48
9 / 23

Repository Engineering

5 evaluations
Benchmark / mode
Score
Rank/total
DeepSWE
LowTools
41.65
76 / 91
DeepSWE
MediumTools
67.48
30 / 91
DeepSWE
HighTools
65.19
41 / 91
DeepSWE
Extra-HighTools
66.74
36 / 91
APEX-SWE
HighTools
56.40
2 / 3

Capability Indices

3 evaluations
Benchmark / mode
Score
Rank/total
ECI
unknown
156.48
18 / 167
Vals Index
HighTools
59.17
17 / 42

Tool Orchestration

1 evaluations
Benchmark / mode
Score
Rank/total
APEX-Agents
HighTools
57.50
2 / 7

Scientific Computing

5 evaluations
Benchmark / mode
Score
Rank/total
49.40
77 / 134
SciCode
Medium
55.90
31 / 134
56.50
25 / 134
SciCode
Extra-High
53
55 / 134
7.10
10 / 12

Service Workflows

6 evaluations
Benchmark / mode
Score
Rank/total
66.85
9 / 26
τ³-Banking
LowTools
38.10
39 / 167
τ³-Banking
MediumTools
44.30
20 / 167
τ³-Banking
HighTools
50.70
3 / 167
τ³-Banking
Extra-HighTools
43.30
22 / 167
SAGE
HighTools
28.90
61 / 64

Legal

3 evaluations
Benchmark / mode
Score
Rank/total
48.08
7 / 42
15.83
5 / 43
Harvey Lab-AA
HighTools
15.80
43 / 44

Finance

3 evaluations
Benchmark / mode
Score
Rank/total
70.79
5 / 20
62.73
22 / 42
53.68
21 / 42

Mathematics

3 evaluations
Benchmark / mode
Score
Rank/total
FrontierMath v2
Extra-High
65.96
18 / 58
51
17 / 28
31.71
17 / 42

Agentic Development

10 evaluations
Benchmark / mode
Score
Rank/total
69.90
3 / 5
33.40
28 / 46
CursorBench 4.0
MediumTools
36.10
23 / 46
40.40
18 / 46
CursorBench 4.0
Extra-HighTools
41.40
15 / 46
26
7 / 11
3
70 / 96
13.10
48 / 96
20.30
36 / 96
Terminal-Bench 4.0
Extra-HighTools
17.20
42 / 96

Code Generation & Editing

2 evaluations
Benchmark / mode
Score
Rank/total
76.24
21 / 60
61.30
4 / 7

Documents & Charts

4 evaluations
Benchmark / mode
Score
Rank/total
15.80
57 / 122
GDP.pdf
Medium
17.80
47 / 122
17
54 / 122
GDP.pdf
Extra-High
17.20
52 / 122

Data Analysis

1 evaluations
Benchmark / mode
Score
Rank/total
AA-AnalystAgent
HighToolsInternet
41.25
12 / 29

Coding Indices

1 evaluations
Benchmark / mode
Score
Rank/total
47
9 / 11

Maintenance & Optimization

1 evaluations
Benchmark / mode
Score
Rank/total
44.57
14 / 43

Algorithmic Coding

1 evaluations
Benchmark / mode
Score
Rank/total
IOI (Vals v2)
HighTools
47.61
19 / 26

Clinical Workflows

2 evaluations
Benchmark / mode
Score
Rank/total
MedScribe
HighTools
86.53
15 / 66
MedCode
HighTools
44.71
29 / 64

Competitor Comparison

Benchmark scores for Grok 4.6 compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkGrok 4.6CurrentClaude Opus 5GPT-5.6 SolGLM-5.3
ARC-AGI-1
Score (%)
Abstract Generalization
87.50Thinking Level · Medium
97.50Thinking Level · Extra High
97.50Thinking Level · Extra High
--
ARC-AGI-2
Score (% solved); cost per task (USD)
Abstract Generalization
67.08Thinking Level · Extra High
90.42Thinking Level · High
92.50Thinking Level · High
--
ARC-AGI-3 (Standard harness)
Action efficiency score(以 ARC Prize Standard harness 口径为准)
Abstract Generalization
2.11Thinking Level · Extra High
30.20Thinking Level · High
7.78Thinking Level · High
--
CritPt
Score
Scientific Reasoning
19.70Thinking Level · Extra High
29.10Thinking Level · High
32.30Thinking Level · High
19.10Thinking Level · High
GPQA Diamond
Accuracy
Scientific Reasoning
94.00Thinking Level · High
87.88Thinking Level · Low
94.60Thinking Level · High
--
MysteryMechanism
Accuracy (%)
Scientific Reasoning
30.63Thinking Level · High | Tools
37.39Thinking Level · High | Tools
33.33Thinking Level · High | Tools
22.97Thinking Level · High | Tools
SimpleBench
Score (AVG@5)
Commonsense
75.90Thinking Level · High
80.60Thinking Level · High
64.80Thinking Level · Extra High
66.20Thinking Level · High
Context Arena
Accuracy (8 needles, 4K-128K context)
Memory & Persistence
81.36Thinking Level · Medium
97.72Thinking Level · High
97.63Thinking Level · High
88.53Thinking Level · High
EBR-bench
Score (% of 21 objectives, best of the two scored playthroughs)
Memory & Persistence
30.48Thinking Level · Extra High | Tools
45.71Thinking Level · High | Tools
44.76Thinking Level · High | Tools
--
DeepSWE
Pass@1 (DeepSWE v1.1)
Repository Engineering
67.48Thinking Level · Medium | Tools
73.65Thinking Level · High | Tools
72.70Thinking Level · Extra High | Tools
68.96Thinking Level · High | Tools
AA Intelligence Index (historical versions)
Historical index (not comparable)
Capability Indices
61.00Thinking Level · High | Tools
63.00Thinking Level · Extra High | Tools
61.00Thinking Level · High | Tools
--
ECI
ECI score (capability index, higher is better)
Capability Indices
156.48Thinking Level · High
162.67Thinking Level · High
161.99Thinking Level · High
155.56Thinking Level · High
29 additional benchmarks remain in the chart above.

Standard API Pricing: Grok 4.6 vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

Grok 4.6: Base price applies to <= 200000
ModelSupplierStandard inputStandard outputBase price applies to
Grok 4.6
xAI$2 / 1M tokens$6 / 1M tokens<= 200000
Claude Opus 5
Anthropic$5 / 1M tokens$25 / 1M tokens—
GPT-5.6 Sol
OpenAI$4 / 1M tokens$20 / 1M tokens—
GLM-5.3
智谱AI$1.4 / 1M tokens$4.4 / 1M tokens—

Version History

How each version of the Grok 4.6 series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkGrok 4.6CurrentGrok 4.5Grok 4.20
ARC-AGI-1
Score (%)
Abstract Generalization
87.50Thinking Level · Medium
87.17Thinking Level · Medium | Tools
--
ARC-AGI-2
Score (% solved); cost per task (USD)
Abstract Generalization
67.08Thinking Level · Extra High
52.64Thinking Level · High | Tools
--
ARC-AGI-3 (Standard harness)
Action efficiency score(以 ARC Prize Standard harness 口径为准)
Abstract Generalization
2.11Thinking Level · Extra High
0.32Thinking Level · Medium
0.09Thinking Enabled
CritPt
Score
Scientific Reasoning
19.70Thinking Level · Extra High
15.40Thinking Level · High
--
GPQA Diamond
Accuracy
Scientific Reasoning
94.00Thinking Level · High
93.43Thinking Level · High
--
SimpleBench
Score (AVG@5)
Commonsense
75.90Thinking Level · High
70.00Thinking Level · High
--
Context Arena
Accuracy (8 needles, 4K-128K context)
Memory & Persistence
81.36Thinking Level · Medium
--
54.00Thinking Enabled
APEX-SWE
Pass@1
Repository Engineering
56.40Thinking Level · High | Tools
53.60Thinking Level · High | Tools
--
DeepSWE
Pass@1 (DeepSWE v1.1)
Repository Engineering
67.48Thinking Level · Medium | Tools
53.76Thinking Level · High | Tools
--
AA Intelligence Index (historical versions)
Historical index (not comparable)
Capability Indices
61.00Thinking Level · High | Tools
56.00Thinking Level · High | Tools
--
ECI
ECI score (capability index, higher is better)
Capability Indices
156.48Thinking Level · High
154.02Thinking Level · High
152.01Thinking Level · High
Vals Index
跨行业任务准确率综合指数
Capability Indices
59.17Thinking Level · High | Tools
51.53Thinking Level · High | Tools
17.55Thinking Enabled | Tools
23 additional benchmarks remain in the chart above.

Single-Benchmark Version Trend

Viewing: ARC-AGI-1 · Abstract Generalization

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the Grok 4.6 Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

Grok 4.6: Base price applies to <= 200000
Grok 4.20: Base price applies to <= 200000
ModelSupplierStandard inputStandard outputBase price applies to
Grok 4.6
xAI$2 / 1M tokens$6 / 1M tokens<= 200000
Grok 4.5
xAI$2 / 1M tokens$6 / 1M tokens—
Grok 4.20
xAI$1.25 / 1M tokens$2.5 / 1M tokens<= 200000