Grok 4.6 Benchmark Details
Grok 4.6 currently shows benchmark results led by τ³-Banking (3 / 167, score 50.70), GPQA Diamond (9 / 253, score 94), ECI (18 / 167, score 156.48). This page also compares it with 3 competitor models and 3 predecessor or same-series models, including performance and pricing views when available.
Benchmark Results
Benchmark Results
Abstract Generalization
10 evaluationsScientific Reasoning
7 evaluationsMemory & Persistence
5 evaluationsRepository Engineering
5 evaluationsCapability Indices
3 evaluationsScientific Computing
5 evaluationsService Workflows
6 evaluationsLegal
3 evaluationsFinance
3 evaluationsMathematics
3 evaluationsAgentic Development
10 evaluationsCode Generation & Editing
2 evaluationsDocuments & Charts
4 evaluationsClinical Workflows
2 evaluationsCompetitor Comparison
Benchmark scores for Grok 4.6 compared against top models in its class
12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.
| Benchmark | Grok 4.6Current | Claude Opus 5 | GPT-5.6 Sol | GLM-5.3 |
|---|---|---|---|---|
87.50Thinking Level · Medium | 97.50Thinking Level · Extra High | 97.50Thinking Level · Extra High | -- | |
67.08Thinking Level · Extra High | 90.42Thinking Level · High | 92.50Thinking Level · High | -- | |
ARC-AGI-3 (Standard harness) Action efficiency score(以 ARC Prize Standard harness 口径为准) Abstract Generalization | 2.11Thinking Level · Extra High | 30.20Thinking Level · High | 7.78Thinking Level · High | -- |
19.70Thinking Level · Extra High | 29.10Thinking Level · High | 32.30Thinking Level · High | 19.10Thinking Level · High | |
94.00Thinking Level · High | 87.88Thinking Level · Low | 94.60Thinking Level · High | -- | |
30.63Thinking Level · High | Tools | 37.39Thinking Level · High | Tools | 33.33Thinking Level · High | Tools | 22.97Thinking Level · High | Tools | |
75.90Thinking Level · High | 80.60Thinking Level · High | 64.80Thinking Level · Extra High | 66.20Thinking Level · High | |
81.36Thinking Level · Medium | 97.72Thinking Level · High | 97.63Thinking Level · High | 88.53Thinking Level · High | |
30.48Thinking Level · Extra High | Tools | 45.71Thinking Level · High | Tools | 44.76Thinking Level · High | Tools | -- | |
67.48Thinking Level · Medium | Tools | 73.65Thinking Level · High | Tools | 72.70Thinking Level · Extra High | Tools | 68.96Thinking Level · High | Tools | |
61.00Thinking Level · High | Tools | 63.00Thinking Level · Extra High | Tools | 61.00Thinking Level · High | Tools | -- | |
156.48Thinking Level · High | 162.67Thinking Level · High | 161.99Thinking Level · High | 155.56Thinking Level · High |
Standard API Pricing: Grok 4.6 vs. Peer Models
Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.
Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens
When a context threshold exists, the charted base price only applies within these limits:
| Model | Supplier | Standard input | Standard output | Base price applies to |
|---|---|---|---|---|
Grok 4.6 | xAI | $2 / 1M tokens | $6 / 1M tokens | <= 200000 |
Claude Opus 5 | Anthropic | $5 / 1M tokens | $25 / 1M tokens | — |
GPT-5.6 Sol | OpenAI | $4 / 1M tokens | $20 / 1M tokens | — |
GLM-5.3 | 智谱AI | $1.4 / 1M tokens | $4.4 / 1M tokens | — |
Version History
How each version of the Grok 4.6 series stacks up on benchmark tests
12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.
| Benchmark | Grok 4.6Current | Grok 4.5 | Grok 4.20 |
|---|---|---|---|
87.50Thinking Level · Medium | 87.17Thinking Level · Medium | Tools | -- | |
67.08Thinking Level · Extra High | 52.64Thinking Level · High | Tools | -- | |
ARC-AGI-3 (Standard harness) Action efficiency score(以 ARC Prize Standard harness 口径为准) Abstract Generalization | 2.11Thinking Level · Extra High | 0.32Thinking Level · Medium | 0.09Thinking Enabled |
19.70Thinking Level · Extra High | 15.40Thinking Level · High | -- | |
94.00Thinking Level · High | 93.43Thinking Level · High | -- | |
75.90Thinking Level · High | 70.00Thinking Level · High | -- | |
81.36Thinking Level · Medium | -- | 54.00Thinking Enabled | |
56.40Thinking Level · High | Tools | 53.60Thinking Level · High | Tools | -- | |
67.48Thinking Level · Medium | Tools | 53.76Thinking Level · High | Tools | -- | |
61.00Thinking Level · High | Tools | 56.00Thinking Level · High | Tools | -- | |
156.48Thinking Level · High | 154.02Thinking Level · High | 152.01Thinking Level · High | |
59.17Thinking Level · High | Tools | 51.53Thinking Level · High | Tools | 17.55Thinking Enabled | Tools |
Single-Benchmark Version Trend
Viewing: ARC-AGI-1 · Abstract Generalization
Standard API Pricing Across the Grok 4.6 Series
Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.
Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens
When a context threshold exists, the charted base price only applies within these limits:
| Model | Supplier | Standard input | Standard output | Base price applies to |
|---|---|---|---|---|
Grok 4.6 | xAI | $2 / 1M tokens | $6 / 1M tokens | <= 200000 |
Grok 4.5 | xAI | $2 / 1M tokens | $6 / 1M tokens | — |
Grok 4.20 | xAI | $1.25 / 1M tokens | $2.5 / 1M tokens | <= 200000 |