DataLearner logo

Grok 4.20 Benchmark Details

Grok 4.20 currently shows benchmark results led by Context Arena (81 / 126, score 54), ARC-AGI-3 (9 / 9, score 0). This page also tracks comparisons against 1 predecessor or same-series models.

Benchmark Results

Grok 4.20

Benchmark Results

Thinking

Text Embedding

2 evaluations
Benchmark / mode
Score
Rank/total
Context Arena
Standard Mode
20.32
123 / 126
Context Arena
Thinking Mode
54
81 / 126

General Knowledge

1 evaluations
Benchmark / mode
Score
Rank/total
ARC-AGI-3
Thinking Mode
0
9 / 9

Version History

How each version of the Grok 4.20 series stacks up on benchmark tests

Grok 4.20Grok 4.1
No benchmark data matches the selected filters.

Standard API Pricing Across the Grok 4.20 Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

Grok 4.20: Base price applies to <= 200000
ModelSupplierStandard inputStandard outputBase price applies to
Grok 4.20
xAI$1.25 / 1M tokens$2.5 / 1M tokens<= 200000