DataLearner logo

Grok 4.3 Beta Benchmark Details

Grok 4.3 Beta currently shows benchmark results led by GPQA Diamond (53 / 253, score 88.83), Terminal Bench Hard (59 / 244, score 37.90), MMMU-Pro (63 / 229, score 78.10). This page also tracks comparisons against 2 predecessor or same-series models.

Benchmark Results

Grok 4.3 Beta

Benchmark Results

Thinking
Tool usage
Internet

Scientific Reasoning

4 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Thinking Level · High
88.83
53 / 253
CritPt
Thinking Level · Low
0.60
172 / 204
CritPt
Thinking Level · Medium
4.90
102 / 204
CritPt
Thinking Level · High
8
85 / 204

Cross-capability Suites

1 evaluations
Benchmark / mode
Score
Rank/total
LiveBench
Standard Mode
62.25
59 / 117

Agentic Development

4 evaluations
Benchmark / mode
Score
Rank/total
Terminal Bench Hard
Standard ModeTools
18.90
150 / 244
Terminal Bench Hard
Thinking Level · LowTools
26.50
123 / 244
Terminal Bench Hard
Thinking Level · MediumTools
30.30
110 / 244
Terminal Bench Hard
Thinking Level · HighTools
37.90
59 / 244

Visual Understanding

4 evaluations
Benchmark / mode
Score
Rank/total
MMMU-Pro
Standard Mode
64.80
155 / 229
MMMU-Pro
Thinking Level · Low
72.80
113 / 229
MMMU-Pro
Thinking Level · Medium
75.80
80 / 229
MMMU-Pro
Thinking Level · High
78.10
63 / 229

Scientific Computing

2 evaluations
Benchmark / mode
Score
Rank/total
SciCode
Standard Mode
39.40
109 / 134
SciCode
Thinking Level · High
48.30
80 / 134

Service Workflows

2 evaluations
Benchmark / mode
Score
Rank/total
τ³-Banking
Standard ModeTools
8
142 / 167
τ³-Banking
Thinking Level · HighTools
12.40
124 / 167

Legal

1 evaluations
Benchmark / mode
Score
Rank/total
Harvey Lab-AA
Thinking Level · HighTools
68.39
35 / 44

Mathematics

2 evaluations
Benchmark / mode
Score
Rank/total
FrontierMath v2
Thinking Level · High
42.81
33 / 58
FrontierMath Tier 4 v2
Thinking Level · High
14.63
30 / 42

Documents & Charts

2 evaluations
Benchmark / mode
Score
Rank/total
GDP.pdf
Standard Mode
2.80
106 / 122
GDP.pdf
Thinking Level · High
5.80
97 / 122

Data Analysis

1 evaluations
Benchmark / mode
Score
Rank/total
AA-AnalystAgent
Thinking Level · HighToolsInternet
8.75
26 / 29

Tool Orchestration

1 evaluations
Benchmark / mode
Score
Rank/total
73.73
22 / 45

Capability Indices

1 evaluations
Benchmark / mode
Score
Rank/total
ECI
unknown
149.13
52 / 167

Version History

How each version of the Grok 4.3 Beta series stacks up on benchmark tests

Grok 4.3 BetaGrok 4.20Grok 4.1
Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

3 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkGrok 4.3 BetaCurrentGrok 4.20
τ³-Banking
Score
Service Workflows
12.40Thinking Level · High | Tools
18.00Thinking Level · High | Tools
PinchBench v2
Average score (%)
Tool Orchestration
73.73Thinking Level · High
80.33Thinking Level · High
ECI
ECI score (capability index, higher is better)
Capability Indices
149.13Thinking Level · High
152.01Thinking Level · High

Single-Benchmark Version Trend

Viewing: τ³-Banking · Service Workflows

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the Grok 4.3 Beta Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

Grok 4.3 Beta: Base price applies to <= 200000
Grok 4.20: Base price applies to <= 200000
ModelSupplierStandard inputStandard outputBase price applies to
Grok 4.3 Beta
xAI$1.25 / 1M tokens$2.5 / 1M tokens<= 200000
Grok 4.20
xAI$1.25 / 1M tokens$2.5 / 1M tokens<= 200000