Grok 4.3 Beta Benchmark Details
Grok 4.3 Beta currently shows benchmark results led by GPQA Diamond (53 / 253, score 88.83), Terminal Bench Hard (59 / 244, score 37.90), MMMU-Pro (63 / 229, score 78.10). This page also tracks comparisons against 2 predecessor or same-series models.
Benchmark Results
Benchmark Results
Scientific Reasoning
4 evaluationsAgentic Development
4 evaluationsVisual Understanding
4 evaluationsScientific Computing
2 evaluationsService Workflows
2 evaluationsMathematics
2 evaluationsDocuments & Charts
2 evaluationsVersion History
How each version of the Grok 4.3 Beta series stacks up on benchmark tests
3 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.
| Benchmark | Grok 4.3 BetaCurrent | Grok 4.20 |
|---|---|---|
12.40Thinking Level · High | Tools | 18.00Thinking Level · High | Tools | |
73.73Thinking Level · High | 80.33Thinking Level · High | |
149.13Thinking Level · High | 152.01Thinking Level · High |
Single-Benchmark Version Trend
Viewing: τ³-Banking · Service Workflows
Standard API Pricing Across the Grok 4.3 Beta Series
Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.
Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens
When a context threshold exists, the charted base price only applies within these limits:
| Model | Supplier | Standard input | Standard output | Base price applies to |
|---|---|---|---|---|
Grok 4.3 Beta | xAI | $1.25 / 1M tokens | $2.5 / 1M tokens | <= 200000 |
Grok 4.20 | xAI | $1.25 / 1M tokens | $2.5 / 1M tokens | <= 200000 |