Grok 4.5 Benchmark Details
Grok 4.5 currently shows benchmark results led by GPQA Diamond (18 / 271, score 93.43), SWE-Bench Pro - Public (7 / 60, score 64.70), SimpleBench (15 / 92, score 70). This page also compares it with 6 competitor models and 3 predecessor or same-series models, including performance and pricing views when available.
Benchmark Results
Benchmark Results
Writing and Creative Capabilities
1 evaluationsCoding and Software Engineer
6 evaluationsAI Agent - Tool Usage
3 evaluationsProductivity Knowledge
2 evaluationsMath and Reasoning
2 evaluationsCompetitor Comparison
Benchmark scores for Grok 4.5 compared against top models in its class
12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.
| Benchmark | Grok 4.5Current | GPT-5.6 Sol | Claude Fable 5 | GLM-5.2 | Claude Sonnet 5 | GPT-5.6 Terra | Kimi K3 |
|---|---|---|---|---|---|---|---|
GPQA Diamond 科学与综合推理 | 93.43Thinking Level · High | 93.50Thinking Level · High | 85.86Thinking Level · High | 91.86Thinking Level · High | 90.53Thinking Level · Extra High | 93.31Thinking Level · High | 93.50Thinking Level · High |
Creative Writing 写作和创作 | 1576.00Standard Mode | 1964.10Standard Mode | 1933.20Standard Mode | 1750.90Standard Mode | 1787.60Standard Mode | 1850.00Standard Mode | 2070.80Standard Mode |
SimpleBench 常识推理 | 70.00Thinking Level · High | 64.80Thinking Level · Extra High | 81.90Standard Mode | 58.80Standard Mode | 60.60Standard Mode | 48.90Thinking Level · Extra High | 60.70Thinking Level · High |
APEX-SWE 编程与软件工程 | 53.60Thinking Level · High | Tools | -- | 58.80Thinking Level · High | Tools | -- | -- | -- | -- |
CursorBench 3.2 编程与软件工程 | 66.70Thinking Level · High | Tools | 67.20Thinking Level · High | Tools | 70.50Thinking Level · High | Tools | -- | -- | -- | -- |
DeepSWE 编程与软件工程 | 53.00Thinking Level · High | Tools | 72.70Thinking Level · Extra High | Tools | 70.00Deep Thinking Mode | Tools | 44.00Deep Thinking Mode | Tools | 54.00Deep Thinking Mode | Tools | 69.60Thinking Level · Extra High | Tools | 67.50Thinking Level · High | Tools |
FrontierCode 1.1 编程与软件工程 | 56.60Thinking Level · High | Tools | 60.60Thinking Level · High | Tools | 63.60Thinking Level · High | Tools | -- | -- | -- | -- |
SWE-Bench Pro - Public 编程与软件工程 | 64.70Thinking Level · High | Tools | 64.60Thinking Level · Extra High | Tools | 80.30Deep Thinking Mode | Tools | 62.10Thinking Enabled | Tools | -- | -- | -- |
SWE-Marathon 编程与软件工程 | 29.00Thinking Level · High | Tools | -- | -- | 13.00Thinking Level · High | Tools | -- | -- | 42.00Thinking Level · High | Tools |
Terminal-Bench 2.1 AI Agent - 工具使用 | 83.30Thinking Level · High | Tools | 88.80Thinking Level · High | 88.00Thinking Level · High | Tools | 81.00Thinking Level · High | Tools | 80.40Thinking Level · Extra High | Tools | 87.40Thinking Level · High | 88.30Thinking Level · High | Tools |
Terminal-Bench 3.0 AI Agent - 工具使用 | 15.70Thinking Level · High | Tools | 34.60Thinking Level · High | Tools | 34.10Thinking Level · High | Tools | -- | -- | -- | -- |
Terminal-Bench 4.0 AI Agent - 工具使用 | 12.42Thinking Level · High | Tools | 37.27Thinking Level · High | Tools | 44.55Thinking Level · High | Tools | -- | 12.42Thinking Level · High | Tools | 21.52Thinking Level · High | Tools | -- |
Standard API Pricing: Grok 4.5 vs. Peer Models
Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.
Source: DataLearnerAI. Standard text prices shown here use the default supplier.
These models use different currencies or billing units, so the page falls back to raw price values instead of a shared bar chart.
| Model | Supplier | Standard input | Standard output | Base price applies to |
|---|---|---|---|---|
Grok 4.5 | xAI | $2 / 1M tokens | $6 / 1M tokens | — |
GPT-5.6 Sol | OpenAI | $4 / 1M tokens | $20 / 1M tokens | — |
Claude Fable 5 | Anthropic | $10 / 1M tokens | $50 / 1M tokens | — |
GLM-5.2 | 智谱AI | $1.4 / 1M tokens | $4.4 / 1M tokens | — |
Claude Sonnet 5 | Anthropic | $2 / 1M tokens | $10 / 1M tokens | — |
GPT-5.6 Terra | OpenAI | $2 / 1M tokens | $12 / 1M tokens | — |
Kimi K3 | Moonshot AI | ¥20 / 1M tokens | ¥100 / 1M tokens | — |
Version History
How each version of the Grok 4.5 series stacks up on benchmark tests
3 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.
| Benchmark | Grok 4.5Current | Grok 4.3 Beta |
|---|---|---|
GPQA Diamond 科学与综合推理 | 93.43Thinking Level · High | 88.83Thinking Level · High |
24.39Thinking Level · High | 14.63Thinking Level · High | |
FrontierMath v2 数学推理 | 57.19Thinking Level · High | 42.81Thinking Level · High |
Single-Benchmark Version Trend
Viewing: GPQA Diamond · 科学与综合推理
Standard API Pricing Across the Grok 4.5 Series
Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.
Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens
When a context threshold exists, the charted base price only applies within these limits:
| Model | Supplier | Standard input | Standard output | Base price applies to |
|---|---|---|---|---|
Grok 4.5 | xAI | $2 / 1M tokens | $6 / 1M tokens | — |
Grok 4.20 | xAI | $1.25 / 1M tokens | $2.5 / 1M tokens | <= 200000 |