Grok 4.5 Benchmark Details
Grok 4.5 currently shows benchmark results led by GPQA Diamond (15 / 225, score 93.43), SWE-Bench Pro - Public (6 / 58, score 64.70), Terminal-Bench 2.1 (15 / 46, score 83.30). This page also compares it with 4 competitor models and 3 predecessor or same-series models, including performance and pricing views when available.
Benchmark Results
Benchmark Results
Writing and Creative Capabilities
1 evaluationsCoding and Software Engineer
6 evaluationsAI Agent - Tool Usage
2 evaluationsProductivity Knowledge
3 evaluationsMath and Reasoning
2 evaluationsCompetitor Comparison
Benchmark scores for Grok 4.5 compared against top models in its class
12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.
| Benchmark | Grok 4.5Current | Claude Sonnet 5 | GPT-5.6 Terra | Kimi K3 | GLM-5.2 |
|---|---|---|---|---|---|
GPQA Diamond 科学与综合推理 | 93.43Thinking Level · High | 90.53Thinking Level · Extra High | 93.31Thinking Level · High | 93.50Thinking Level · High | 91.86Thinking Level · High |
Creative Writing 写作和创作 | 1576.00Standard Mode | 1787.60Standard Mode | 1850.00Standard Mode | 2070.80Standard Mode | 1750.90Standard Mode |
DeepSWE 编程与软件工程 | 53.00Thinking Level · High | Tools | 54.00Deep Thinking Mode | Tools | 69.60Thinking Level · Extra High | Tools | 67.50Thinking Level · High | Tools | 44.00Deep Thinking Mode | Tools |
SWE-Bench Pro - Public 编程与软件工程 | 64.70Thinking Level · High | Tools | -- | -- | -- | 62.10Thinking Enabled | Tools |
SWE-Marathon 编程与软件工程 | 29.00Thinking Level · High | Tools | -- | -- | 42.00Thinking Level · High | Tools | 13.00Thinking Level · High | Tools |
Terminal-Bench 2.1 AI Agent - 工具使用 | 83.30Thinking Level · High | Tools | 80.40Thinking Level · Extra High | Tools | 87.40Thinking Level · High | 88.30Thinking Level · High | Tools | 81.00Thinking Level · High | Tools |
56.00Thinking Level · High | Tools | -- | 55.00Thinking Level · High | -- | -- | |
AA-Briefcase 生产力知识 | 1313.00Thinking Level · High | Tools | -- | -- | 1548.00Thinking Level · High | Tools | -- |
GDPval-AA v2 生产力知识 | 1526.00Thinking Level · High | Tools | -- | -- | 1686.00Thinking Level · High | Tools | -- |
Harvey Lab-AA 生产力知识 | 12.90Thinking Level · High | Tools | -- | -- | 94.60Thinking Level · High | Tools | -- |
APEX-Agents Agent能力评测 | 47.10Thinking Level · High | Tools | -- | -- | 41.00Thinking Level · High | Tools | -- |
24.39Thinking Level · High | 29.27Thinking Level · High | 70.73Thinking Level · High | 39.02Thinking Level · High | 29.27Thinking Level · High |
Standard API Pricing: Grok 4.5 vs. Peer Models
Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.
Source: DataLearnerAI. Standard text prices shown here use the default supplier.
These models use different currencies or billing units, so the page falls back to raw price values instead of a shared bar chart.
| Model | Supplier | Standard input | Standard output | Base price applies to |
|---|---|---|---|---|
Grok 4.5 | xAI | $2 / 1M tokens | $6 / 1M tokens | — |
Claude Sonnet 5 | Anthropic | $2 / 1M tokens | $10 / 1M tokens | — |
GPT-5.6 Terra | OpenAI | $2.5 / 1M tokens | $15 / 1M tokens | — |
Kimi K3 | Moonshot AI | ¥20 / 1M tokens | ¥100 / 1M tokens | — |
GLM-5.2 | 智谱AI | $1.4 / 1M tokens | $4.4 / 1M tokens | — |
Version History
How each version of the Grok 4.5 series stacks up on benchmark tests
2 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.
| Benchmark | Grok 4.5Current | Grok 4.3 Beta |
|---|---|---|
24.39Thinking Level · High | 14.63Thinking Level · High | |
FrontierMath v2 数学推理 | 57.19Thinking Level · High | 42.81Thinking Level · High |
Single-Benchmark Version Trend
Viewing: FrontierMath Tier 4 v2 · 数学推理
Standard API Pricing Across the Grok 4.5 Series
Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.
Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens
When a context threshold exists, the charted base price only applies within these limits:
| Model | Supplier | Standard input | Standard output | Base price applies to |
|---|---|---|---|---|
Grok 4.5 | xAI | $2 / 1M tokens | $6 / 1M tokens | — |
Grok 4.20 | xAI | $1.25 / 1M tokens | $2.5 / 1M tokens | <= 200000 |