DataLearner logo

Grok 4.5 Benchmark Details

Grok 4.5 currently shows benchmark results led by GPQA Diamond (15 / 225, score 93.43), SWE-Bench Pro - Public (6 / 58, score 64.70), Terminal-Bench 2.1 (15 / 46, score 83.30). This page also compares it with 4 competitor models and 3 predecessor or same-series models, including performance and pricing views when available.

Benchmark Results

Grok 4.5

Benchmark Results

Thinking
Tool usage

Other

1 evaluations
Benchmark / mode
Score
Rank/total
93.43
15 / 225

Writing and Creative Capabilities

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
1576
38 / 99

Coding and Software Engineer

6 evaluations
Benchmark / mode
Score
Rank/total
66.70
4 / 4
64.70
6 / 58
56.60
4 / 5
APEX-SWE
HighTools
53.60
3 / 3
DeepSWE
HighTools
53
21 / 30
SWE-Marathon
HighTools
29
3 / 4

AI Agent - Tool Usage

2 evaluations
Benchmark / mode
Score
Rank/total
83.30
15 / 46
15.70
5 / 6

General Knowledge

1 evaluations
Benchmark / mode
Score
Rank/total
56
4 / 8

Productivity Knowledge

3 evaluations
Benchmark / mode
Score
Rank/total
GDPval-AA v2
HighTools
1526
9 / 15
AA-Briefcase
HighTools
1313
6 / 6
Harvey Lab-AA
HighTools
12.90
4 / 6

Agent Level Benchmark

1 evaluations
Benchmark / mode
Score
Rank/total
APEX-Agents
HighTools
47.10
4 / 5

Math and Reasoning

2 evaluations
Benchmark / mode
Score
Rank/total
57.19
21 / 34
24.39
20 / 34

Competitor Comparison

Benchmark scores for Grok 4.5 compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkGrok 4.5CurrentClaude Sonnet 5GPT-5.6 TerraKimi K3GLM-5.2
GPQA Diamond
科学与综合推理
93.43Thinking Level · High
90.53Thinking Level · Extra High
93.31Thinking Level · High
93.50Thinking Level · High
91.86Thinking Level · High
Creative Writing
写作和创作
1576.00Standard Mode
1787.60Standard Mode
1850.00Standard Mode
2070.80Standard Mode
1750.90Standard Mode
DeepSWE
编程与软件工程
53.00Thinking Level · High | Tools
54.00Deep Thinking Mode | Tools
69.60Thinking Level · Extra High | Tools
67.50Thinking Level · High | Tools
44.00Deep Thinking Mode | Tools
SWE-Bench Pro - Public
编程与软件工程
64.70Thinking Level · High | Tools
--
--
--
62.10Thinking Enabled | Tools
SWE-Marathon
编程与软件工程
29.00Thinking Level · High | Tools
--
--
42.00Thinking Level · High | Tools
13.00Thinking Level · High | Tools
Terminal-Bench 2.1
AI Agent - 工具使用
83.30Thinking Level · High | Tools
80.40Thinking Level · Extra High | Tools
87.40Thinking Level · High
88.30Thinking Level · High | Tools
81.00Thinking Level · High | Tools
56.00Thinking Level · High | Tools
--
55.00Thinking Level · High
--
--
AA-Briefcase
生产力知识
1313.00Thinking Level · High | Tools
--
--
1548.00Thinking Level · High | Tools
--
GDPval-AA v2
生产力知识
1526.00Thinking Level · High | Tools
--
--
1686.00Thinking Level · High | Tools
--
Harvey Lab-AA
生产力知识
12.90Thinking Level · High | Tools
--
--
94.60Thinking Level · High | Tools
--
APEX-Agents
Agent能力评测
47.10Thinking Level · High | Tools
--
--
41.00Thinking Level · High | Tools
--
24.39Thinking Level · High
29.27Thinking Level · High
70.73Thinking Level · High
39.02Thinking Level · High
29.27Thinking Level · High
1 additional benchmarks remain in the chart above.

Standard API Pricing: Grok 4.5 vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier.

These models use different currencies or billing units, so the page falls back to raw price values instead of a shared bar chart.

Grok 4.5
Supplier: xAI
Standard input: $2 / 1M tokens
Standard output: $6 / 1M tokens
Claude Sonnet 5
Supplier: Anthropic
Standard input: $2 / 1M tokens
Standard output: $10 / 1M tokens
GPT-5.6 Terra
Supplier: OpenAI
Standard input: $2.5 / 1M tokens
Standard output: $15 / 1M tokens
Kimi K3
Supplier: Moonshot AI
Standard input: ¥20 / 1M tokens
Standard output: ¥100 / 1M tokens
GLM-5.2
Supplier: 智谱AI
Standard input: $1.4 / 1M tokens
Standard output: $4.4 / 1M tokens
ModelSupplierStandard inputStandard outputBase price applies to
Grok 4.5
xAI$2 / 1M tokens$6 / 1M tokens
Claude Sonnet 5
Anthropic$2 / 1M tokens$10 / 1M tokens
GPT-5.6 Terra
OpenAI$2.5 / 1M tokens$15 / 1M tokens
Kimi K3
Moonshot AI¥20 / 1M tokens¥100 / 1M tokens
GLM-5.2
智谱AI$1.4 / 1M tokens$4.4 / 1M tokens

Version History

How each version of the Grok 4.5 series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

2 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkGrok 4.5CurrentGrok 4.3 Beta
24.39Thinking Level · High
14.63Thinking Level · High
FrontierMath v2
数学推理
57.19Thinking Level · High
42.81Thinking Level · High

Single-Benchmark Version Trend

Viewing: FrontierMath Tier 4 v2 · 数学推理

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the Grok 4.5 Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

Grok 4.20: Base price applies to <= 200000
ModelSupplierStandard inputStandard outputBase price applies to
Grok 4.5
xAI$2 / 1M tokens$6 / 1M tokens
Grok 4.20
xAI$1.25 / 1M tokens$2.5 / 1M tokens<= 200000