DataLearner logo

Grok 4.5 Benchmark Details

Grok 4.5 currently shows benchmark results led by GPQA Diamond (18 / 271, score 93.43), SWE-Bench Pro - Public (7 / 60, score 64.70), SimpleBench (15 / 92, score 70). This page also compares it with 6 competitor models and 3 predecessor or same-series models, including performance and pricing views when available.

Benchmark Results

Grok 4.5

Benchmark Results

Thinking
Tool usage

Other

1 evaluations
Benchmark / mode
Score
Rank/total
93.43
18 / 271

Writing and Creative Capabilities

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
1576
38 / 99

Common Sense Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
70
15 / 92

Coding and Software Engineer

6 evaluations
Benchmark / mode
Score
Rank/total
66.70
5 / 5
64.70
7 / 60
56.60
5 / 6
APEX-SWE
HighTools
53.60
3 / 3
DeepSWE
HighTools
53
26 / 35
SWE-Marathon
HighTools
29
5 / 6

AI Agent - Tool Usage

3 evaluations
Benchmark / mode
Score
Rank/total
83.30
18 / 49
15.70
6 / 7
12.42
12 / 13

General Knowledge

1 evaluations
Benchmark / mode
Score
Rank/total
56
16 / 28

Productivity Knowledge

2 evaluations
Benchmark / mode
Score
Rank/total
GDPval-AA v2
HighTools
1526
15 / 25
Harvey Lab-AA
HighTools
12.90
4 / 6

Agent Level Benchmark

1 evaluations
Benchmark / mode
Score
Rank/total
APEX-Agents
HighTools
47.10
4 / 6

Math and Reasoning

2 evaluations
Benchmark / mode
Score
Rank/total
57.19
24 / 58
24.39
24 / 41

Competitor Comparison

Benchmark scores for Grok 4.5 compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkGrok 4.5CurrentGPT-5.6 SolClaude Fable 5GLM-5.2Claude Sonnet 5GPT-5.6 TerraKimi K3
GPQA Diamond
科学与综合推理
93.43Thinking Level · High
93.50Thinking Level · High
85.86Thinking Level · High
91.86Thinking Level · High
90.53Thinking Level · Extra High
93.31Thinking Level · High
93.50Thinking Level · High
Creative Writing
写作和创作
1576.00Standard Mode
1964.10Standard Mode
1933.20Standard Mode
1750.90Standard Mode
1787.60Standard Mode
1850.00Standard Mode
2070.80Standard Mode
SimpleBench
常识推理
70.00Thinking Level · High
64.80Thinking Level · Extra High
81.90Standard Mode
58.80Standard Mode
60.60Standard Mode
48.90Thinking Level · Extra High
60.70Thinking Level · High
APEX-SWE
编程与软件工程
53.60Thinking Level · High | Tools
--
58.80Thinking Level · High | Tools
--
--
--
--
CursorBench 3.2
编程与软件工程
66.70Thinking Level · High | Tools
67.20Thinking Level · High | Tools
70.50Thinking Level · High | Tools
--
--
--
--
DeepSWE
编程与软件工程
53.00Thinking Level · High | Tools
72.70Thinking Level · Extra High | Tools
70.00Deep Thinking Mode | Tools
44.00Deep Thinking Mode | Tools
54.00Deep Thinking Mode | Tools
69.60Thinking Level · Extra High | Tools
67.50Thinking Level · High | Tools
FrontierCode 1.1
编程与软件工程
56.60Thinking Level · High | Tools
60.60Thinking Level · High | Tools
63.60Thinking Level · High | Tools
--
--
--
--
SWE-Bench Pro - Public
编程与软件工程
64.70Thinking Level · High | Tools
64.60Thinking Level · Extra High | Tools
80.30Deep Thinking Mode | Tools
62.10Thinking Enabled | Tools
--
--
--
SWE-Marathon
编程与软件工程
29.00Thinking Level · High | Tools
--
--
13.00Thinking Level · High | Tools
--
--
42.00Thinking Level · High | Tools
Terminal-Bench 2.1
AI Agent - 工具使用
83.30Thinking Level · High | Tools
88.80Thinking Level · High
88.00Thinking Level · High | Tools
81.00Thinking Level · High | Tools
80.40Thinking Level · Extra High | Tools
87.40Thinking Level · High
88.30Thinking Level · High | Tools
Terminal-Bench 3.0
AI Agent - 工具使用
15.70Thinking Level · High | Tools
34.60Thinking Level · High | Tools
34.10Thinking Level · High | Tools
--
--
--
--
Terminal-Bench 4.0
AI Agent - 工具使用
12.42Thinking Level · High | Tools
37.27Thinking Level · High | Tools
44.55Thinking Level · High | Tools
--
12.42Thinking Level · High | Tools
21.52Thinking Level · High | Tools
--
6 additional benchmarks remain in the chart above.

Standard API Pricing: Grok 4.5 vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier.

These models use different currencies or billing units, so the page falls back to raw price values instead of a shared bar chart.

Grok 4.5
Supplier: xAI
Standard input: $2 / 1M tokens
Standard output: $6 / 1M tokens
GPT-5.6 Sol
Supplier: OpenAI
Standard input: $4 / 1M tokens
Standard output: $20 / 1M tokens
Claude Fable 5
Supplier: Anthropic
Standard input: $10 / 1M tokens
Standard output: $50 / 1M tokens
GLM-5.2
Supplier: 智谱AI
Standard input: $1.4 / 1M tokens
Standard output: $4.4 / 1M tokens
Claude Sonnet 5
Supplier: Anthropic
Standard input: $2 / 1M tokens
Standard output: $10 / 1M tokens
GPT-5.6 Terra
Supplier: OpenAI
Standard input: $2 / 1M tokens
Standard output: $12 / 1M tokens
Kimi K3
Supplier: Moonshot AI
Standard input: ¥20 / 1M tokens
Standard output: ¥100 / 1M tokens
ModelSupplierStandard inputStandard outputBase price applies to
Grok 4.5
xAI$2 / 1M tokens$6 / 1M tokens
GPT-5.6 Sol
OpenAI$4 / 1M tokens$20 / 1M tokens
Claude Fable 5
Anthropic$10 / 1M tokens$50 / 1M tokens
GLM-5.2
智谱AI$1.4 / 1M tokens$4.4 / 1M tokens
Claude Sonnet 5
Anthropic$2 / 1M tokens$10 / 1M tokens
GPT-5.6 Terra
OpenAI$2 / 1M tokens$12 / 1M tokens
Kimi K3
Moonshot AI¥20 / 1M tokens¥100 / 1M tokens

Version History

How each version of the Grok 4.5 series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

3 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkGrok 4.5CurrentGrok 4.3 Beta
GPQA Diamond
科学与综合推理
93.43Thinking Level · High
88.83Thinking Level · High
24.39Thinking Level · High
14.63Thinking Level · High
FrontierMath v2
数学推理
57.19Thinking Level · High
42.81Thinking Level · High

Single-Benchmark Version Trend

Viewing: GPQA Diamond · 科学与综合推理

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the Grok 4.5 Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

When a context threshold exists, the charted base price only applies within these limits:

Grok 4.20: Base price applies to <= 200000
ModelSupplierStandard inputStandard outputBase price applies to
Grok 4.5
xAI$2 / 1M tokens$6 / 1M tokens
Grok 4.20
xAI$1.25 / 1M tokens$2.5 / 1M tokens<= 200000