DataLearner logo

GPT-5.4 mini Benchmark Details

GPT-5.4 mini currently shows benchmark results led by GPQA Diamond (73 / 274, score 88), Creative Writing (34 / 106, score 1661.80), HLE (70 / 197, score 41.50). This page also compares it with 2 competitor models and 1 predecessor or same-series models, including performance and pricing views when available.

Benchmark Results

GPT-5.4 mini

Benchmark Results

Thinking
Tool usage

General Knowledge

7 evaluations
Benchmark / mode
Score
Rank/total
LiveBench
Standard Mode
36.95
114 / 117
49.54
95 / 117
LiveBench
Medium
58.33
78 / 117
63.57
56 / 117
LiveBench
Deep Thinking Mode
66.37
50 / 117
HLE
Extra-High
28.20
115 / 197
HLE
Extra-HighTools
41.50
70 / 197

Other

2 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Standard Mode
64.14
221 / 274
GPQA Diamond
Extra-High
88
73 / 274

Writing and Creative Capabilities

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
1661.80
34 / 106

Math and Reasoning

5 evaluations
Benchmark / mode
Score
Rank/total
FrontierMath v2
Standard Mode
17.19
54 / 58
24.56
47 / 58
FrontierMath v2
Extra-High
51.23
29 / 58
9.76
33 / 42
2.10
56 / 80

Coding and Software Engineer

1 evaluations
Benchmark / mode
Score
Rank/total
SWE-Bench Pro - Public
Extra-HighTools
54.40
34 / 62

Agent Level Benchmark

1 evaluations
Benchmark / mode
Score
Rank/total
τ²-Bench - Telecom
Extra-HighTools
93.40
18 / 36

AI Agent - Tool Usage

4 evaluations
Benchmark / mode
Score
Rank/total
OSWorld-Verified
Extra-HighTools
72.10
18 / 26
Terminal Bench 2.0
Extra-HighTools
60
19 / 48
MCP-Atlas
Extra-HighTools
56.70
37 / 41
Tool Decathlon
Extra-HighTools
42.90
5 / 10

Text Embedding

5 evaluations
Benchmark / mode
Score
Rank/total
Context Arena
Standard Mode
23.34
119 / 126
39.26
99 / 126
47.64
91 / 126
49.60
90 / 126
Context Arena
Extra-High
50.92
88 / 126

Claw-style Agent Evaluation

1 evaluations
Benchmark / mode
Score
Rank/total
Claw Bench
Thinking EnabledTools
75.30
25 / 29

Competitor Comparison

Benchmark scores for GPT-5.4 mini compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

9 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkGPT-5.4 miniCurrentHaiku 4.5Gemini 3.0 Flash
HLE
Accuracy
综合评估
41.50Thinking Level · Extra High | Tools
9.70Extended Thinking
43.50Thinking Enabled | Tools
LiveBench
Accuracy
综合评估
66.37Deep Thinking Mode
61.3264K
72.40Thinking Level · High
GPQA Diamond
Accuracy
科学与综合推理
88.00Thinking Level · Extra High
73.30Extended Thinking
90.40Thinking Enabled
FrontierMath - Tier 4
Accuracy
数学推理
2.10Thinking Level · High
2.1032K
4.20Standard Mode
SWE-Bench Pro - Public
Accuracy
编程与软件工程
54.40Thinking Level · Extra High | Tools
39.45Extended Thinking | Tools
49.60Thinking Level · High | Tools
MCP-Atlas
Pass rate / claim coverage
AI Agent - 工具使用
56.70Thinking Level · Extra High | Tools
40.20Standard Mode | Tools
62.00Standard Mode | Tools
Terminal Bench 2.0
Accuracy
AI Agent - 工具使用
60.00Thinking Level · Extra High | Tools
--
47.60Thinking Enabled | Tools
Context Arena
Accuracy (8 needles, 4K-128K context)
文本向量检索
50.92Thinking Level · Extra High
36.27Thinking Enabled
--
Claw Bench
Accuracy
OpenClaw智能体能力综合测评
75.30Thinking Enabled | Tools
89.40Thinking Enabled | Tools
85.70Thinking Enabled | Tools

Standard API Pricing: GPT-5.4 mini vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
GPT-5.4 mini
OpenAI$0.75 / 1M tokens$4.5 / 1M tokens
Haiku 4.5
Anthropic$1 / 1M tokens$5 / 1M tokens
Gemini 3.0 Flash
Google Deep Mind$0.5 / 1M tokens$3 / 1M tokens

Version History

How each version of the GPT-5.4 mini series stacks up on benchmark tests

GPT-5.4 miniGPT-5-mini
Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

7 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkGPT-5.4 miniCurrentGPT-5-mini
HLE
Accuracy
综合评估
41.50Thinking Level · Extra High | Tools
5.00Thinking Enabled
LiveBench
Accuracy
综合评估
66.37Deep Thinking Mode
65.91Thinking Level · High
GPQA Diamond
Accuracy
科学与综合推理
88.00Thinking Level · Extra High
69.00Thinking Enabled
Creative Writing
Elo、大模型评判两两对战
写作和创作
1661.80Standard Mode
1310.30Standard Mode
FrontierMath - Tier 4
Accuracy
数学推理
2.10Thinking Level · High
6.30Thinking Level · High
FrontierMath Tier 4 v2
Accuracy (verification_code)
数学推理
9.76Thinking Level · Extra High
12.20Thinking Level · High
FrontierMath v2
Accuracy (verification_code)
数学推理
51.23Thinking Level · Extra High
46.67Thinking Level · High

Single-Benchmark Version Trend

Viewing: HLE · 综合评估

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the GPT-5.4 mini Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
GPT-5.4 mini
OpenAI$0.75 / 1M tokens$4.5 / 1M tokens
GPT-5-mini
OpenAI$0.25 / 1M tokens$2 / 1M tokens