DataLearner logo

Step 5 Preview Benchmark Details

Step 5 Preview currently shows benchmark results led by AA-LCR (2 / 171, score 88.30), HLE (9 / 565, score 59.40), GPQA Diamond (25 / 463, score 93.50). This page also compares it with 3 competitor models and 3 predecessor or same-series models, including performance and pricing views when available.

Benchmark Results

Step 5 Preview

Benchmark Results

Thinking
Tool usage
Internet

General Knowledge

3 evaluations
Benchmark / mode
Score
Rank/total
HLE
High
46.50
74 / 565
HLE
HighTools
59.40
9 / 565
CritPt
High
20.90
33 / 201

Other

1 evaluations
Benchmark / mode
Score
Rank/total
93.50
25 / 463

AI Agent - Information Search

1 evaluations
Benchmark / mode
Score
Rank/total
BrowseComp
HighToolsInternet
88.70
7 / 58

Long Context

1 evaluations
Benchmark / mode
Score
Rank/total
AA-LCR
High
88.30
2 / 171

AI Agent - Tool Usage

6 evaluations
Benchmark / mode
Score
Rank/total
MCP-Atlas
HighTools
85.60
5 / 44
85
34 / 194
CyberGym
HighTools
84.70
2 / 9
74.10
5 / 12
51
14 / 14
33.30
18 / 88

Coding and Software Engineer

6 evaluations
Benchmark / mode
Score
Rank/total
Program Bench
HighTools
80.50
2 / 12
SWE-Marathon
HighTools
72.70
1 / 7
DeepSWE
HighTools
67.70
25 / 86
SWE-Atlas-QnA
HighTools
63.60
1 / 6
58.90
9 / 131
MLS Bench
HighTools
40.50
4 / 6

Agent Level Benchmark

4 evaluations
Benchmark / mode
Score
Rank/total
Job Bench
HighTools
59
4 / 6
τ³-Banking
HighTools
42.50
26 / 167
APEX-Agents
HighTools
37.80
6 / 7
29.50
7 / 20

Productivity Knowledge

5 evaluations
Benchmark / mode
Score
Rank/total
GDPval-AA v2
HighTools
1571
21 / 106
AA-Briefcase
HighTools
1417
25 / 84
Office QA Pro
HighTools
60.30
5 / 5
44
9 / 18
29.40
2 / 2

Multimodal Understanding

2 evaluations
Benchmark / mode
Score
Rank/total
76
78 / 229
14.80
59 / 119

Competitor Comparison

Benchmark scores for Step 5 Preview compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkStep 5 PreviewCurrentQwen3.8-27BDeepSeek-V4.1-FlashGLM-5.3-Flash
CritPt
Score
综合评估
20.90Thinking Level · High
5.40Thinking Level · Extra High
14.30Thinking Level · High
15.40Thinking Enabled
HLE
Accuracy
综合评估
59.40Thinking Level · High | Tools
33.90Thinking Level · Extra High
63.90Thinking Level · High | Tools
55.30Thinking Level · High | Tools
GPQA Diamond
Accuracy
科学与综合推理
93.50Thinking Level · High
90.50Thinking Level · Extra High
90.90Thinking Level · High
91.20Thinking Enabled
AA-LCR
Accuracy
长上下文能力
88.30Thinking Level · High
82.00Thinking Level · Extra High
84.00Thinking Level · High
--
AutomationBench-AA
Guardrail-safe objective completion rate (%)
AI Agent - 工具使用
51.00Thinking Level · High | Tools
--
68.90Thinking Level · High
60.40Thinking Level · High
CyberGym
Score
AI Agent - 工具使用
84.70Thinking Level · High | Tools
--
88.10Thinking Level · High | Tools
--
Terminal-Bench 2.1
Accuracy
AI Agent - 工具使用
85.00Thinking Level · High | Tools
79.80Thinking Level · Extra High | Tools
90.60Thinking Level · High | Tools
84.30Thinking Level · High | Tools
Terminal-Bench 4.0
Resolution rate (%)
AI Agent - 工具使用
33.30Thinking Level · High | Tools
5.60Thinking Level · Extra High | Tools
26.80Thinking Level · High | Tools
32.80Thinking Enabled | Tools
Toolathlon-Verified
Score
AI Agent - 工具使用
74.10Thinking Level · High | Tools
--
--
78.40Thinking Level · High | Tools
DeepSWE
Pass@1 (DeepSWE v1.1)
编程与软件工程
67.70Thinking Level · High | Tools
42.20Thinking Enabled | Tools
74.20Thinking Level · High | Tools
63.39Thinking Level · High | Tools
Program Bench
Score
编程与软件工程
80.50Thinking Level · High | Tools
--
20.30Thinking Level · High | Tools
--
SciCode
Score
编程与软件工程
58.90Thinking Level · High
46.60Thinking Level · Extra High
51.90Thinking Level · High
51.60Thinking Enabled
8 additional benchmarks remain in the chart above.

Standard API Pricing: Step 5 Preview vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
Step 5 Preview
StepFunAI$1 / 1M tokens$2.7 / 1M tokens
GLM-5.3-Flash
智谱AI$0.075 / 1M tokens$0.25 / 1M tokens

Version History

How each version of the Step 5 Preview series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

8 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkStep 5 PreviewCurrentStep 3.7 FlashStep 3.5 FlashStep3
CritPt
Score
综合评估
20.90Thinking Level · High
2.30Thinking Enabled
2.50Thinking Enabled
--
HLE
Accuracy
综合评估
59.40Thinking Level · High | Tools
47.20Thinking Enabled | Tools
24.50Thinking Enabled
--
GPQA Diamond
Accuracy
科学与综合推理
93.50Thinking Level · High
80.90Thinking Enabled
83.10Thinking Enabled
73.00Standard Mode
BrowseComp
Accuracy
AI Agent - 信息收集
88.70Thinking Level · High | Tools
75.82Thinking Enabled | Tools
69.00Thinking Enabled | Tools
--
Terminal-Bench 2.1
Accuracy
AI Agent - 工具使用
85.00Thinking Level · High | Tools
59.50Thinking Enabled | Tools
--
--
SciCode
Score
编程与软件工程
58.90Thinking Level · High
43.90Thinking Enabled
--
--
τ³-Banking
Score
Agent能力评测
42.50Thinking Level · High | Tools
12.00Thinking Enabled | Tools
--
--
MMMU-Pro
Accuracy
多模态理解
76.00Thinking Level · High
75.30Thinking Enabled
--
--

Single-Benchmark Version Trend

Viewing: CritPt · 综合评估

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the Step 5 Preview Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier.

These models use different currencies or billing units, so the page falls back to raw price values instead of a shared bar chart.

Step 5 Preview
Supplier: StepFunAI
Standard input: $1 / 1M tokens
Standard output: $2.7 / 1M tokens
Step 3.7 Flash
Supplier: StepFunAI
Standard input: ¥1.35 / 1M tokens
Standard output: ¥8.1 / 1M tokens
ModelSupplierStandard inputStandard outputBase price applies to
Step 5 Preview
StepFunAI$1 / 1M tokens$2.7 / 1M tokens
Step 3.7 Flash
StepFunAI¥1.35 / 1M tokens¥8.1 / 1M tokens