DataLearner logo

Step 3.7 Flash Benchmark Details

Step 3.7 Flash currently shows benchmark results led by τ²-Bench - Telecom (4 / 264, score 98.50), HLE (69 / 563, score 47.20), Terminal Bench Hard (70 / 244, score 35.60). This page also compares it with 3 competitor models and 2 predecessor or same-series models, including performance and pricing views when available.

Benchmark Results

Step 3.7 Flash

Benchmark Results

Thinking
Tool usage

General Knowledge

3 evaluations
Benchmark / mode
Score
Rank/total
HLE
Thinking Enabled
21.40
277 / 563
HLE
Thinking EnabledTools
47.20
69 / 563
CritPt
Thinking Enabled
2.30
126 / 200

Other

1 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Thinking Enabled
80.90
225 / 462

Multimodal Understanding

2 evaluations
Benchmark / mode
Score
Rank/total
SimpleVQA
Thinking EnabledTools
79.20
1 / 3
MMMU-Pro
Thinking Enabled
75.30
86 / 227

Coding and Software Engineer

2 evaluations
Benchmark / mode
Score
Rank/total
SWE-Bench Pro - Public
Thinking EnabledTools
56.30
28 / 62
SciCode
Thinking Enabled
43.90
93 / 130

Agent Level Benchmark

3 evaluations
Benchmark / mode
Score
Rank/total
τ²-Bench - Telecom
Thinking EnabledTools
98.50
4 / 264
Terminal Bench Hard
Thinking EnabledTools
35.60
70 / 244
τ³-Banking
Thinking EnabledTools
12
123 / 164

Instruction Following

1 evaluations
Benchmark / mode
Score
Rank/total
IF Bench
Thinking Enabled
67.30
90 / 282

AI Agent - Information Search

1 evaluations
Benchmark / mode
Score
Rank/total
BrowseComp
Thinking EnabledTools
75.82
27 / 57

Text Embedding

3 evaluations
Benchmark / mode
Score
Rank/total
Context Arena
Thinking Level · Low
33.97
108 / 126
Context Arena
Thinking Level · Medium
37.65
101 / 126
Context Arena
Thinking Level · High
35.33
104 / 126

AI Agent - Tool Usage

1 evaluations
Benchmark / mode
Score
Rank/total
Terminal-Bench 2.1
Thinking EnabledTools
59.50
112 / 192

Productivity Knowledge

1 evaluations
Benchmark / mode
Score
Rank/total
Harvey Lab-AA
Thinking EnabledTools
72.73
32 / 43

Competitor Comparison

Benchmark scores for Step 3.7 Flash compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkStep 3.7 FlashCurrentGLM-5.2MiniMax M3Qwen3.6-35B-A3B
CritPt
Score
综合评估
2.30Thinking Enabled
20.90Thinking Level · High
3.70Thinking Enabled
0.30Thinking Enabled
HLE
Accuracy
综合评估
47.20Thinking Enabled | Tools
54.70Thinking Enabled | Tools
39.00Thinking Enabled
22.20Thinking Enabled
GPQA Diamond
Accuracy
科学与综合推理
80.90Thinking Enabled
91.86Thinking Level · High
92.90Thinking Enabled
84.85Standard Mode
MMMU-Pro
Accuracy
多模态理解
75.30Thinking Enabled
--
78.60Thinking Enabled
75.00Thinking Enabled
SciCode
Score
编程与软件工程
43.90Thinking Enabled
51.20Thinking Level · High
45.37Thinking Enabled
36.60Thinking Enabled
SWE-Bench Pro - Public
Accuracy
编程与软件工程
56.30Thinking Enabled | Tools
62.10Thinking Enabled | Tools
59.00Thinking Enabled | Tools
49.50Thinking Enabled
Terminal Bench Hard
Accuracy
Agent能力评测
35.60Thinking Enabled | Tools
50.80Thinking Level · High | Tools
42.40Thinking Enabled | Tools
34.80Thinking Enabled | Tools
τ²-Bench - Telecom
Accuracy
Agent能力评测
98.50Thinking Enabled | Tools
99.10Thinking Level · High | Tools
88.90Thinking Enabled | Tools
95.30Thinking Enabled | Tools
τ³-Banking
Score
Agent能力评测
12.00Thinking Enabled | Tools
37.11Thinking Level · Extra High | Tools
15.30Thinking Enabled | Tools
9.30Thinking Enabled | Tools
IF Bench
Accuracy
指令跟随
67.30Thinking Enabled
73.30Thinking Level · High
82.90Thinking Enabled
64.40Thinking Enabled
BrowseComp
Accuracy
AI Agent - 信息收集
75.82Thinking Enabled | Tools
--
83.50Thinking Enabled | Tools
--
Context Arena
Accuracy (8 needles, 4K-128K context)
文本向量检索
37.65Thinking Level · Medium
72.34Thinking Level · High
51.15Thinking Enabled
83.53Thinking Enabled
2 additional benchmarks remain in the chart above.

Standard API Pricing: Step 3.7 Flash vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier.

These models use different currencies or billing units, so the page falls back to raw price values instead of a shared bar chart.

Step 3.7 Flash
Supplier: StepFunAI
Standard input: ¥1.35 / 1M tokens
Standard output: ¥8.1 / 1M tokens
GLM-5.2
Supplier: 智谱AI
Standard input: $1.4 / 1M tokens
Standard output: $4.4 / 1M tokens
MiniMax M3
Supplier: MiniMaxAI
Standard input: ¥2.1 / 1M tokens
Standard output: ¥8.4 / 1M tokens
ModelSupplierStandard inputStandard outputBase price applies to
Step 3.7 Flash
StepFunAI¥1.35 / 1M tokens¥8.1 / 1M tokens
GLM-5.2
智谱AI$1.4 / 1M tokens$4.4 / 1M tokens
MiniMax M3
MiniMaxAI¥2.1 / 1M tokens¥8.4 / 1M tokens

Version History

How each version of the Step 3.7 Flash series stacks up on benchmark tests

Step 3.7 FlashStep 3.5 FlashStep3
Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

8 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkStep 3.7 FlashCurrentStep 3.5 FlashStep3
CritPt
Score
综合评估
2.30Thinking Enabled
2.50Thinking Enabled
--
HLE
Accuracy
综合评估
47.20Thinking Enabled | Tools
24.50Thinking Enabled
--
GPQA Diamond
Accuracy
科学与综合推理
80.90Thinking Enabled
83.10Thinking Enabled
73.00Standard Mode
SimpleVQA
Accuracy
多模态理解
79.20Thinking Enabled | Tools
--
62.20Standard Mode
Terminal Bench Hard
Accuracy
Agent能力评测
35.60Thinking Enabled | Tools
32.60Thinking Enabled | Tools
--
τ²-Bench - Telecom
Accuracy
Agent能力评测
98.50Thinking Enabled | Tools
94.40Thinking Enabled | Tools
--
IF Bench
Accuracy
指令跟随
67.30Thinking Enabled
66.50Thinking Enabled
--
BrowseComp
Accuracy
AI Agent - 信息收集
75.82Thinking Enabled | Tools
69.00Thinking Enabled | Tools
--

Single-Benchmark Version Trend

Viewing: CritPt · 综合评估

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the Step 3.7 Flash Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · CNY / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
Step 3.7 Flash
StepFunAI¥1.35 / 1M tokens¥8.1 / 1M tokens