DataLearner logo

Step 3.7 Flash Benchmark Details

Step 3.7 Flash currently shows benchmark results led by HLE (46 / 197, score 47.20), SimpleVQA (1 / 3, score 79.20), SWE-Bench Pro - Public (28 / 62, score 56.30). This page also compares it with 3 competitor models and 2 predecessor or same-series models, including performance and pricing views when available.

Benchmark Results

Step 3.7 Flash

Benchmark Results

Thinking
Tool usage

General Knowledge

1 evaluations
Benchmark / mode
Score
Rank/total
HLE
Thinking EnabledTools
47.20
46 / 197

Multimodal Understanding

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleVQA
Thinking EnabledTools
79.20
1 / 3

Coding and Software Engineer

1 evaluations
Benchmark / mode
Score
Rank/total
SWE-Bench Pro - Public
Thinking EnabledTools
56.30
28 / 62

AI Agent - Information Search

1 evaluations
Benchmark / mode
Score
Rank/total
BrowseComp
Thinking EnabledTools
75.82
27 / 57

Text Embedding

3 evaluations
Benchmark / mode
Score
Rank/total
Context Arena
Thinking Level · Low
33.97
108 / 126
Context Arena
Thinking Level · Medium
37.65
101 / 126
Context Arena
Thinking Level · High
35.33
104 / 126

AI Agent - Tool Usage

1 evaluations
Benchmark / mode
Score
Rank/total
Terminal-Bench 2.1
Thinking EnabledTools
59.50
45 / 53

Competitor Comparison

Benchmark scores for Step 3.7 Flash compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

5 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkStep 3.7 FlashCurrentGLM-5.2MiniMax M3Qwen3.6-35B-A3B
HLE
Accuracy
综合评估
47.20Thinking Enabled | Tools
54.70Thinking Enabled | Tools
--
21.40Thinking Enabled
SWE-Bench Pro - Public
Accuracy
编程与软件工程
56.30Thinking Enabled | Tools
62.10Thinking Enabled | Tools
59.00Thinking Enabled | Tools
49.50Thinking Enabled
BrowseComp
Accuracy
AI Agent - 信息收集
75.82Thinking Enabled | Tools
--
83.50Thinking Enabled | Tools
--
Context Arena
Accuracy (8 needles, 4K-128K context)
文本向量检索
37.65Thinking Level · Medium
72.34Thinking Level · High
51.15Thinking Enabled
83.53Thinking Enabled
Terminal-Bench 2.1
Accuracy
AI Agent - 工具使用
59.50Thinking Enabled | Tools
81.00Thinking Level · High | Tools
66.00Thinking Enabled | Tools
--

Standard API Pricing: Step 3.7 Flash vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier.

These models use different currencies or billing units, so the page falls back to raw price values instead of a shared bar chart.

Step 3.7 Flash
Supplier: StepFunAI
Standard input: ¥1.35 / 1M tokens
Standard output: ¥8.1 / 1M tokens
GLM-5.2
Supplier: 智谱AI
Standard input: $1.4 / 1M tokens
Standard output: $4.4 / 1M tokens
MiniMax M3
Supplier: MiniMaxAI
Standard input: ¥2.1 / 1M tokens
Standard output: ¥8.4 / 1M tokens
ModelSupplierStandard inputStandard outputBase price applies to
Step 3.7 Flash
StepFunAI¥1.35 / 1M tokens¥8.1 / 1M tokens
GLM-5.2
智谱AI$1.4 / 1M tokens$4.4 / 1M tokens
MiniMax M3
MiniMaxAI¥2.1 / 1M tokens¥8.4 / 1M tokens

Version History

How each version of the Step 3.7 Flash series stacks up on benchmark tests

Step 3.7 FlashStep 3.5 FlashStep3
Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

2 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkStep 3.7 FlashCurrentStep 3.5 FlashStep3
SimpleVQA
Accuracy
多模态理解
79.20Thinking Enabled | Tools
--
62.20Standard Mode
BrowseComp
Accuracy
AI Agent - 信息收集
75.82Thinking Enabled | Tools
69.00Thinking Enabled | Tools
--

Single-Benchmark Version Trend

Viewing: SimpleVQA · 多模态理解

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the Step 3.7 Flash Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · CNY / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
Step 3.7 Flash
StepFunAI¥1.35 / 1M tokens¥8.1 / 1M tokens