DataLearner logo

Step 3.7 FlashvsStep 3.5 Flash

Across 6 shared benchmarks, Step 3.7 Flash leads overall: Step 3.7 Flash wins 4, Step 3.5 Flash wins 2, with 0 ties and an average score difference of +2.05.

StepFunAI
Step 3.7 Flash

StepFunAI · 2026-05-29 · Reasoning model

StepFunAI
Step 3.5 Flash

StepFunAI · 2026-02-02 · Chat model

Step 3.7 Flash4 wins(67%)(33%)2 winsStep 3.5 Flash

Benchmark scores

Grouped by capability, sorted by largest gap within each. 6 shared benchmarks.

Agent Level Benchmark

Step 3.7 Flash 2/2
BenchmarkStep 3.7 FlashStep 3.5 FlashDiff
τ²-Bench - Telecom98.504 / 264Thinking (With Tools)94.4035 / 264Thinking (With Tools)+4.10
Terminal Bench Hard35.6070 / 244Thinking (With Tools)32.6093 / 244Thinking (With Tools)+3

AI Agent - Information Search

Step 3.7 Flash 1/1
BenchmarkStep 3.7 FlashStep 3.5 FlashDiff
BrowseComp75.8227 / 57Thinking (With Tools)6931 / 57Thinking (With Tools)+6.82

General Evaluation

Step 3.5 Flash 1/1
BenchmarkStep 3.7 FlashStep 3.5 FlashDiff
GPQA Diamond80.90225 / 462Thinking (No Tools)83.10191 / 462Thinking (No Tools)-2.20

General Knowledge

Step 3.5 Flash 1/1
BenchmarkStep 3.7 FlashStep 3.5 FlashDiff
CritPt2.30126 / 200Thinking (No Tools)2.50125 / 200Thinking (No Tools)-0.20

Instruction Following

Step 3.7 Flash 1/1
BenchmarkStep 3.7 FlashStep 3.5 FlashDiff
IF Bench67.3090 / 282Thinking (No Tools)66.5093 / 282Thinking (No Tools)+0.80

Specs

FieldStep 3.7 FlashStep 3.5 Flash
PublisherStepFunAIStepFunAI
Release date2026-05-292026-02-02
Model typeReasoning modelChat model
ArchitectureMoEMoE
Parameters198B196B
Context length256K256K
Max outputNot available16K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemStep 3.7 FlashStep 3.5 Flash
Text input¥1.35 / 1M tokensNot public
Text output¥8.1 / 1M tokensNot public
Cache read¥0.27 / 1M tokensNot public

One or both models have incomplete public pricing.

Summary

  • Step 3.7 Flashleads in:Agent Level Benchmark (2/2), AI Agent - Information Search (1/1), Instruction Following (1/1)
  • Step 3.5 Flashleads in:General Evaluation (1/1), General Knowledge (1/1)

On average across the 6 shared benchmarks, Step 3.7 Flash scores 2.05 higher.

Largest single-benchmark gap: BrowseComp — Step 3.7 Flash 75.82 vs Step 3.5 Flash 69 (+6.82).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.