DataLearner logo

Step 3.5 Flash Benchmark Details

Step 3.5 Flash currently shows benchmark results led by AIME2025 (7 / 110, score 99.80), τ²-Bench (5 / 44, score 88.20), LiveCodeBench (17 / 125, score 86.40).

Benchmark Results

Step 3.5 Flash

Benchmark Results

Thinking
Tool usage

Abstract Generalization

2 evaluations
Benchmark / mode
Score
Rank/total
ARC-AGI-1
Thinking Mode
53.50
130 / 176
ARC-AGI-1
Thinking ModeTools
56.50
124 / 176

Repository Engineering

1 evaluations
Benchmark / mode
Score
Rank/total
SWE-bench Verified
Thinking Mode
74.40
43 / 116

Algorithmic Coding

1 evaluations
Benchmark / mode
Score
Rank/total
LiveCodeBench
Thinking Mode
86.40
17 / 125

Mathematics

4 evaluations
Benchmark / mode
Score
Rank/total
AIME2025
Thinking Mode
97.30
19 / 110
AIME2025
Thinking ModeTools
99.80
7 / 110
IMO-AnswerBench
Thinking Mode
85.40
11 / 24
IMO-AnswerBench
Thinking ModeTools
86.70
9 / 24

Service Workflows

1 evaluations
Benchmark / mode
Score
Rank/total
τ²-Bench
Thinking ModeTools
88.20
5 / 44

Fact Finding

1 evaluations
Benchmark / mode
Score
Rank/total
BrowseComp
Thinking ModeTools
69
32 / 58

Agentic Development

2 evaluations
Benchmark / mode
Score
Rank/total
Terminal Bench 2.0
Thinking ModeTools
51
34 / 48
Terminal Bench Hard
Thinking ModeTools
32.60
93 / 244

Tool Orchestration

3 evaluations
Benchmark / mode
Score
Rank/total
Pinch Bench
Thinking ModeTools
85.30
16 / 38
Claw Bench
Thinking ModeTools
84.90
16 / 29
79.35
12 / 45

Scientific Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
CritPt
Thinking Mode
2.50
129 / 204