Step 3.7 FlashvsQwen3.6-35B-A3B
Across 3 shared benchmarks, Step 3.7 Flash leads overall: Step 3.7 Flash wins 2, Qwen3.6-35B-A3B wins 1, with 0 ties and an average score difference of +5.68.
Step 3.7 Flash
StepFunAI · 2026-05-29 · Reasoning model
Qwen3.6-35B-A3B
阿里巴巴 · 2026-04-16 · Reasoning model
Step 3.7 Flash2 wins(67%)(33%)1 winQwen3.6-35B-A3B
Benchmark scores
Grouped by capability, sorted by largest gap within each. 3 shared benchmarks.
Coding and Software Engineer
Step 3.7 Flash 1/1| Benchmark | Step 3.7 Flash | Qwen3.6-35B-A3B | Diff |
|---|---|---|---|
| SWE-Bench Pro - Public | 56.3027 / 59Thinking (With Tools) | 49.5048 / 59Thinking (No Tools) | +6.80 |
General Knowledge
Step 3.7 Flash 1/1| Benchmark | Step 3.7 Flash | Qwen3.6-35B-A3B | Diff |
|---|---|---|---|
| HLE | 47.2041 / 185Thinking (With Tools) | 21.40127 / 185Thinking (No Tools) | +25.80 |
Text Embedding
Qwen3.6-35B-A3B 1/1| Benchmark | Step 3.7 Flash | Qwen3.6-35B-A3B | Diff |
|---|---|---|---|
| Context Arena | 37.65101 / 126Thinking Medium (No Tools) | 53.2084 / 126Normal (No Tools) | -15.55 |
Specs
| Field | Step 3.7 Flash | Qwen3.6-35B-A3B |
|---|---|---|
| Publisher | StepFunAI | 阿里巴巴 |
| Release date | 2026-05-29 | 2026-04-16 |
| Model type | Reasoning model | Reasoning model |
| Architecture | MoE | MoE |
| Parameters | 198B | 35B |
| Context length | 256K | 200K |
| Max output | Not available | 80K |
API pricing
Prices use DataLearner records when available; missing fields are not inferred.
| Item | Step 3.7 Flash | Qwen3.6-35B-A3B |
|---|---|---|
| Text input | ¥1.35 / 1M tokens | Not public |
| Text output | ¥8.1 / 1M tokens | Not public |
| Cache read | ¥0.27 / 1M tokens | Not public |
One or both models have incomplete public pricing.
Summary
- Step 3.7 Flashleads in:Coding and Software Engineer (1/1), General Knowledge (1/1)
- Qwen3.6-35B-A3Bleads in:Text Embedding (1/1)
On average across the 3 shared benchmarks, Step 3.7 Flash scores 5.68 higher.
Largest single-benchmark gap: HLE — Step 3.7 Flash 47.20 vs Qwen3.6-35B-A3B 21.40 (+25.80).
Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.