DataLearner logo

Step 3.7 FlashvsGLM-5V-Turbo

Across 6 shared benchmarks, Step 3.7 Flash leads overall: Step 3.7 Flash wins 4, GLM-5V-Turbo wins 0, with 2 ties and an average score difference of +2.23.

StepFunAI
Step 3.7 Flash

StepFunAI · 2026-05-29 · Reasoning model

智谱AI
GLM-5V-Turbo

智谱AI · 2026-04-01 · Reasoning model

Step 3.7 Flash4 wins(67%)Ties2(0%)0 winsGLM-5V-Turbo

Benchmark scores

Grouped by capability, sorted by largest gap within each. 6 shared benchmarks.

Agent Level Benchmark

Step 3.7 Flash 1/2
BenchmarkStep 3.7 FlashGLM-5V-TurboDiff
Terminal Bench Hard35.6070 / 244Thinking (With Tools)32.6093 / 244Thinking (With Tools)+3
τ²-Bench - Telecom98.504 / 264Thinking (With Tools)98.504 / 264Thinking (With Tools)

General Evaluation

Even 1/1
BenchmarkStep 3.7 FlashGLM-5V-TurboDiff
GPQA Diamond80.90225 / 462Thinking (No Tools)80.90225 / 462Thinking (No Tools)

General Knowledge

Step 3.7 Flash 1/1
BenchmarkStep 3.7 FlashGLM-5V-TurboDiff
CritPt2.30126 / 200Thinking (No Tools)0.60168 / 200Thinking (No Tools)+1.70

Instruction Following

Step 3.7 Flash 1/1
BenchmarkStep 3.7 FlashGLM-5V-TurboDiff
IF Bench67.3090 / 282Thinking (No Tools)61.10116 / 282Thinking (No Tools)+6.20

Multimodal Understanding

Step 3.7 Flash 1/1
BenchmarkStep 3.7 FlashGLM-5V-TurboDiff
MMMU-Pro75.3086 / 227Thinking (No Tools)72.80111 / 227Thinking (No Tools)+2.50

Specs

FieldStep 3.7 FlashGLM-5V-Turbo
PublisherStepFunAI智谱AI
Release date2026-05-292026-04-01
Model typeReasoning modelReasoning model
ArchitectureMoEDense
Parameters198BNot available
Context length256K200K
Max outputNot available128K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemStep 3.7 FlashGLM-5V-Turbo
Text input¥1.35 / 1M tokens$1.2 / 1M tokens
Text output¥8.1 / 1M tokens$4 / 1M tokens
Cache read¥0.27 / 1M tokens$0.24 / 1M tokens

Summary

  • Step 3.7 Flashleads in:Agent Level Benchmark (1/2), General Knowledge (1/1), Instruction Following (1/1), Multimodal Understanding (1/1)
  • Tied in:General Evaluation

On average across the 6 shared benchmarks, Step 3.7 Flash scores 2.23 higher.

Largest single-benchmark gap: IF Bench — Step 3.7 Flash 67.30 vs GLM-5V-Turbo 61.10 (+6.20).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.