DataLearner logo

GLM-5.2vsStep 3.7 Flash

Across 3 shared benchmarks, GLM-5.2 leads overall: GLM-5.2 wins 3, Step 3.7 Flash wins 0, with 0 ties and an average score difference of +11.60.

智谱AI
GLM-5.2

智谱AI · 2026-06-13 · Reasoning model

StepFunAI
Step 3.7 Flash

StepFunAI · 2026-05-29 · Reasoning model

GLM-5.23 wins(100%)(0%)0 winsStep 3.7 Flash

Benchmark scores

Grouped by capability, sorted by largest gap within each. 3 shared benchmarks.

AI Agent - Tool Usage

GLM-5.2 1/1
BenchmarkGLM-5.2Step 3.7 FlashDiff
Terminal-Bench 2.18117 / 44Thinking High (With Tools)59.5037 / 44Thinking (With Tools)+21.50

Coding and Software Engineer

GLM-5.2 1/1
BenchmarkGLM-5.2Step 3.7 FlashDiff
SWE-Bench Pro - Public62.109 / 57Thinking (With Tools)56.3025 / 57Thinking (With Tools)+5.80

General Knowledge

GLM-5.2 1/1
BenchmarkGLM-5.2Step 3.7 FlashDiff
HLE54.7015 / 181Thinking (With Tools)47.2039 / 181Thinking (With Tools)+7.50

Specs

FieldGLM-5.2Step 3.7 Flash
Publisher智谱AIStepFunAI
Release date2026-06-132026-05-29
Model typeReasoning modelReasoning model
ArchitectureMoEMoE
Parameters753.33B198B
Context length1M256K
Max output128KNot available

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemGLM-5.2Step 3.7 Flash
Text input$1.4 / 1M tokens¥1.35 / 1M tokens
Text output$4.4 / 1M tokens¥8.1 / 1M tokens
Cache read$0.26 / 1M tokens¥0.27 / 1M tokens

Summary

  • GLM-5.2leads in:AI Agent - Tool Usage (1/1), Coding and Software Engineer (1/1), General Knowledge (1/1)

On average across the 3 shared benchmarks, GLM-5.2 scores 11.60 higher.

Largest single-benchmark gap: Terminal-Bench 2.1 — GLM-5.2 81 vs Step 3.7 Flash 59.50 (+21.50).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.