DataLearner logo

GLM-5.2vsGLM 5.1

Across 7 shared benchmarks, GLM-5.2 leads overall: GLM-5.2 wins 7, GLM 5.1 wins 0, with 0 ties and an average score difference of +7.22.

智谱AI
GLM-5.2

智谱AI · 2026-06-13 · Reasoning model

智谱AI
GLM 5.1

智谱AI · 2026-03-27 · Reasoning model

GLM-5.27 wins(100%)(0%)0 winsGLM 5.1

Benchmark scores

Grouped by capability, sorted by largest gap within each. 7 shared benchmarks.

General Knowledge

GLM-5.2 3/3
BenchmarkGLM-5.2GLM 5.1Diff
LiveBench76.249 / 115Normal (No Tools)70.1837 / 115Normal (No Tools)+6.06
GPQA Diamond91.2016 / 187Thinking (No Tools)86.2047 / 187Thinking (No Tools)+5
HLE54.7013 / 172Thinking (With Tools)52.3019 / 172Thinking (With Tools)+2.40

Math and Reasoning

GLM-5.2 2/2
BenchmarkGLM-5.2GLM 5.1Diff
IMO-AnswerBench911 / 21Thinking (No Tools)83.8012 / 21Thinking (No Tools)+7.20
AIME 202699.201 / 18Thinking (No Tools)95.304 / 18Thinking (No Tools)+3.90

AI Agent - Tool Usage

GLM-5.2 1/1
BenchmarkGLM-5.2GLM 5.1Diff
TerminalBench 2.18110 / 28Thinking High (With Tools)58.7025 / 28Thinking High (With Tools)+22.30

Coding and Software Engineer

GLM-5.2 1/1
BenchmarkGLM-5.2GLM 5.1Diff
SWE-Bench Pro - Public62.108 / 54Thinking (With Tools)58.4015 / 54Thinking (With Tools)+3.70

Specs

FieldGLM-5.2GLM 5.1
Publisher智谱AI智谱AI
Release date2026-06-132026-03-27
Model typeReasoning modelReasoning model
ArchitectureMoEMoE
Parameters753.33B75.4B
Context length1M200K
Max output128K125K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemGLM-5.2GLM 5.1
Text input$1.4 / 1M tokens$1.4 / 1M tokens
Text output$4.4 / 1M tokens$4.4 / 1M tokens
Cache read$0.26 / 1M tokens$4.4 / 1M tokens
Cache writeNot public$0.26 / 1M tokens

Summary

  • GLM-5.2leads in:General Knowledge (3/3), Math and Reasoning (2/2), AI Agent - Tool Usage (1/1), Coding and Software Engineer (1/1)

On average across the 7 shared benchmarks, GLM-5.2 scores 7.22 higher.

Largest single-benchmark gap: TerminalBench 2.1 — GLM-5.2 81 vs GLM 5.1 58.70 (+22.30).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.