DataLearner logo

GLM-5vsGLM-4.7

Across 15 shared benchmarks, GLM-5 leads overall: GLM-5 wins 12, GLM-4.7 wins 2, with 1 ties and an average score difference of +17.25.

智谱AI
GLM-5

智谱AI · 2026-02-11 · Chat model

智谱AI
GLM-4.7

智谱AI · 2025-12-22 · Chat model

GLM-512 wins(80%)Ties1(13%)2 winsGLM-4.7

Benchmark scores

Grouped by capability, sorted by largest gap within each. 15 shared benchmarks.

Agent Level Benchmark

GLM-5 3/4
BenchmarkGLM-5GLM-4.7Diff
Terminal Bench Hard39.4053 / 244Normal (With Tools)30.30110 / 244Normal (With Tools)+9.10
τ²-Bench - Telecom97.4016 / 264Normal (With Tools)94.2038 / 264Normal (With Tools)+3.20
τ³-Banking9.79130 / 164Thinking (With Tools)12.20122 / 164Thinking (With Tools)-2.41
τ²-Bench89.704 / 44Thinking (With Tools)87.406 / 44Thinking (With Tools)+2.30

General Knowledge

GLM-5 3/3
BenchmarkGLM-5GLM-4.7Diff
LiveBench68.8543 / 117Normal (No Tools)58.0980 / 117Normal (No Tools)+10.76
HLE7.60427 / 563Normal (No Tools) · Text only6.40453 / 563Normal (No Tools) · Text only+1.20
CritPt2127 / 200Thinking (No Tools)1.70131 / 200Thinking (No Tools)+0.30

Math and Reasoning

GLM-4.7 1/2
BenchmarkGLM-5GLM-4.7Diff
AIME 202692.7018 / 29Thinking (No Tools)92.9017 / 29Thinking (No Tools)-0.20
FrontierMath - Tier 42.1056 / 80Normal (No Tools)2.1056 / 80Normal (No Tools)

AI Agent - Information Search

GLM-5 1/1
BenchmarkGLM-5GLM-4.7Diff
BrowseComp75.9026 / 57Thinking (With Tools)5243 / 57Thinking (With Tools)+23.90

AI Agent - Tool Usage

GLM-5 1/1
BenchmarkGLM-5GLM-4.7Diff
Terminal Bench 2.061.1018 / 48Thinking (With Tools)4145 / 48Thinking (With Tools)+20.10

General Evaluation

GLM-5 1/1
BenchmarkGLM-5GLM-4.7Diff
GPQA Diamond66.60349 / 462Normal (No Tools)66.40350 / 462Normal (No Tools)+0.20

Instruction Following

GLM-5 1/1
BenchmarkGLM-5GLM-4.7Diff
IF Bench55.20132 / 282Normal (No Tools)54.60138 / 282Normal (No Tools)+0.60

Long Context

GLM-5 1/1
BenchmarkGLM-5GLM-4.7Diff
AA-LCR43.70151 / 170Normal (No Tools)40.70157 / 170Normal (No Tools)+3

Writing and Creative Capabilities

GLM-5 1/1
BenchmarkGLM-5GLM-4.7Diff
Creative Writing1,59839 / 106Normal (No Tools)1,41161 / 106Normal (No Tools)+186.70

Specs

FieldGLM-5GLM-4.7
Publisher智谱AI智谱AI
Release date2026-02-112025-12-22
Model typeChat modelChat model
ArchitectureMoEMoE
Parameters744B358B
Context length200K200K
Max output128K132072

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemGLM-5GLM-4.7
Text input$1 / 1M tokens¥4 / 1M tokens
Text output$3.2 / 1M tokens¥16 / 1M tokens
Cache readNot public¥2 / 1M tokens
Cache write$0.2 / 1M tokens¥0 / 1M tokens

Summary

  • GLM-5leads in:Agent Level Benchmark (3/4), General Knowledge (3/3), AI Agent - Information Search (1/1), AI Agent - Tool Usage (1/1), General Evaluation (1/1), Instruction Following (1/1), Long Context (1/1), Writing and Creative Capabilities (1/1)
  • GLM-4.7leads in:Math and Reasoning (1/2)

On average across the 15 shared benchmarks, GLM-5 scores 17.25 higher.

Largest single-benchmark gap: Creative Writing — GLM-5 1,598 vs GLM-4.7 1,411 (+186.70).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.