DataLearner logo

GLM-5vsGLM-4.7

Across 11 shared benchmarks, GLM-5 leads overall: GLM-5 wins 9, GLM-4.7 wins 1, with 1 ties and an average score difference of +7.63.

智谱AI
GLM-5

智谱AI · 2026-02-11 · Chat model

智谱AI
GLM-4.7

智谱AI · 2025-12-22 · Chat model

GLM-59 wins(82%)Ties1(9%)1 winGLM-4.7

Benchmark scores

Grouped by capability, sorted by largest gap within each. 11 shared benchmarks.

General Knowledge

GLM-5 3/3
BenchmarkGLM-5GLM-4.7Diff
LiveBench68.8543 / 115Normal (No Tools)58.0978 / 115Normal (No Tools)+10.76
HLE50.4025 / 17242.8052 / 172+7.60
GPQA Diamond8648 / 187Thinking (No Tools)85.7049 / 187+0.30

Agent Level Benchmark

GLM-5 2/2
BenchmarkGLM-5GLM-4.7Diff
Terminal Bench Hard432 / 1333.307 / 13+9.70
τ²-Bench89.704 / 4387.406 / 43+2.30

Math and Reasoning

GLM-4.7 1/2
BenchmarkGLM-5GLM-4.7Diff
AIME 202692.709 / 18Thinking (No Tools)92.908 / 18-0.20
FrontierMath - Tier 42.1056 / 80Normal (No Tools)2.1056 / 80Normal (No Tools)

AI Agent - Information Search

GLM-5 1/1
BenchmarkGLM-5GLM-4.7Diff
BrowseComp75.9024 / 535241 / 53+23.90

AI Agent - Tool Usage

GLM-5 1/1
BenchmarkGLM-5GLM-4.7Diff
Terminal Bench 2.061.1018 / 474144 / 47+20.10

Coding and Software Engineer

GLM-5 1/1
BenchmarkGLM-5GLM-4.7Diff
SWE-bench Verified77.8025 / 112Thinking (No Tools)73.8043 / 112+4

Commonsense Reasoning

GLM-5 1/1
BenchmarkGLM-5GLM-4.7Diff
Simple Bench53.2023 / 63Normal (No Tools)47.7029 / 63Thinking (No Tools)+5.50

Specs

FieldGLM-5GLM-4.7
Publisher智谱AI智谱AI
Release date2026-02-112025-12-22
Model typeChat modelChat model
ArchitectureMoEMoE
Parameters744B358B
Context length200K200K
Max output128K132072

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemGLM-5GLM-4.7
Text input$1 / 1M tokens¥4 / 1M tokens
Text output$3.2 / 1M tokens¥16 / 1M tokens
Cache readNot public¥2 / 1M tokens
Cache write$0.2 / 1M tokens¥0 / 1M tokens

Summary

  • GLM-5leads in:General Knowledge (3/3), Agent Level Benchmark (2/2), AI Agent - Information Search (1/1), AI Agent - Tool Usage (1/1), Coding and Software Engineer (1/1), Commonsense Reasoning (1/1)
  • GLM-4.7leads in:Math and Reasoning (1/2)

On average across the 11 shared benchmarks, GLM-5 scores 7.63 higher.

Largest single-benchmark gap: BrowseComp — GLM-5 75.90 vs GLM-4.7 52 (+23.90).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.