DataLearner logo

GLM-5vsKimi K2.5

Across 13 shared benchmarks, GLM-5 leads overall: GLM-5 wins 7, Kimi K2.5 wins 6, with 0 ties and an average score difference of -1.07.

智谱AI
GLM-5

智谱AI · 2026-02-11 · Chat model

Moonshot AI
Kimi K2.5

Moonshot AI · 2026-01-27 · Multimodal model

GLM-57 wins(54%)(46%)6 winsKimi K2.5

Benchmark scores

Grouped by capability, sorted by largest gap within each. 13 shared benchmarks.

Math and Reasoning

GLM-5 2/3
BenchmarkGLM-5Kimi K2.5Diff
FrontierMath - Tier 42.1056 / 80Normal (No Tools)4.2040 / 80Normal (No Tools)-2.10
IMO-AnswerBench82.5017 / 24Thinking (No Tools)81.8018 / 24Thinking (No Tools)+0.70
AIME 202692.7018 / 29Thinking (No Tools)92.5021 / 29Thinking (No Tools)+0.20

Claw-style Agent Evaluation

GLM-5 2/2
BenchmarkGLM-5Kimi K2.5Diff
Claw Bench91.705 / 29Thinking (With Tools)81.7018 / 29Thinking (With Tools)+10
Pinch Bench86.4013 / 38Thinking (With Tools)84.8018 / 38Thinking (With Tools)+1.60

General Knowledge

Kimi K2.5 2/2
BenchmarkGLM-5Kimi K2.5Diff
ARC-AGI-144.67111 / 147Thinking (No Tools)65.3390 / 147Thinking (No Tools)-20.66
HLE7.60427 / 563Normal (No Tools) · Text only13.20357 / 563Normal (No Tools) · Text only-5.60

Long Context

Kimi K2.5 2/2
BenchmarkGLM-5Kimi K2.5Diff
AA-LCR43.70151 / 170Normal (No Tools)67.30113 / 170Normal (No Tools)-23.60
LongBench v260.807 / 14Normal (No Tools)616 / 14Normal (No Tools)-0.20

AI Agent - Tool Usage

GLM-5 1/1
BenchmarkGLM-5Kimi K2.5Diff
Terminal Bench 2.061.1018 / 48Thinking (With Tools)50.8035 / 48Thinking (With Tools)+10.30

General Evaluation

Kimi K2.5 1/1
BenchmarkGLM-5Kimi K2.5Diff
GPQA Diamond66.60349 / 462Normal (No Tools)78.90248 / 462Normal (No Tools)-12.30

Productivity Knowledge

GLM-5 1/1
BenchmarkGLM-5Kimi K2.5Diff
GDPval-AA468 / 15Thinking (No Tools)409 / 15Thinking (No Tools)+6

Writing and Creative Capabilities

GLM-5 1/1
BenchmarkGLM-5Kimi K2.5Diff
Creative Writing1,59839 / 106Normal (No Tools)1,57643 / 106Normal (No Tools)+21.70

Specs

FieldGLM-5Kimi K2.5
Publisher智谱AIMoonshot AI
Release date2026-02-112026-01-27
Model typeChat modelMultimodal model
ArchitectureMoEMoE
Parameters744B1T
Context length200K256K
Max output128K16K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemGLM-5Kimi K2.5
Text input$1 / 1M tokens$0.6 / 1M tokens
Text output$3.2 / 1M tokens$3 / 1M tokens
Cache readNot public$0.1 / 1M tokens
Cache write$0.2 / 1M tokensNot public

Summary

  • GLM-5leads in:Math and Reasoning (2/3), Claw-style Agent Evaluation (2/2), AI Agent - Tool Usage (1/1), Productivity Knowledge (1/1), Writing and Creative Capabilities (1/1)
  • Kimi K2.5leads in:General Knowledge (2/2), Long Context (2/2), General Evaluation (1/1)

On average across the 13 shared benchmarks, Kimi K2.5 scores 1.07 higher.

Largest single-benchmark gap: AA-LCR — GLM-5 43.70 vs Kimi K2.5 67.30 (-23.60).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.