DataLearner logo

GLM-5.2vsKimi K2.6

Across 10 shared benchmarks, GLM-5.2 leads overall: GLM-5.2 wins 9, Kimi K2.6 wins 1, with 0 ties and an average score difference of +6.59.

智谱AI
GLM-5.2

智谱AI · 2026-06-13 · Reasoning model

Moonshot AI
Kimi K2.6

Moonshot AI · 2026-04-20 · Reasoning model

GLM-5.29 wins(90%)(10%)1 winKimi K2.6

Benchmark scores

Grouped by capability, sorted by largest gap within each. 10 shared benchmarks.

AI Agent - Tool Usage

GLM-5.2 2/3
BenchmarkGLM-5.2Kimi K2.6Diff
Terminal-Bench 2.18117 / 44Thinking High (With Tools)53.5643 / 44Thinking (No Tools)+27.44
MCP-Atlas76.8013 / 38Thinking (With Tools)69.4028 / 38Thinking (With Tools)+7.40
Tool Decathlon48.204 / 10Thinking (With Tools)502 / 10Thinking (With Tools)-1.80

Coding and Software Engineer

GLM-5.2 2/2
BenchmarkGLM-5.2Kimi K2.6Diff
Program Bench63.702 / 5Thinking (With Tools)48.304 / 5Thinking (With Tools)+15.40
SWE-Bench Pro - Public62.109 / 57Thinking (With Tools)58.6015 / 57Thinking (With Tools)+3.50

General Knowledge

GLM-5.2 2/2
BenchmarkGLM-5.2Kimi K2.6Diff
LiveBench76.249 / 115Normal (No Tools)72.1728 / 115Thinking (No Tools)+4.07
HLE54.7015 / 181Thinking (With Tools)5417 / 181Thinking (With Tools + Internet)+0.70

Math and Reasoning

GLM-5.2 2/2
BenchmarkGLM-5.2Kimi K2.6Diff
IMO-AnswerBench912 / 23Thinking (No Tools)8610 / 23Thinking (No Tools)+5
AIME 202699.201 / 19Thinking (No Tools)96.403 / 19Thinking (No Tools)+2.80

General Evaluation

GLM-5.2 1/1
BenchmarkGLM-5.2Kimi K2.6Diff
GPQA Diamond91.8626 / 226最高(无工具)90.5035 / 226Thinking (No Tools)+1.36

Specs

FieldGLM-5.2Kimi K2.6
Publisher智谱AIMoonshot AI
Release date2026-06-132026-04-20
Model typeReasoning modelReasoning model
ArchitectureMoEMoE
Parameters753.33B1T
Context length1M256K
Max output128KNot available

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemGLM-5.2Kimi K2.6
Text input$1.4 / 1M tokens$0.95 / 1M tokens
Text output$4.4 / 1M tokens$4 / 1M tokens
Cache read$0.26 / 1M tokens$0.16 / 1M tokens
Cache writeNot public$0.95 / 1M tokens

Summary

  • GLM-5.2leads in:AI Agent - Tool Usage (2/3), Coding and Software Engineer (2/2), General Knowledge (2/2), Math and Reasoning (2/2), General Evaluation (1/1)

On average across the 10 shared benchmarks, GLM-5.2 scores 6.59 higher.

Largest single-benchmark gap: Terminal-Bench 2.1 — GLM-5.2 81 vs Kimi K2.6 53.56 (+27.44).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.