GLM-5.2vsKimi K2.6
Across 10 shared benchmarks, GLM-5.2 leads overall: GLM-5.2 wins 9, Kimi K2.6 wins 1, with 0 ties and an average score difference of +6.59.
GLM-5.29 wins(90%)(10%)1 winKimi K2.6
Benchmark scores
Grouped by capability, sorted by largest gap within each. 10 shared benchmarks.
AI Agent - Tool Usage
GLM-5.2 2/3| Benchmark | GLM-5.2 | Kimi K2.6 | Diff |
|---|---|---|---|
| Terminal-Bench 2.1 | 8117 / 44Thinking High (With Tools) | 53.5643 / 44Thinking (No Tools) | +27.44 |
| MCP-Atlas | 76.8013 / 38Thinking (With Tools) | 69.4028 / 38Thinking (With Tools) | +7.40 |
| Tool Decathlon | 48.204 / 10Thinking (With Tools) | 502 / 10Thinking (With Tools) | -1.80 |
Coding and Software Engineer
GLM-5.2 2/2| Benchmark | GLM-5.2 | Kimi K2.6 | Diff |
|---|---|---|---|
| Program Bench | 63.702 / 5Thinking (With Tools) | 48.304 / 5Thinking (With Tools) | +15.40 |
| SWE-Bench Pro - Public | 62.109 / 57Thinking (With Tools) | 58.6015 / 57Thinking (With Tools) | +3.50 |
General Knowledge
GLM-5.2 2/2| Benchmark | GLM-5.2 | Kimi K2.6 | Diff |
|---|---|---|---|
| LiveBench | 76.249 / 115Normal (No Tools) | 72.1728 / 115Thinking (No Tools) | +4.07 |
| HLE | 54.7015 / 181Thinking (With Tools) | 5417 / 181Thinking (With Tools + Internet) | +0.70 |
Math and Reasoning
GLM-5.2 2/2| Benchmark | GLM-5.2 | Kimi K2.6 | Diff |
|---|---|---|---|
| IMO-AnswerBench | 912 / 23Thinking (No Tools) | 8610 / 23Thinking (No Tools) | +5 |
| AIME 2026 | 99.201 / 19Thinking (No Tools) | 96.403 / 19Thinking (No Tools) | +2.80 |
General Evaluation
GLM-5.2 1/1| Benchmark | GLM-5.2 | Kimi K2.6 | Diff |
|---|---|---|---|
| GPQA Diamond | 91.8626 / 226最高(无工具) | 90.5035 / 226Thinking (No Tools) | +1.36 |
Specs
| Field | GLM-5.2 | Kimi K2.6 |
|---|---|---|
| Publisher | 智谱AI | Moonshot AI |
| Release date | 2026-06-13 | 2026-04-20 |
| Model type | Reasoning model | Reasoning model |
| Architecture | MoE | MoE |
| Parameters | 753.33B | 1T |
| Context length | 1M | 256K |
| Max output | 128K | Not available |
API pricing
Prices use DataLearner records when available; missing fields are not inferred.
| Item | GLM-5.2 | Kimi K2.6 |
|---|---|---|
| Text input | $1.4 / 1M tokens | $0.95 / 1M tokens |
| Text output | $4.4 / 1M tokens | $4 / 1M tokens |
| Cache read | $0.26 / 1M tokens | $0.16 / 1M tokens |
| Cache write | Not public | $0.95 / 1M tokens |
Summary
- GLM-5.2leads in:AI Agent - Tool Usage (2/3), Coding and Software Engineer (2/2), General Knowledge (2/2), Math and Reasoning (2/2), General Evaluation (1/1)
On average across the 10 shared benchmarks, GLM-5.2 scores 6.59 higher.
Largest single-benchmark gap: Terminal-Bench 2.1 — GLM-5.2 81 vs Kimi K2.6 53.56 (+27.44).
Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.