Kimi K3vsGLM-5.2
Across 15 shared benchmarks, Kimi K3 leads overall: Kimi K3 wins 14, GLM-5.2 wins 1, with 0 ties and an average score difference of +37.55.
Kimi K314 wins(93%)(7%)1 winGLM-5.2
Benchmark scores
Grouped by capability, sorted by largest gap within each. 15 shared benchmarks.
Coding and Software Engineer
Kimi K3 6/6| Benchmark | Kimi K3 | GLM-5.2 | Diff |
|---|---|---|---|
| Text Arena (Coding) | 1,6822 / 35Max (No Tools) | 1,5935 / 35Max (No Tools) | +88.50 |
| SWE-Marathon | 423 / 6Max (With Tools) | 136 / 6Max (With Tools) | +29 |
| DeepSWE | 67.509 / 35Max (With Tools) | 4429 / 35Deep Thinking (With Tools) | +23.50 |
| Program Bench | 77.801 / 7Max (With Tools) | 63.702 / 7Thinking (With Tools) | +14.10 |
| FrontierSWE | 81.201 / 4Max (With Tools) | 74.403 / 4Max (With Tools) | +6.80 |
| PostTrain Bench | 36.603 / 5Max (With Tools) | 34.305 / 5Max (With Tools) | +2.30 |
AI Agent - Tool Usage
Kimi K3 2/2| Benchmark | Kimi K3 | GLM-5.2 | Diff |
|---|---|---|---|
| MCP-Atlas | 84.204 / 41Max (With Tools) | 76.8015 / 41Thinking (With Tools) | +7.40 |
| Terminal-Bench 2.1 | 88.304 / 49Max (With Tools) | 8122 / 49Thinking High (With Tools) | +7.30 |
Math and Reasoning
Kimi K3 2/2| Benchmark | Kimi K3 | GLM-5.2 | Diff |
|---|---|---|---|
| FrontierMath v2 | 72.1813 / 58Max (No Tools) | 42.4634 / 58Normal (No Tools) | +29.73 |
| FrontierMath Tier 4 v2 | 39.0214 / 41Max (No Tools) | 29.2719 / 41Max (No Tools) | +9.76 |
Commonsense Reasoning
Kimi K3 1/1| Benchmark | Kimi K3 | GLM-5.2 | Diff |
|---|---|---|---|
| SimpleBench | 60.7029 / 92Max (No Tools) | 58.8034 / 92Normal (No Tools) | +1.90 |
General Evaluation
Kimi K3 1/1| Benchmark | Kimi K3 | GLM-5.2 | Diff |
|---|---|---|---|
| GPQA Diamond | 93.5016 / 271Max (No Tools) | 71.21188 / 271Normal (No Tools) | +22.29 |
General Knowledge
Kimi K3 1/1| Benchmark | Kimi K3 | GLM-5.2 | Diff |
|---|---|---|---|
| HLE | 5617 / 190Max (With Tools) | 54.7020 / 190Thinking (With Tools) | +1.30 |
Text Embedding
GLM-5.2 1/1| Benchmark | Kimi K3 | GLM-5.2 | Diff |
|---|---|---|---|
| Context Arena | 71.7559 / 126Max (No Tools) | 72.3458 / 126Max (No Tools) | -0.59 |
Writing and Creative Capabilities
Kimi K3 1/1| Benchmark | Kimi K3 | GLM-5.2 | Diff |
|---|---|---|---|
| Creative Writing | 2,0712 / 99Normal (No Tools) | 1,75119 / 99Normal (No Tools) | +319.90 |
Specs
| Field | Kimi K3 | GLM-5.2 |
|---|---|---|
| Publisher | Moonshot AI | 智谱AI |
| Release date | 2026-07-16 | 2026-06-13 |
| Model type | Reasoning model | Reasoning model |
| Architecture | MoE | MoE |
| Parameters | 2.8T | 753.33B |
| Context length | 1M | 1M |
| Max output | 1M | 128K |
API pricing
Prices use DataLearner records when available; missing fields are not inferred.
| Item | Kimi K3 | GLM-5.2 |
|---|---|---|
| Text input | ¥20 / 1M tokens | $1.4 / 1M tokens |
| Text output | ¥100 / 1M tokens | $4.4 / 1M tokens |
| Cache read | ¥2 / 1M tokens | $0.26 / 1M tokens |
Summary
- Kimi K3leads in:Coding and Software Engineer (6/6), AI Agent - Tool Usage (2/2), Math and Reasoning (2/2), Commonsense Reasoning (1/1), General Evaluation (1/1), General Knowledge (1/1), Writing and Creative Capabilities (1/1)
- GLM-5.2leads in:Text Embedding (1/1)
On average across the 15 shared benchmarks, Kimi K3 scores 37.55 higher.
Largest single-benchmark gap: Creative Writing — Kimi K3 2,071 vs GLM-5.2 1,751 (+319.90).
Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.