DataLearner logo

Kimi K3vsGLM-5.2

Across 4 shared benchmarks, Kimi K3 leads overall: Kimi K3 wins 4, GLM-5.2 wins 0, with 0 ties and an average score difference of +8.60.

Moonshot AI
Kimi K3

Moonshot AI · 2026-07-16 · Reasoning model

智谱AI
GLM-5.2

智谱AI · 2026-06-13 · Reasoning model

Kimi K34 wins(100%)(0%)0 winsGLM-5.2

Benchmark scores

Grouped by capability, sorted by largest gap within each. 4 shared benchmarks.

General Knowledge

Kimi K3 2/2
BenchmarkKimi K3GLM-5.2Diff
GPQA Diamond93.508 / 187最高(无工具)91.2016 / 187Thinking (No Tools)+2.30
HLE5610 / 170Max (With Tools)54.7011 / 170Thinking (With Tools)+1.30

AI Agent - Tool Usage

Kimi K3 1/1
BenchmarkKimi K3GLM-5.2Diff
TerminalBench 2.188.302 / 25Max (With Tools)819 / 25Thinking High (With Tools)+7.30

Coding and Software Engineer

Kimi K3 1/1
BenchmarkKimi K3GLM-5.2Diff
DeepSWE67.504 / 17Max (With Tools)4412 / 17Deep Thinking (With Tools)+23.50

Specs

FieldKimi K3GLM-5.2
PublisherMoonshot AI智谱AI
Release date2026-07-162026-06-13
Model typeReasoning modelReasoning model
ArchitectureMoEMoE
Parameters2.8T753.33B
Context length1M1M
Max output1M128K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemKimi K3GLM-5.2
Text input¥20 / 1M tokens$1.4 / 1M tokens
Text output¥100 / 1M tokens$4.4 / 1M tokens
Cache read¥2 / 1M tokens$0.26 / 1M tokens

Summary

  • Kimi K3leads in:General Knowledge (2/2), AI Agent - Tool Usage (1/1), Coding and Software Engineer (1/1)

On average across the 4 shared benchmarks, Kimi K3 scores 8.60 higher.

Largest single-benchmark gap: DeepSWE — Kimi K3 67.50 vs GLM-5.2 44 (+23.50).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.