DataLearner logo

GLM-5.3vsKimi K3

Across 11 shared benchmarks, GLM-5.3 leads overall: GLM-5.3 wins 6, Kimi K3 wins 5, with 0 ties and an average score difference of +4.06.

智谱AI
GLM-5.3

智谱AI · 2026-08-14 · Reasoning model

Moonshot AI
Kimi K3

Moonshot AI · 2026-07-16 · Reasoning model

GLM-5.36 wins(55%)(45%)5 winsKimi K3

Benchmark scores

Grouped by capability, sorted by largest gap within each. 11 shared benchmarks.

Coding and Software Engineer

Kimi K3 3/5
BenchmarkGLM-5.3Kimi K3Diff
Program Bench195 / 5Max (With Tools)77.801 / 5Max (With Tools)-58.80
PostTrain Bench39.801 / 4Max (With Tools)36.603 / 4Max (With Tools)+3.20
FrontierSWE78.102 / 4Max (With Tools)81.201 / 4Max (With Tools)-3.10
DeepSWE66.908 / 26Max (With Tools)67.505 / 26Max (With Tools)-0.60
SWE-Marathon42.501 / 4Max (With Tools)422 / 4Max (With Tools)+0.50

AI Agent - Tool Usage

Kimi K3 2/3
BenchmarkGLM-5.3Kimi K3Diff
Automation Bench48.201 / 7Max (With Tools)30.803 / 7Max (With Tools)+17.40
Toolathlon-Verified733 / 5Max (With Tools)76.501 / 5Max (With Tools)-3.50
Terminal-Bench 2.188.203 / 43Max (With Tools)88.302 / 43Max (With Tools)-0.10

Agent Level Benchmark

GLM-5.3 1/1
BenchmarkGLM-5.3Kimi K3Diff
Agents' Last Exam28.504 / 10Max (With Tools)28.305 / 10Max (With Tools)+0.20

General Knowledge

GLM-5.3 1/1
BenchmarkGLM-5.3Kimi K3Diff
HLE62.503 / 181Max (With Tools)5614 / 181Max (With Tools)+6.50

Productivity Knowledge

GLM-5.3 1/1
BenchmarkGLM-5.3Kimi K3Diff
GDPval-AA v21,7692 / 13Max (With Tools)1,6866 / 13Max (With Tools)+83

Specs

FieldGLM-5.3Kimi K3
Publisher智谱AIMoonshot AI
Release date2026-08-142026-07-16
Model typeReasoning modelReasoning model
ArchitectureMoEMoE
Parameters753.33B2.8T
Context length1M1M
Max output128K1M

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemGLM-5.3Kimi K3
Text inputNot public¥20 / 1M tokens
Text outputNot public¥100 / 1M tokens
Cache readNot public¥2 / 1M tokens

One or both models have incomplete public pricing.

Summary

  • GLM-5.3leads in:Agent Level Benchmark (1/1), General Knowledge (1/1), Productivity Knowledge (1/1)
  • Kimi K3leads in:Coding and Software Engineer (3/5), AI Agent - Tool Usage (2/3)

On average across the 11 shared benchmarks, GLM-5.3 scores 4.06 higher.

Largest single-benchmark gap: GDPval-AA v2 — GLM-5.3 1,769 vs Kimi K3 1,686 (+83).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.