GLM-5.3vsGLM-5.2
Across 8 shared benchmarks, GLM-5.3 leads overall: GLM-5.3 wins 7, GLM-5.2 wins 1, with 0 ties and an average score difference of +5.12.
GLM-5.37 wins(88%)(13%)1 winGLM-5.2
Benchmark scores
Grouped by capability, sorted by largest gap within each. 8 shared benchmarks.
Coding and Software Engineer
GLM-5.3 5/6| Benchmark | GLM-5.3 | GLM-5.2 | Diff |
|---|---|---|---|
| Program Bench | 195 / 5Max (With Tools) | 63.702 / 5Thinking (With Tools) | -44.70 |
| SWE-Marathon | 42.501 / 4Max (With Tools) | 134 / 4Max (With Tools) | +29.50 |
| DeepSWE | 66.908 / 26Max (With Tools) | 4420 / 26Deep Thinking (With Tools) | +22.90 |
| NL2Repo-Bench | 582 / 7Max (With Tools) | 48.905 / 7Thinking (With Tools) | +9.10 |
| PostTrain Bench | 39.801 / 4Max (With Tools) | 34.304 / 4Max (With Tools) | +5.50 |
| FrontierSWE | 78.102 / 4Max (With Tools) | 74.403 / 4Max (With Tools) | +3.70 |
AI Agent - Tool Usage
GLM-5.3 1/1| Benchmark | GLM-5.3 | GLM-5.2 | Diff |
|---|---|---|---|
| Terminal-Bench 2.1 | 88.203 / 43Max (With Tools) | 8116 / 43Thinking High (With Tools) | +7.20 |
General Knowledge
GLM-5.3 1/1| Benchmark | GLM-5.3 | GLM-5.2 | Diff |
|---|---|---|---|
| HLE | 62.503 / 181Max (With Tools) | 54.7015 / 181Thinking (With Tools) | +7.80 |
Specs
| Field | GLM-5.3 | GLM-5.2 |
|---|---|---|
| Publisher | 智谱AI | 智谱AI |
| Release date | 2026-08-14 | 2026-06-13 |
| Model type | Reasoning model | Reasoning model |
| Architecture | MoE | MoE |
| Parameters | 753.33B | 753.33B |
| Context length | 1M | 1M |
| Max output | 128K | 128K |
API pricing
Prices use DataLearner records when available; missing fields are not inferred.
| Item | GLM-5.3 | GLM-5.2 |
|---|---|---|
| Text input | Not public | $1.4 / 1M tokens |
| Text output | Not public | $4.4 / 1M tokens |
| Cache read | Not public | $0.26 / 1M tokens |
One or both models have incomplete public pricing.
Summary
- GLM-5.3leads in:Coding and Software Engineer (5/6), AI Agent - Tool Usage (1/1), General Knowledge (1/1)
On average across the 8 shared benchmarks, GLM-5.3 scores 5.12 higher.
Largest single-benchmark gap: Program Bench — GLM-5.3 19 vs GLM-5.2 63.70 (-44.70).
Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.