Grok 4.5vsGLM-5.2
Across 7 shared benchmarks, Grok 4.5 leads overall: Grok 4.5 wins 5, GLM-5.2 wins 2, with 0 ties and an average score difference of +3.51.
Grok 4.55 wins(71%)(29%)2 winsGLM-5.2
Benchmark scores
Grouped by capability, sorted by largest gap within each. 7 shared benchmarks.
Coding and Software Engineer
Grok 4.5 3/3| Benchmark | Grok 4.5 | GLM-5.2 | Diff |
|---|---|---|---|
| SWE-Marathon | 293 / 4Thinking High (With Tools) | 134 / 4Max (With Tools) | +16 |
| DeepSWE | 5318 / 27Thinking High (With Tools) | 4421 / 27Deep Thinking (With Tools) | +9 |
| SWE-Bench Pro - Public | 64.706 / 57Thinking High (With Tools) | 62.109 / 57Thinking (With Tools) | +2.60 |
Math and Reasoning
GLM-5.2 2/2| Benchmark | Grok 4.5 | GLM-5.2 | Diff |
|---|---|---|---|
| FrontierMath Tier 4 v2 | 24.3920 / 34Thinking High (No Tools) | 29.2716 / 34最高(无工具) | -4.88 |
| FrontierMath v2 | 57.1921 / 34Thinking High (No Tools) | 59.2119 / 34最高(无工具) | -2.01 |
AI Agent - Tool Usage
Grok 4.5 1/1| Benchmark | Grok 4.5 | GLM-5.2 | Diff |
|---|---|---|---|
| Terminal-Bench 2.1 | 83.3014 / 44Thinking High (With Tools) | 8117 / 44Thinking High (With Tools) | +2.30 |
General Evaluation
Grok 4.5 1/1| Benchmark | Grok 4.5 | GLM-5.2 | Diff |
|---|---|---|---|
| GPQA Diamond | 93.4315 / 226Thinking High (No Tools) | 91.8626 / 226最高(无工具) | +1.58 |
Specs
| Field | Grok 4.5 | GLM-5.2 |
|---|---|---|
| Publisher | xAI | 智谱AI |
| Release date | 2026-07-08 | 2026-06-13 |
| Model type | Coding model | Reasoning model |
| Architecture | Dense | MoE |
| Parameters | Not available | 753.33B |
| Context length | 500K | 1M |
| Max output | Not available | 128K |
API pricing
Prices use DataLearner records when available; missing fields are not inferred.
| Item | Grok 4.5 | GLM-5.2 |
|---|---|---|
| Text input | $2 / 1M tokens | $1.4 / 1M tokens |
| Text output | $6 / 1M tokens | $4.4 / 1M tokens |
| Cache read | $0.5 / 1M tokens | $0.26 / 1M tokens |
Summary
- Grok 4.5leads in:Coding and Software Engineer (3/3), AI Agent - Tool Usage (1/1), General Evaluation (1/1)
- GLM-5.2leads in:Math and Reasoning (2/2)
On average across the 7 shared benchmarks, Grok 4.5 scores 3.51 higher.
Largest single-benchmark gap: SWE-Marathon — Grok 4.5 29 vs GLM-5.2 13 (+16).
Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.