DataLearner logo

Grok 4.5vsGLM-5.2

Across 7 shared benchmarks, Grok 4.5 leads overall: Grok 4.5 wins 5, GLM-5.2 wins 2, with 0 ties and an average score difference of +3.51.

xAI
Grok 4.5

xAI · 2026-07-08 · Coding model

智谱AI
GLM-5.2

智谱AI · 2026-06-13 · Reasoning model

Grok 4.55 wins(71%)(29%)2 winsGLM-5.2

Benchmark scores

Grouped by capability, sorted by largest gap within each. 7 shared benchmarks.

Coding and Software Engineer

Grok 4.5 3/3
BenchmarkGrok 4.5GLM-5.2Diff
SWE-Marathon293 / 4Thinking High (With Tools)134 / 4Max (With Tools)+16
DeepSWE5318 / 27Thinking High (With Tools)4421 / 27Deep Thinking (With Tools)+9
SWE-Bench Pro - Public64.706 / 57Thinking High (With Tools)62.109 / 57Thinking (With Tools)+2.60

Math and Reasoning

GLM-5.2 2/2
BenchmarkGrok 4.5GLM-5.2Diff
FrontierMath Tier 4 v224.3920 / 34Thinking High (No Tools)29.2716 / 34最高(无工具)-4.88
FrontierMath v257.1921 / 34Thinking High (No Tools)59.2119 / 34最高(无工具)-2.01

AI Agent - Tool Usage

Grok 4.5 1/1
BenchmarkGrok 4.5GLM-5.2Diff
Terminal-Bench 2.183.3014 / 44Thinking High (With Tools)8117 / 44Thinking High (With Tools)+2.30

General Evaluation

Grok 4.5 1/1
BenchmarkGrok 4.5GLM-5.2Diff
GPQA Diamond93.4315 / 226Thinking High (No Tools)91.8626 / 226最高(无工具)+1.58

Specs

FieldGrok 4.5GLM-5.2
PublisherxAI智谱AI
Release date2026-07-082026-06-13
Model typeCoding modelReasoning model
ArchitectureDenseMoE
ParametersNot available753.33B
Context length500K1M
Max outputNot available128K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemGrok 4.5GLM-5.2
Text input$2 / 1M tokens$1.4 / 1M tokens
Text output$6 / 1M tokens$4.4 / 1M tokens
Cache read$0.5 / 1M tokens$0.26 / 1M tokens

Summary

  • Grok 4.5leads in:Coding and Software Engineer (3/3), AI Agent - Tool Usage (1/1), General Evaluation (1/1)
  • GLM-5.2leads in:Math and Reasoning (2/2)

On average across the 7 shared benchmarks, Grok 4.5 scores 3.51 higher.

Largest single-benchmark gap: SWE-Marathon — Grok 4.5 29 vs GLM-5.2 13 (+16).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.