DataLearner logo

GLM-5.3vsGrok 4.6

Across 3 shared benchmarks, GLM-5.3 leads overall: GLM-5.3 wins 3, Grok 4.6 wins 0, with 0 ties and an average score difference of +6.43.

智谱AI
GLM-5.3

智谱AI · 2026-08-14 · Reasoning model

xAI
Grok 4.6

xAI · 2026-08-12 · Coding model

GLM-5.33 wins(100%)(0%)0 winsGrok 4.6

Benchmark scores

Grouped by capability, sorted by largest gap within each. 3 shared benchmarks.

AI Agent - Tool Usage

GLM-5.3 1/1
BenchmarkGLM-5.3Grok 4.6Diff
Terminal-Bench 3.028.303 / 6Max (With Tools)264 / 6Thinking High (With Tools)+2.30

Coding and Software Engineer

GLM-5.3 1/1
BenchmarkGLM-5.3Grok 4.6Diff
DeepSWE66.908 / 27Max (With Tools)65.909 / 27Thinking High (With Tools)+1

Productivity Knowledge

GLM-5.3 1/1
BenchmarkGLM-5.3Grok 4.6Diff
GDPval-AA v21,7692 / 13Max (With Tools)1,7533 / 13Thinking High (With Tools)+16

Specs

FieldGLM-5.3Grok 4.6
Publisher智谱AIxAI
Release date2026-08-142026-08-12
Model typeReasoning modelCoding model
ArchitectureMoEDense
Parameters753.33BNot available
Context length1M500K
Max output128KNot available

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemGLM-5.3Grok 4.6
Text inputNot public$2 / 1M tokens
Text outputNot public$6 / 1M tokens
Cache readNot public$0.5 / 1M tokens

One or both models have incomplete public pricing.

Summary

  • GLM-5.3leads in:AI Agent - Tool Usage (1/1), Coding and Software Engineer (1/1), Productivity Knowledge (1/1)

On average across the 3 shared benchmarks, GLM-5.3 scores 6.43 higher.

Largest single-benchmark gap: GDPval-AA v2 — GLM-5.3 1,769 vs Grok 4.6 1,753 (+16).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.