DataLearner logo

Claude Sonnet 5vsGLM-5.2

Claude Sonnet 5 and GLM-5.2 are tied across 9 shared benchmarks: Claude Sonnet 5 leads on 4, GLM-5.2 leads on 4, with 1 ties and an average score difference of -3.48.

Anthropic
Claude Sonnet 5

Anthropic · 2026-06-30 · Multimodal model

智谱AI
GLM-5.2

智谱AI · 2026-06-13 · Reasoning model

Claude Sonnet 54 wins(44%)Ties1(44%)4 winsGLM-5.2

Benchmark scores

Grouped by capability, sorted by largest gap within each. 9 shared benchmarks.

Coding and Software Engineer

Claude Sonnet 5 2/3
BenchmarkClaude Sonnet 5GLM-5.2Diff
Text Arena (Coding)1,54410 / 35Thinking High (No Tools)1,5935 / 35最高(无工具)-49.09
DeepSWE5416 / 27Deep Thinking (With Tools)4421 / 27Deep Thinking (With Tools)+10
WeirdML v268.7820 / 52Thinking High (With Tools)67.3121 / 52Thinking High (With Tools)+1.47

Math and Reasoning

Claude Sonnet 5 1/2
BenchmarkClaude Sonnet 5GLM-5.2Diff
FrontierMath v265.6117 / 34最高(无工具)59.2119 / 34最高(无工具)+6.41
FrontierMath Tier 4 v229.2716 / 34最高(无工具)29.2716 / 34最高(无工具)

AI Agent - Tool Usage

GLM-5.2 1/1
BenchmarkClaude Sonnet 5GLM-5.2Diff
Terminal-Bench 2.180.4018 / 44极高强度思考(工具)8117 / 44Thinking High (With Tools)-0.60

Commonsense Reasoning

GLM-5.2 1/1
BenchmarkClaude Sonnet 5GLM-5.2Diff
SimpleBench57.9023 / 67Normal (No Tools)58.8020 / 67Normal (No Tools)-0.90

General Evaluation

GLM-5.2 1/1
BenchmarkClaude Sonnet 5GLM-5.2Diff
GPQA Diamond90.5333 / 226极高强度思考(无工具)91.8626 / 226最高(无工具)-1.33

General Knowledge

Claude Sonnet 5 1/1
BenchmarkClaude Sonnet 5GLM-5.2Diff
HLE57.409 / 181极高强度思考(工具)54.7015 / 181Thinking (With Tools)+2.70

Specs

FieldClaude Sonnet 5GLM-5.2
PublisherAnthropic智谱AI
Release date2026-06-302026-06-13
Model typeMultimodal modelReasoning model
ArchitectureDenseMoE
ParametersNot available753.33B
Context length1M1M
Max output128K128K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemClaude Sonnet 5GLM-5.2
Text input$2 / 1M tokens$1.4 / 1M tokens
Text output$10 / 1M tokens$4.4 / 1M tokens
Cache read$0.3 / 1M tokens$0.26 / 1M tokens
Cache write$3.75 / 1M tokensNot public

Summary

  • Claude Sonnet 5leads in:Coding and Software Engineer (2/3), Math and Reasoning (1/2), General Knowledge (1/1)
  • GLM-5.2leads in:AI Agent - Tool Usage (1/1), Commonsense Reasoning (1/1), General Evaluation (1/1)

On average across the 9 shared benchmarks, GLM-5.2 scores 3.48 higher.

Largest single-benchmark gap: Text Arena (Coding) — Claude Sonnet 5 1,544 vs GLM-5.2 1,593 (-49.09).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.