DataLearner logo

Hy3vsGLM-5.2

Across 8 shared benchmarks, GLM-5.2 leads overall: Hy3 wins 2, GLM-5.2 wins 6, with 0 ties and an average score difference of -3.86.

腾讯AI实验室
Hy3

腾讯AI实验室 · 2026-07-06 · Reasoning model

智谱AI
GLM-5.2

智谱AI · 2026-06-13 · Reasoning model

Hy32 wins(25%)(75%)6 winsGLM-5.2

Benchmark scores

Grouped by capability, sorted by largest gap within each. 8 shared benchmarks.

AI Agent - Tool Usage

Hy3 2/3
BenchmarkHy3GLM-5.2Diff
Terminal-Bench 2.171.7030 / 44Thinking High (With Tools)8117 / 44Thinking High (With Tools)-9.30
MCP-Atlas79.109 / 38Thinking High (With Tools)76.8013 / 38Thinking (With Tools)+2.30
Tool Decathlon48.503 / 10Thinking High (With Tools)48.204 / 10Thinking (With Tools)+0.30

Coding and Software Engineer

GLM-5.2 2/2
BenchmarkHy3GLM-5.2Diff
DeepSWE2826 / 27Thinking High (With Tools)4421 / 27Deep Thinking (With Tools)-16
SWE-Bench Pro - Public57.9018 / 57Thinking High (With Tools)62.109 / 57Thinking (With Tools)-4.20

General Evaluation

GLM-5.2 1/1
BenchmarkHy3GLM-5.2Diff
GPQA Diamond90.4036 / 226Thinking High (No Tools)91.8626 / 226最高(无工具)-1.46

General Knowledge

GLM-5.2 1/1
BenchmarkHy3GLM-5.2Diff
HLE53.2019 / 181Thinking High (With Tools)54.7015 / 181Thinking (With Tools)-1.50

Math and Reasoning

GLM-5.2 1/1
BenchmarkHy3GLM-5.2Diff
IMO-AnswerBench903 / 23Thinking High (No Tools)912 / 23Thinking (No Tools)-1

Specs

FieldHy3GLM-5.2
Publisher腾讯AI实验室智谱AI
Release date2026-07-062026-06-13
Model typeReasoning modelReasoning model
ArchitectureMoEMoE
Parameters295B753.33B
Context length256K1M
Max outputNot available128K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemHy3GLM-5.2
Text input¥1.2 / 1M tokens$1.4 / 1M tokens
Text output¥4 / 1M tokens$4.4 / 1M tokens
Cache read¥0.4 / 1M tokens$0.26 / 1M tokens

Summary

  • Hy3leads in:AI Agent - Tool Usage (2/3)
  • GLM-5.2leads in:Coding and Software Engineer (2/2), General Evaluation (1/1), General Knowledge (1/1), Math and Reasoning (1/1)

On average across the 8 shared benchmarks, GLM-5.2 scores 3.86 higher.

Largest single-benchmark gap: DeepSWE — Hy3 28 vs GLM-5.2 44 (-16).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.