Hy3vsGLM-5.2
Across 8 shared benchmarks, GLM-5.2 leads overall: Hy3 wins 2, GLM-5.2 wins 6, with 0 ties and an average score difference of -3.86.
Hy32 wins(25%)(75%)6 winsGLM-5.2
Benchmark scores
Grouped by capability, sorted by largest gap within each. 8 shared benchmarks.
AI Agent - Tool Usage
Hy3 2/3| Benchmark | Hy3 | GLM-5.2 | Diff |
|---|---|---|---|
| Terminal-Bench 2.1 | 71.7030 / 44Thinking High (With Tools) | 8117 / 44Thinking High (With Tools) | -9.30 |
| MCP-Atlas | 79.109 / 38Thinking High (With Tools) | 76.8013 / 38Thinking (With Tools) | +2.30 |
| Tool Decathlon | 48.503 / 10Thinking High (With Tools) | 48.204 / 10Thinking (With Tools) | +0.30 |
Coding and Software Engineer
GLM-5.2 2/2| Benchmark | Hy3 | GLM-5.2 | Diff |
|---|---|---|---|
| DeepSWE | 2826 / 27Thinking High (With Tools) | 4421 / 27Deep Thinking (With Tools) | -16 |
| SWE-Bench Pro - Public | 57.9018 / 57Thinking High (With Tools) | 62.109 / 57Thinking (With Tools) | -4.20 |
General Evaluation
GLM-5.2 1/1| Benchmark | Hy3 | GLM-5.2 | Diff |
|---|---|---|---|
| GPQA Diamond | 90.4036 / 226Thinking High (No Tools) | 91.8626 / 226最高(无工具) | -1.46 |
General Knowledge
GLM-5.2 1/1| Benchmark | Hy3 | GLM-5.2 | Diff |
|---|---|---|---|
| HLE | 53.2019 / 181Thinking High (With Tools) | 54.7015 / 181Thinking (With Tools) | -1.50 |
Math and Reasoning
GLM-5.2 1/1| Benchmark | Hy3 | GLM-5.2 | Diff |
|---|---|---|---|
| IMO-AnswerBench | 903 / 23Thinking High (No Tools) | 912 / 23Thinking (No Tools) | -1 |
Specs
| Field | Hy3 | GLM-5.2 |
|---|---|---|
| Publisher | 腾讯AI实验室 | 智谱AI |
| Release date | 2026-07-06 | 2026-06-13 |
| Model type | Reasoning model | Reasoning model |
| Architecture | MoE | MoE |
| Parameters | 295B | 753.33B |
| Context length | 256K | 1M |
| Max output | Not available | 128K |
API pricing
Prices use DataLearner records when available; missing fields are not inferred.
| Item | Hy3 | GLM-5.2 |
|---|---|---|
| Text input | ¥1.2 / 1M tokens | $1.4 / 1M tokens |
| Text output | ¥4 / 1M tokens | $4.4 / 1M tokens |
| Cache read | ¥0.4 / 1M tokens | $0.26 / 1M tokens |
Summary
- Hy3leads in:AI Agent - Tool Usage (2/3)
- GLM-5.2leads in:Coding and Software Engineer (2/2), General Evaluation (1/1), General Knowledge (1/1), Math and Reasoning (1/1)
On average across the 8 shared benchmarks, GLM-5.2 scores 3.86 higher.
Largest single-benchmark gap: DeepSWE — Hy3 28 vs GLM-5.2 44 (-16).
Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.