DataLearner logo

Hy3 PrevsGLM 5.1

Across 5 shared benchmarks, GLM 5.1 leads overall: Hy3 Pre wins 0, GLM 5.1 wins 5, with 0 ties and an average score difference of -13.80.

腾讯AI实验室
Hy3 Pre

腾讯AI实验室 · 2026-04-23 · Reasoning model

智谱AI
GLM 5.1

智谱AI · 2026-03-27 · Reasoning model

Hy3 Pre0 wins(0%)(100%)5 winsGLM 5.1

Benchmark scores

Grouped by capability, sorted by largest gap within each. 5 shared benchmarks.

Agent Level Benchmark

GLM 5.1 2/2
BenchmarkHy3 PreGLM 5.1Diff
τ²-Bench - Telecom67.50143 / 264Normal (With Tools)97.1017 / 264Normal (With Tools)-29.60
Terminal Bench Hard31.80100 / 244Normal (With Tools)35.6070 / 244Normal (With Tools)-3.80

General Evaluation

GLM 5.1 1/1
BenchmarkHy3 PreGLM 5.1Diff
GPQA Diamond73.20299 / 462Normal (No Tools)83.90180 / 462Normal (No Tools)-10.70

General Knowledge

GLM 5.1 1/1
BenchmarkHy3 PreGLM 5.1Diff
HLE7440 / 563Normal (No Tools) · Text only27.90226 / 563Normal (No Tools) · Text only-20.90

Instruction Following

GLM 5.1 1/1
BenchmarkHy3 PreGLM 5.1Diff
IF Bench48166 / 282Normal (No Tools)52148 / 282Normal (No Tools)-4

Specs

FieldHy3 PreGLM 5.1
Publisher腾讯AI实验室智谱AI
Release date2026-04-232026-03-27
Model typeReasoning modelReasoning model
ArchitectureMoEMoE
Parameters295B754B
Context length256K200K
Max outputNot available125K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemHy3 PreGLM 5.1
Text inputNot public$1.4 / 1M tokens
Text outputNot public$4.4 / 1M tokens
Cache readNot public$4.4 / 1M tokens
Cache writeNot public$0.26 / 1M tokens

One or both models have incomplete public pricing.

Summary

  • GLM 5.1leads in:Agent Level Benchmark (2/2), General Evaluation (1/1), General Knowledge (1/1), Instruction Following (1/1)

On average across the 5 shared benchmarks, GLM 5.1 scores 13.80 higher.

Largest single-benchmark gap: τ²-Bench - Telecom — Hy3 Pre 67.50 vs GLM 5.1 97.10 (-29.60).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.