DataLearner 标志

GLM-5.2vsGLM 5.1

在 11 个共同 benchmark 中,GLM-5.2 整体领先:GLM-5.2 领先 8 项,GLM 5.1 领先 3 项,持平 0 项,平均分差 +16.17。

智谱AI
GLM-5.2

智谱AI · 2026-06-13 · 推理大模型

智谱AI
GLM 5.1

智谱AI · 2026-03-27 · 推理大模型

GLM-5.28 (73%)(27%)3 GLM 5.1

评测分数

按能力类目分组,每组内按分差大小排列;共 11 项。

General Knowledge

胶着 2/2
评测项GLM-5.2GLM 5.1分差
HLE9.80392 / 563Normal (No Tools) · Text only27.90226 / 563Normal (No Tools) · Text only-18.10
LiveBench73.1821 / 117Normal (No Tools)70.1837 / 117Normal (No Tools)+3

Math and Reasoning

GLM-5.2 领先 2/2
评测项GLM-5.2GLM 5.1分差
IMO-AnswerBench912 / 24Thinking (No Tools)83.8014 / 24Thinking (No Tools)+7.20
AIME 202699.203 / 30Thinking (No Tools)95.3013 / 30Thinking (No Tools)+3.90

AI Agent - Tool Usage

GLM-5.2 领先 1/1
评测项GLM-5.2GLM 5.1分差
Tool Decathlon48.204 / 10Thinking (With Tools)40.706 / 10Thinking (With Tools)+7.50

Claw-style Agent Evaluation

GLM-5.2 领先 1/1
评测项GLM-5.2GLM 5.1分差
PinchBench v286.985 / 45Reported best (effort unspecified)59.9533 / 45Reported best (effort unspecified)+27.03

Coding and Software Engineer

GLM-5.2 领先 1/1
评测项GLM-5.2GLM 5.1分差
SWE-Bench Pro - Public62.1012 / 62Thinking (With Tools)58.4020 / 62Thinking (With Tools)+3.70

Long Context

GLM 5.1 领先 1/1
评测项GLM-5.2GLM 5.1分差
AA-LCR42.30154 / 170Normal (No Tools)53.30134 / 170Normal (No Tools)-11

Writing and Creative Capabilities

GLM-5.2 领先 1/1
评测项GLM-5.2GLM 5.1分差
Creative Writing1,75322 / 106Normal (No Tools)1,58940 / 106Normal (No Tools)+163.60

常识推理

GLM-5.2 领先 1/1
评测项GLM-5.2GLM 5.1分差
SimpleBench58.8036 / 93Normal (No Tools)55.1042 / 93Normal (No Tools)+3.70

科学与综合推理

GLM 5.1 领先 1/1
评测项GLM-5.2GLM 5.1分差
GPQA Diamond71.21310 / 462Normal (No Tools)83.90180 / 462Normal (No Tools)-12.69

规格对比

字段GLM-5.2GLM 5.1
发布机构智谱AI智谱AI
发布时间2026-06-132026-03-27
模型类型推理大模型推理大模型
架构MoE 架构MoE 架构
参数规模7533.3亿7540亿
上下文长度1M200K
最大输出128K125K

API 调用价格

价格优先使用 DataLearner 配置的 API 记录;缺失项不做推测。

价格项GLM-5.2GLM 5.1
文本输入$1.4 / 1M tokens$1.4 / 1M tokens
文本输出$4.4 / 1M tokens$4.4 / 1M tokens
缓存读取$0.26 / 1M tokens$4.4 / 1M tokens
缓存写入暂无公开价格$0.26 / 1M tokens

小结

  • GLM-5.2在以下类目领先:Math and Reasoning (2/2)、AI Agent - Tool Usage (1/1)、Claw-style Agent Evaluation (1/1)、Coding and Software Engineer (1/1)、Writing and Creative Capabilities (1/1)、常识推理 (1/1)
  • GLM 5.1在以下类目领先:Long Context (1/1)、科学与综合推理 (1/1)
  • 胶着类目:General Knowledge

11 个共同 benchmark 上,GLM-5.2 平均高出 16.17 分。

单项差距最大的 benchmark:Creative Writing — GLM-5.2 1,753,GLM 5.1 1,589(分差 +163.60)。

本页正文由结构化模型、价格与 benchmark 数据生成,不使用实时 LLM 撰写。