DataLearner 标志

GLM-5vsKimi K2.5

在 17 个共同 benchmark 中,GLM-5 整体领先:GLM-5 领先 10 项,Kimi K2.5 领先 7 项,持平 0 项,平均分差 +1.75。

智谱AI
GLM-5

智谱AI · 2026-02-11 · 聊天大模型

Moonshot AI
Kimi K2.5

Moonshot AI · 2026-01-27 · 多模态大模型

GLM-510 (59%)(41%)7 Kimi K2.5

评测分数

按能力类目分组,每组内按分差大小排列;共 17 项。

Agent Level Benchmark

GLM-5 领先 2/3
评测项GLM-5Kimi K2.5分差
Terminal Bench Hard39.4053 / 244Normal (With Tools)18.90150 / 244Normal (With Tools)+20.50
τ²-Bench - Telecom97.4016 / 264Normal (With Tools)81.30108 / 264Normal (With Tools)+16.10
τ³-Banking9.79130 / 164Thinking (With Tools)14.20113 / 164Thinking (With Tools)-4.41

Math and Reasoning

GLM-5 领先 2/3
评测项GLM-5Kimi K2.5分差
FrontierMath - Tier 42.1056 / 80Normal (No Tools)4.2040 / 80Normal (No Tools)-2.10
IMO-AnswerBench82.5017 / 24Thinking (No Tools)81.8018 / 24Thinking (No Tools)+0.70
AIME 202692.7018 / 29Thinking (No Tools)92.5021 / 29Thinking (No Tools)+0.20

Claw-style Agent Evaluation

GLM-5 领先 2/2
评测项GLM-5Kimi K2.5分差
Claw Bench91.705 / 29Thinking (With Tools)81.7018 / 29Thinking (With Tools)+10
Pinch Bench86.4013 / 38Thinking (With Tools)84.8018 / 38Thinking (With Tools)+1.60

General Knowledge

Kimi K2.5 领先 2/2
评测项GLM-5Kimi K2.5分差
ARC-AGI-144.67111 / 147Thinking (No Tools)65.3390 / 147Thinking (No Tools)-20.66
HLE7.60427 / 563Normal (No Tools) · Text only13.20357 / 563Normal (No Tools) · Text only-5.60

Long Context

Kimi K2.5 领先 2/2
评测项GLM-5Kimi K2.5分差
AA-LCR43.70151 / 170Normal (No Tools)67.30113 / 170Normal (No Tools)-23.60
LongBench v260.807 / 14Normal (No Tools)616 / 14Normal (No Tools)-0.20

AI Agent - Tool Usage

GLM-5 领先 1/1
评测项GLM-5Kimi K2.5分差
Terminal Bench 2.061.1018 / 48Thinking (With Tools)50.8035 / 48Thinking (With Tools)+10.30

Instruction Following

GLM-5 领先 1/1
评测项GLM-5Kimi K2.5分差
IF Bench55.20132 / 282Normal (No Tools)43.70188 / 282Normal (No Tools)+11.50

Productivity Knowledge

GLM-5 领先 1/1
评测项GLM-5Kimi K2.5分差
GDPval-AA468 / 15Thinking (No Tools)409 / 15Thinking (No Tools)+6

Writing and Creative Capabilities

GLM-5 领先 1/1
评测项GLM-5Kimi K2.5分差
Creative Writing1,59839 / 106Normal (No Tools)1,57643 / 106Normal (No Tools)+21.70

科学与综合推理

Kimi K2.5 领先 1/1
评测项GLM-5Kimi K2.5分差
GPQA Diamond66.60349 / 462Normal (No Tools)78.90248 / 462Normal (No Tools)-12.30

规格对比

字段GLM-5Kimi K2.5
发布机构智谱AIMoonshot AI
发布时间2026-02-112026-01-27
模型类型聊天大模型多模态大模型
架构MoE 架构MoE 架构
参数规模7440亿1万亿
上下文长度200K256K
最大输出128K16K

API 调用价格

价格优先使用 DataLearner 配置的 API 记录;缺失项不做推测。

价格项GLM-5Kimi K2.5
文本输入$1 / 1M tokens$0.6 / 1M tokens
文本输出$3.2 / 1M tokens$3 / 1M tokens
缓存读取暂无公开价格$0.1 / 1M tokens
缓存写入$0.2 / 1M tokens暂无公开价格

小结

  • GLM-5在以下类目领先:Agent Level Benchmark (2/3)、Math and Reasoning (2/3)、Claw-style Agent Evaluation (2/2)、AI Agent - Tool Usage (1/1)、Instruction Following (1/1)、Productivity Knowledge (1/1)、Writing and Creative Capabilities (1/1)
  • Kimi K2.5在以下类目领先:General Knowledge (2/2)、Long Context (2/2)、科学与综合推理 (1/1)

17 个共同 benchmark 上,GLM-5 平均高出 1.75 分。

单项差距最大的 benchmark:AA-LCR — GLM-5 43.70,Kimi K2.5 67.30(分差 -23.60)。

本页正文由结构化模型、价格与 benchmark 数据生成,不使用实时 LLM 撰写。