DataLearner logo

DeepSeek-V4-ProvsGLM 5.1

Across 11 shared benchmarks, GLM 5.1 leads overall: DeepSeek-V4-Pro wins 4, GLM 5.1 wins 7, with 0 ties and an average score difference of -7.27.

DeepSeek-AI
DeepSeek-V4-Pro

DeepSeek-AI · 2026-08-13 · Reasoning model

智谱AI
GLM 5.1

智谱AI · 2026-03-27 · Reasoning model

DeepSeek-V4-Pro4 wins(36%)(64%)7 winsGLM 5.1

Benchmark scores

Grouped by capability, sorted by largest gap within each. 11 shared benchmarks.

Agent Level Benchmark

Even 2/2
BenchmarkDeepSeek-V4-ProGLM 5.1Diff
τ²-Bench - Telecom91.2059 / 264Normal (With Tools)97.1017 / 264Normal (With Tools)-5.90
Terminal Bench Hard36.4067 / 244Normal (With Tools)35.6070 / 244Normal (With Tools)+0.80

General Knowledge

Even 2/2
BenchmarkDeepSeek-V4-ProGLM 5.1Diff
HLE8.20416 / 563Normal (No Tools) · Text only27.90226 / 563Normal (No Tools) · Text only-19.70
LiveBench71.5732 / 117Normal (No Tools)70.1837 / 117Normal (No Tools)+1.39

Claw-style Agent Evaluation

DeepSeek-V4-Pro 1/1
BenchmarkDeepSeek-V4-ProGLM 5.1Diff
PinchBench v261.0732 / 45Reported best (effort unspecified)59.9533 / 45Reported best (effort unspecified)+1.12

Commonsense Reasoning

GLM 5.1 1/1
BenchmarkDeepSeek-V4-ProGLM 5.1Diff
SimpleBench50.9050 / 93Normal (No Tools)55.1042 / 93Normal (No Tools)-4.20

General Evaluation

GLM 5.1 1/1
BenchmarkDeepSeek-V4-ProGLM 5.1Diff
GPQA Diamond72.90301 / 462Normal (No Tools)83.90180 / 462Normal (No Tools)-11

Instruction Following

GLM 5.1 1/1
BenchmarkDeepSeek-V4-ProGLM 5.1Diff
IF Bench45.80176 / 282Normal (No Tools)52148 / 282Normal (No Tools)-6.20

Long Context

GLM 5.1 1/1
BenchmarkDeepSeek-V4-ProGLM 5.1Diff
AA-LCR53135 / 170Normal (No Tools)53.30134 / 170Normal (No Tools)-0.30

Text Embedding

DeepSeek-V4-Pro 1/1
BenchmarkDeepSeek-V4-ProGLM 5.1Diff
Context Arena31.43115 / 126Normal (No Tools)30.29116 / 126Normal (No Tools)+1.14

Writing and Creative Capabilities

GLM 5.1 1/1
BenchmarkDeepSeek-V4-ProGLM 5.1Diff
Creative Writing1,55247 / 106Normal (No Tools)1,58940 / 106Normal (No Tools)-37.10

Specs

FieldDeepSeek-V4-ProGLM 5.1
PublisherDeepSeek-AI智谱AI
Release date2026-08-132026-03-27
Model typeReasoning modelReasoning model
ArchitectureMoEMoE
Parameters1.6T754B
Context length1M200K
Max output384K125K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemDeepSeek-V4-ProGLM 5.1
Text input$0.435 / 1M tokens$1.4 / 1M tokens
Text output$0.87 / 1M tokens$4.4 / 1M tokens
Cache read$0.003625 / 1M tokens$4.4 / 1M tokens
Cache writeNot public$0.26 / 1M tokens

Summary

  • DeepSeek-V4-Proleads in:Claw-style Agent Evaluation (1/1), Text Embedding (1/1)
  • GLM 5.1leads in:Commonsense Reasoning (1/1), General Evaluation (1/1), Instruction Following (1/1), Long Context (1/1), Writing and Creative Capabilities (1/1)
  • Tied in:Agent Level Benchmark, General Knowledge

On average across the 11 shared benchmarks, GLM 5.1 scores 7.27 higher.

Largest single-benchmark gap: Creative Writing — DeepSeek-V4-Pro 1,552 vs GLM 5.1 1,589 (-37.10).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.