DataLearner logo

DeepSeek-V4.1-FlashvsGLM-5.3

Across 12 shared benchmarks, DeepSeek-V4.1-Flash leads overall: DeepSeek-V4.1-Flash wins 11, GLM-5.3 wins 1, with 0 ties and an average score difference of +2.57.

DeepSeek-AI
DeepSeek-V4.1-Flash

DeepSeek-AI · 2026-09-10 · Multimodal model

智谱AI
GLM-5.3

智谱AI · 2026-08-14 · Reasoning model

DeepSeek-V4.1-Flash11 wins(92%)(8%)1 winGLM-5.3

Benchmark scores

Grouped by capability, sorted by largest gap within each. 12 shared benchmarks.

AI Agent - Tool Usage

DeepSeek-V4.1-Flash 3/4
BenchmarkDeepSeek-V4.1-FlashGLM-5.3Diff
Terminal-Bench 4.031.207 / 16Max (With Tools)37.906 / 16Max (With Tools)-6.70
CyberGym88.101 / 8Max (With Tools)84.502 / 8Max (With Tools)+3.60
Terminal-Bench 2.190.601 / 53Max (With Tools)88.208 / 53Max (With Tools)+2.40
Terminal-Bench 3.0304 / 11Max (With Tools)28.306 / 11Max (With Tools)+1.70

Coding and Software Engineer

DeepSeek-V4.1-Flash 3/3
BenchmarkDeepSeek-V4.1-FlashGLM-5.3Diff
NL2Repo-Bench65.402 / 16Max (With Tools)586 / 16Max (With Tools)+7.40
DeepSWE74.202 / 38Max (With Tools)66.9015 / 38Max (With Tools)+7.30
Program Bench20.308 / 11Max (With Tools)199 / 11Max (With Tools)+1.30

Agent Capability

DeepSeek-V4.1-Flash 1/1
BenchmarkDeepSeek-V4.1-FlashGLM-5.3Diff
ExploitGym (budget unspecified)15.302 / 4Max (With Tools)153 / 4Max (With Tools)+0.30

Agent Level Benchmark

DeepSeek-V4.1-Flash 1/1
BenchmarkDeepSeek-V4.1-FlashGLM-5.3Diff
Agents' Last Exam31.806 / 19Max (With Tools)28.508 / 19Max (With Tools)+3.30

General Evaluation

DeepSeek-V4.1-Flash 1/1
BenchmarkDeepSeek-V4.1-FlashGLM-5.3Diff
GPQA Diamond90.9039 / 274Max (No Tools)88.1070 / 274Max (No Tools)+2.80

General Knowledge

DeepSeek-V4.1-Flash 1/1
BenchmarkDeepSeek-V4.1-FlashGLM-5.3Diff
HLE63.903 / 197Max (With Tools)62.505 / 197Max (With Tools)+1.40

Productivity Knowledge

DeepSeek-V4.1-Flash 1/1
BenchmarkDeepSeek-V4.1-FlashGLM-5.3Diff
AutomationBench54.801 / 17Max (With Tools)48.805 / 17Max (With Tools)+6

Specs

FieldDeepSeek-V4.1-FlashGLM-5.3
PublisherDeepSeek-AI智谱AI
Release date2026-09-102026-08-14
Model typeMultimodal modelReasoning model
ArchitectureMoEMoE
Parameters552B744B
Context length1M1M
Max output384K128K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemDeepSeek-V4.1-FlashGLM-5.3
Text input¥1 / 1M tokens$1.4 / 1M tokens
Text output¥4 / 1M tokens$4.4 / 1M tokens
Cache read¥0.02 / 1M tokens$0.26 / 1M tokens

Summary

  • DeepSeek-V4.1-Flashleads in:AI Agent - Tool Usage (3/4), Coding and Software Engineer (3/3), Agent Capability (1/1), Agent Level Benchmark (1/1), General Evaluation (1/1), General Knowledge (1/1), Productivity Knowledge (1/1)

On average across the 12 shared benchmarks, DeepSeek-V4.1-Flash scores 2.57 higher.

Largest single-benchmark gap: NL2Repo-Bench — DeepSeek-V4.1-Flash 65.40 vs GLM-5.3 58 (+7.40).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.