DataLearner logo

DeepSeek-V4-ProvsQwen3.8-Max

Across 11 shared benchmarks, Qwen3.8-Max leads overall: DeepSeek-V4-Pro wins 5, Qwen3.8-Max wins 6, with 0 ties and an average score difference of -12.67.

DeepSeek-AI
DeepSeek-V4-Pro

DeepSeek-AI · 2026-08-13 · Reasoning model

阿里巴巴
Qwen3.8-Max

阿里巴巴 · 2026-08-03 · Reasoning model

DeepSeek-V4-Pro5 wins(45%)(55%)6 winsQwen3.8-Max

Benchmark scores

Grouped by capability, sorted by largest gap within each. 11 shared benchmarks.

AI Agent - Tool Usage

DeepSeek-V4-Pro 3/3
BenchmarkDeepSeek-V4-ProQwen3.8-MaxDiff
AutomationBench31.802 / 8极高强度思考(工具)27.305 / 8极高强度思考(工具)+4.50
Toolathlon-Verified74.102 / 5极高强度思考(工具)72.504 / 5极高强度思考(工具)+1.60
Terminal-Bench 2.187.906 / 44极高强度思考(工具)86.608 / 44极高强度思考(工具)+1.30

Coding and Software Engineer

DeepSeek-V4-Pro 2/3
BenchmarkDeepSeek-V4-ProQwen3.8-MaxDiff
SWE-Bench Pro - Public52.1040 / 57Normal (With Tools)67.705 / 57极高强度思考(工具)-15.60
DeepSWE62.7011 / 27极高强度思考(工具)56.6014 / 27极高强度思考(工具)+6.10
NL2Repo-Bench61.501 / 8极高强度思考(工具)55.904 / 8极高强度思考(工具)+5.60

Math and Reasoning

Qwen3.8-Max 2/2
BenchmarkDeepSeek-V4-ProQwen3.8-MaxDiff
FrontierMath Tier 4 v22.4430 / 34最高(无工具)46.3411 / 34极高强度思考(无工具)-43.90
FrontierMath v245.2626 / 34最高(无工具)74.7411 / 34极高强度思考(无工具)-29.47

Agent Level Benchmark

Qwen3.8-Max 1/1
BenchmarkDeepSeek-V4-ProQwen3.8-MaxDiff
Agents' Last Exam25.709 / 11极高强度思考(工具)277 / 11极高强度思考(工具)-1.30

General Evaluation

Qwen3.8-Max 1/1
BenchmarkDeepSeek-V4-ProQwen3.8-MaxDiff
GPQA Diamond72.90147 / 226Normal (No Tools)92.6021 / 226极高强度思考(无工具)-19.70

General Knowledge

Qwen3.8-Max 1/1
BenchmarkDeepSeek-V4-ProQwen3.8-MaxDiff
HLE7.70165 / 181Normal (No Tools)56.2013 / 181极高强度思考(工具)-48.50

Specs

FieldDeepSeek-V4-ProQwen3.8-Max
PublisherDeepSeek-AI阿里巴巴
Release date2026-08-132026-08-03
Model typeReasoning modelReasoning model
ArchitectureMoEMoE
Parameters1.6T2.4T
Context length1M1M
Max output384K128K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemDeepSeek-V4-ProQwen3.8-Max
Text input$0.435 / 1M tokens¥12 / 1M tokens
Text output$0.87 / 1M tokens¥36 / 1M tokens
Cache read$0.003625 / 1M tokens¥1.5 / 1M tokens

Summary

  • DeepSeek-V4-Proleads in:AI Agent - Tool Usage (3/3), Coding and Software Engineer (2/3)
  • Qwen3.8-Maxleads in:Math and Reasoning (2/2), Agent Level Benchmark (1/1), General Evaluation (1/1), General Knowledge (1/1)

On average across the 11 shared benchmarks, Qwen3.8-Max scores 12.67 higher.

Largest single-benchmark gap: HLE — DeepSeek-V4-Pro 7.70 vs Qwen3.8-Max 56.20 (-48.50).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.