DeepSeek-V4-ProvsQwen3.8-Max
Across 11 shared benchmarks, Qwen3.8-Max leads overall: DeepSeek-V4-Pro wins 5, Qwen3.8-Max wins 6, with 0 ties and an average score difference of -12.67.
DeepSeek-V4-Pro
DeepSeek-AI · 2026-08-13 · Reasoning model
Qwen3.8-Max
阿里巴巴 · 2026-08-03 · Reasoning model
DeepSeek-V4-Pro5 wins(45%)(55%)6 winsQwen3.8-Max
Benchmark scores
Grouped by capability, sorted by largest gap within each. 11 shared benchmarks.
AI Agent - Tool Usage
DeepSeek-V4-Pro 3/3| Benchmark | DeepSeek-V4-Pro | Qwen3.8-Max | Diff |
|---|---|---|---|
| AutomationBench | 31.802 / 8极高强度思考(工具) | 27.305 / 8极高强度思考(工具) | +4.50 |
| Toolathlon-Verified | 74.102 / 5极高强度思考(工具) | 72.504 / 5极高强度思考(工具) | +1.60 |
| Terminal-Bench 2.1 | 87.906 / 44极高强度思考(工具) | 86.608 / 44极高强度思考(工具) | +1.30 |
Coding and Software Engineer
DeepSeek-V4-Pro 2/3| Benchmark | DeepSeek-V4-Pro | Qwen3.8-Max | Diff |
|---|---|---|---|
| SWE-Bench Pro - Public | 52.1040 / 57Normal (With Tools) | 67.705 / 57极高强度思考(工具) | -15.60 |
| DeepSWE | 62.7011 / 27极高强度思考(工具) | 56.6014 / 27极高强度思考(工具) | +6.10 |
| NL2Repo-Bench | 61.501 / 8极高强度思考(工具) | 55.904 / 8极高强度思考(工具) | +5.60 |
Math and Reasoning
Qwen3.8-Max 2/2| Benchmark | DeepSeek-V4-Pro | Qwen3.8-Max | Diff |
|---|---|---|---|
| FrontierMath Tier 4 v2 | 2.4430 / 34最高(无工具) | 46.3411 / 34极高强度思考(无工具) | -43.90 |
| FrontierMath v2 | 45.2626 / 34最高(无工具) | 74.7411 / 34极高强度思考(无工具) | -29.47 |
Agent Level Benchmark
Qwen3.8-Max 1/1| Benchmark | DeepSeek-V4-Pro | Qwen3.8-Max | Diff |
|---|---|---|---|
| Agents' Last Exam | 25.709 / 11极高强度思考(工具) | 277 / 11极高强度思考(工具) | -1.30 |
General Evaluation
Qwen3.8-Max 1/1| Benchmark | DeepSeek-V4-Pro | Qwen3.8-Max | Diff |
|---|---|---|---|
| GPQA Diamond | 72.90147 / 226Normal (No Tools) | 92.6021 / 226极高强度思考(无工具) | -19.70 |
General Knowledge
Qwen3.8-Max 1/1| Benchmark | DeepSeek-V4-Pro | Qwen3.8-Max | Diff |
|---|---|---|---|
| HLE | 7.70165 / 181Normal (No Tools) | 56.2013 / 181极高强度思考(工具) | -48.50 |
Specs
| Field | DeepSeek-V4-Pro | Qwen3.8-Max |
|---|---|---|
| Publisher | DeepSeek-AI | 阿里巴巴 |
| Release date | 2026-08-13 | 2026-08-03 |
| Model type | Reasoning model | Reasoning model |
| Architecture | MoE | MoE |
| Parameters | 1.6T | 2.4T |
| Context length | 1M | 1M |
| Max output | 384K | 128K |
API pricing
Prices use DataLearner records when available; missing fields are not inferred.
| Item | DeepSeek-V4-Pro | Qwen3.8-Max |
|---|---|---|
| Text input | $0.435 / 1M tokens | ¥12 / 1M tokens |
| Text output | $0.87 / 1M tokens | ¥36 / 1M tokens |
| Cache read | $0.003625 / 1M tokens | ¥1.5 / 1M tokens |
Summary
- DeepSeek-V4-Proleads in:AI Agent - Tool Usage (3/3), Coding and Software Engineer (2/3)
- Qwen3.8-Maxleads in:Math and Reasoning (2/2), Agent Level Benchmark (1/1), General Evaluation (1/1), General Knowledge (1/1)
On average across the 11 shared benchmarks, Qwen3.8-Max scores 12.67 higher.
Largest single-benchmark gap: HLE — DeepSeek-V4-Pro 7.70 vs Qwen3.8-Max 56.20 (-48.50).
Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.