DataLearner logo

DeepSeek-V4-ProvsQwen3.7-Max-Preview

Across 9 shared benchmarks, Qwen3.7-Max-Preview leads overall: DeepSeek-V4-Pro wins 0, Qwen3.7-Max-Preview wins 9, with 0 ties and an average score difference of -21.77.

DeepSeek-AI
DeepSeek-V4-Pro

DeepSeek-AI · 2026-08-13 · Reasoning model

阿里巴巴
Qwen3.7-Max-Preview

阿里巴巴 · 2026-05-20 · Reasoning model

DeepSeek-V4-Pro0 wins(0%)(100%)9 winsQwen3.7-Max-Preview

Benchmark scores

Grouped by capability, sorted by largest gap within each. 9 shared benchmarks.

Coding and Software Engineer

Qwen3.7-Max-Preview 4/4
BenchmarkDeepSeek-V4-ProQwen3.7-Max-PreviewDiff
LiveCodeBench56.8076 / 123Normal (No Tools)91.604 / 123最高(无工具)-34.80
SWE-bench Multilingual69.8018 / 23Normal (With Tools)78.304 / 23Thinking (With Tools)-8.50
SWE-Bench Pro - Public52.1038 / 55Normal (With Tools)60.6010 / 55Thinking (With Tools)-8.50
SWE-bench Verified73.6046 / 113Normal (With Tools)80.4013 / 113Thinking (With Tools)-6.80

General Knowledge

Qwen3.7-Max-Preview 2/2
BenchmarkDeepSeek-V4-ProQwen3.7-Max-PreviewDiff
HLE7.70159 / 175Normal (No Tools)53.5016 / 175Thinking (With Tools)-45.80
MMLU Pro82.9048 / 132Normal (No Tools)89.604 / 132最高(无工具)-6.70

AI Agent - Tool Usage

Qwen3.7-Max-Preview 1/1
BenchmarkDeepSeek-V4-ProQwen3.7-Max-PreviewDiff
Terminal Bench 2.059.1022 / 47Normal (With Tools)69.705 / 47Thinking (With Tools)-10.60

General Evaluation

Qwen3.7-Max-Preview 1/1
BenchmarkDeepSeek-V4-ProQwen3.7-Max-PreviewDiff
GPQA Diamond72.90146 / 225Normal (No Tools)92.4022 / 225最高(无工具)-19.50

Math and Reasoning

Qwen3.7-Max-Preview 1/1
BenchmarkDeepSeek-V4-ProQwen3.7-Max-PreviewDiff
IMO-AnswerBench35.3021 / 21Normal (No Tools)902 / 21最高(无工具)-54.70

Specs

FieldDeepSeek-V4-ProQwen3.7-Max-Preview
PublisherDeepSeek-AI阿里巴巴
Release date2026-08-132026-05-20
Model typeReasoning modelReasoning model
ArchitectureMoEDense
Parameters1.6TNot available
Context length1M1M
Max output384K64K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemDeepSeek-V4-ProQwen3.7-Max-Preview
Text input$0.435 / 1M tokens$2.5 / 1M tokens
Text output$0.87 / 1M tokens$7.5 / 1M tokens
Cache read$0.003625 / 1M tokens$0.25 / 1M tokens
Cache writeNot public$3.125 / 1M tokens

Summary

  • Qwen3.7-Max-Previewleads in:Coding and Software Engineer (4/4), General Knowledge (2/2), AI Agent - Tool Usage (1/1), General Evaluation (1/1), Math and Reasoning (1/1)

On average across the 9 shared benchmarks, Qwen3.7-Max-Preview scores 21.77 higher.

Largest single-benchmark gap: IMO-AnswerBench — DeepSeek-V4-Pro 35.30 vs Qwen3.7-Max-Preview 90 (-54.70).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.