DataLearner logo

DeepSeek-V4-ProvsNemotron 3 Ultra

Across 9 shared benchmarks, DeepSeek-V4-Pro leads overall: DeepSeek-V4-Pro wins 5, Nemotron 3 Ultra wins 4, with 0 ties and an average score difference of -2.83.

DeepSeek-AI
DeepSeek-V4-Pro

DeepSeek-AI · 2026-08-13 · Reasoning model

NVIDIA
Nemotron 3 Ultra

NVIDIA · 2026-06-04 · Reasoning model

DeepSeek-V4-Pro5 wins(56%)(44%)4 winsNemotron 3 Ultra

Benchmark scores

Grouped by capability, sorted by largest gap within each. 9 shared benchmarks.

Coding and Software Engineer

DeepSeek-V4-Pro 2/3
BenchmarkDeepSeek-V4-ProNemotron 3 UltraDiff
LiveCodeBench56.8079 / 126Normal (No Tools)899 / 126Thinking (No Tools)-32.20
SWE-bench Verified73.6046 / 114Normal (With Tools)70.7058 / 114Thinking (With Tools)+2.90
SWE-bench Multilingual69.8019 / 25Normal (With Tools)67.7022 / 25Thinking (With Tools)+2.10

General Knowledge

Nemotron 3 Ultra 2/3
BenchmarkDeepSeek-V4-ProNemotron 3 UltraDiff
HLE7.70165 / 181Normal (No Tools)37.4073 / 181Thinking (With Tools)-29.70
LiveBench73.5823 / 115Normal (No Tools)51.7888 / 115Normal (No Tools)+21.80
MMLU Pro82.9049 / 133Normal (No Tools)86.8016 / 133Thinking (No Tools)-3.90

AI Agent - Information Search

DeepSeek-V4-Pro 1/1
BenchmarkDeepSeek-V4-ProNemotron 3 UltraDiff
BrowseComp83.4013 / 54极高强度思考(工具)44.4046 / 54Thinking (With Tools + Internet)+39

AI Agent - Tool Usage

DeepSeek-V4-Pro 1/1
BenchmarkDeepSeek-V4-ProNemotron 3 UltraDiff
Terminal-Bench 2.187.906 / 44极高强度思考(工具)56.4041 / 44Thinking (With Tools)+31.50

Math and Reasoning

Nemotron 3 Ultra 1/1
BenchmarkDeepSeek-V4-ProNemotron 3 UltraDiff
IMO-AnswerBench35.3023 / 23Normal (No Tools)92.301 / 23Thinking (With Tools)-57

Specs

FieldDeepSeek-V4-ProNemotron 3 Ultra
PublisherDeepSeek-AINVIDIA
Release date2026-08-132026-06-04
Model typeReasoning modelReasoning model
ArchitectureMoEMoE
Parameters1.6T550B
Context length1M1M
Max output384KNot available

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemDeepSeek-V4-ProNemotron 3 Ultra
Text input$0.435 / 1M tokensNot public
Text output$0.87 / 1M tokensNot public
Cache read$0.003625 / 1M tokensNot public

One or both models have incomplete public pricing.

Summary

  • DeepSeek-V4-Proleads in:Coding and Software Engineer (2/3), AI Agent - Information Search (1/1), AI Agent - Tool Usage (1/1)
  • Nemotron 3 Ultraleads in:General Knowledge (2/3), Math and Reasoning (1/1)

On average across the 9 shared benchmarks, Nemotron 3 Ultra scores 2.83 higher.

Largest single-benchmark gap: IMO-AnswerBench — DeepSeek-V4-Pro 35.30 vs Nemotron 3 Ultra 92.30 (-57).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.