DataLearner logo

Nemotron 3 Ultra Benchmark Details

Nemotron 3 Ultra currently shows benchmark results led by IF Bench (4 / 282, score 81.70), IMO-AnswerBench (1 / 24, score 92.30), Pinch Bench (2 / 38, score 90). This page also compares it with 3 competitor models, including performance and pricing views when available.

Benchmark Results

Nemotron 3 Ultra

Benchmark Results

Thinking
Tool usage
Internet

General Knowledge

8 evaluations
Benchmark / mode
Score
Rank/total
GPQA
Thinking Mode
87
1 / 17
MMLU-Pro
Thinking Mode
86.80
17 / 176
LiveBench
Standard Mode
51.78
90 / 117
HLE
Thinking Mode
28.40
222 / 565
HLE
Thinking Mode
26.70
236 / 565
HLE
Thinking ModeTools
37.40
154 / 565
CritPt
Thinking Mode
3.10
111 / 201

Other

1 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Thinking Mode
86.70
137 / 463

Coding and Software Engineer

4 evaluations
Benchmark / mode
Score
Rank/total
LiveCodeBench
Thinking Mode
89
15 / 251
SWE-bench Verified
Thinking ModeTools
70.70
58 / 116
SWE-bench Multilingual
Thinking ModeTools
67.70
26 / 30
SciCode
Thinking Mode
44.60
93 / 131

Writing and Creative Capabilities

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
1689.30
27 / 106

Common Sense Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleBench
Thinking Mode
41.70
62 / 93

Agent Level Benchmark

3 evaluations
Benchmark / mode
Score
Rank/total
τ²-Bench - Telecom
Thinking ModeTools
83.30
102 / 264
Terminal Bench Hard
Thinking ModeTools
36.40
67 / 244
τ³-Banking
Thinking ModeTools
22.60
85 / 167

Instruction Following

1 evaluations
Benchmark / mode
Score
Rank/total
IF Bench
Thinking Mode
81.70
4 / 282

AI Agent - Information Search

1 evaluations
Benchmark / mode
Score
Rank/total
BrowseComp
Thinking ModeToolsInternet
44.40
49 / 58

Math and Reasoning

2 evaluations
Benchmark / mode
Score
Rank/total
IMO-AnswerBench
Thinking Mode
88.60
6 / 24
IMO-AnswerBench
Thinking ModeTools
92.30
1 / 24

Long Context

2 evaluations
Benchmark / mode
Score
Rank/total
AA-LCR
Thinking Mode
79.30
54 / 171
LongBench v2
Standard Mode
61.90
5 / 14

Claw-style Agent Evaluation

2 evaluations
Benchmark / mode
Score
Rank/total
Pinch Bench
Thinking ModeTools
90
2 / 38
89.92
3 / 45

AI Agent - Tool Usage

2 evaluations
Benchmark / mode
Score
Rank/total
Terminal-Bench 2.1
Thinking ModeTools
56.40
122 / 196
Terminal-Bench 4.0
Thinking ModeTools
0.50
86 / 91

Productivity Knowledge

4 evaluations
Benchmark / mode
Score
Rank/total
GDPval-AA v2
Thinking ModeTools
1091
84 / 107
AA-Briefcase
Thinking ModeTools
878
71 / 85
Harvey Lab-AA
Thinking ModeTools
81.72
30 / 44
AA-AnalystAgent
Thinking ModeToolsInternet
6.25
29 / 29

Multimodal Understanding

1 evaluations
Benchmark / mode
Score
Rank/total
GDP.pdf
Thinking Mode
5
97 / 119

Competitor Comparison

Benchmark scores for Nemotron 3 Ultra compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

12 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkNemotron 3 UltraCurrentGLM-5.2DeepSeek-V4-ProMiniMax M3
AA Intelligence Index (historical versions)
Historical index (not comparable)
综合评估
38.32Thinking Enabled
--
--
45.40Thinking Enabled
CritPt
Score
综合评估
3.10Thinking Enabled
20.90Thinking Level · High
12.90Thinking Level · High
3.70Thinking Enabled
HLE
Accuracy
综合评估
37.40Thinking Enabled | Tools
54.70Thinking Enabled | Tools
48.20Thinking Level · Extra High | Tools
39.00Thinking Enabled
LiveBench
Accuracy
综合评估
51.78Standard Mode
73.18Standard Mode
71.57Standard Mode
67.26Deep Thinking Mode
MMLU-Pro
Accuracy
综合评估
86.80Thinking Enabled
--
87.50Thinking Level · High
--
GPQA Diamond
Accuracy
科学与综合推理
86.70Thinking Enabled
91.86Thinking Level · High
90.50Thinking Level · High
92.90Thinking Enabled
LiveCodeBench
Pass @K
编程与软件工程
89.00Thinking Enabled
--
93.50Thinking Level · High
--
SciCode
Score
编程与软件工程
44.60Thinking Enabled
51.20Thinking Level · High
50.80Thinking Level · High
45.37Thinking Enabled
SWE-bench Multilingual
Accuracy
编程与软件工程
67.70Thinking Enabled | Tools
--
76.20Thinking Level · Extra High | Tools
--
SWE-bench Verified
Accuracy
编程与软件工程
70.70Thinking Enabled | Tools
--
80.60Thinking Level · Extra High | Tools
--
Creative Writing
Elo、大模型评判两两对战
写作和创作
1689.30Standard Mode
1752.80Standard Mode
1552.10Standard Mode
--
SimpleBench
Score (AVG@5)
常识推理
41.70Thinking Enabled
58.80Standard Mode
50.90Standard Mode
45.80Thinking Enabled
15 additional benchmarks remain in the chart above.

Standard API Pricing: Nemotron 3 Ultra vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier.

These models use different currencies or billing units, so the page falls back to raw price values instead of a shared bar chart.

GLM-5.2
Supplier: 智谱AI
Standard input: $1.4 / 1M tokens
Standard output: $4.4 / 1M tokens
DeepSeek-V4-Pro
Supplier: DeepSeek-AI
Standard input: $0.435 / 1M tokens
Standard output: $0.87 / 1M tokens
MiniMax M3
Supplier: MiniMaxAI
Standard input: ¥2.1 / 1M tokens
Standard output: ¥8.4 / 1M tokens
ModelSupplierStandard inputStandard outputBase price applies to
GLM-5.2
智谱AI$1.4 / 1M tokens$4.4 / 1M tokens
DeepSeek-V4-Pro
DeepSeek-AI$0.435 / 1M tokens$0.87 / 1M tokens
MiniMax M3
MiniMaxAI¥2.1 / 1M tokens¥8.4 / 1M tokens