DataLearner logo

DeepSeek-V3-0324 Benchmark Details

DeepSeek-V3-0324 currently shows benchmark results led by GSM8K (3 / 70, score 96.30), MMLU (29 / 124, score 86.50), GPQA (4 / 17, score 68.40).

Benchmark Results

DeepSeek-V3-0324

Benchmark Results

Thinking
Tool usage

General Knowledge

5 evaluations
Benchmark / mode
Score
Rank/total
MMLU
Standard Mode
86.50
29 / 124
MMLU Pro
Standard Mode
81.20
56 / 134
GPQA
Standard Mode
68.40
4 / 17
ARC-AGI-1
Standard Mode
9
86 / 92
HLE
Standard Mode
5.20
190 / 197

Math and Reasoning

7 evaluations
Benchmark / mode
Score
Rank/total
GSM8K
Standard Mode
96.30
3 / 70
MATH-500
Standard Mode
94
29 / 45
AIME 2024
Standard Mode
59.40
43 / 62
AIME2025
Standard Mode
47.70
89 / 107
IMO-ProofBench
Standard Mode
4.30
15 / 16
IMO 2024
Standard Mode
1.70
9 / 10
IMO 2025
Standard Mode
1.70
9 / 9

Reading Comprehension

1 evaluations
Benchmark / mode
Score
Rank/total
DROP
Standard Mode
89.70
3 / 9

Other

1 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Standard Mode
68.40
205 / 274

Common Sense

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleQA
Standard Mode
27.20
28 / 47

Coding and Software Engineer

2 evaluations
Benchmark / mode
Score
Rank/total
LiveCodeBench
Standard Mode
49.20
100 / 128
SWE-bench Verified
Standard Mode
38.80
107 / 116

Writing and Creative Capabilities

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
1470.20
55 / 106

AI Agent - Tool Usage

1 evaluations
Benchmark / mode
Score
Rank/total
Terminal-Bench
Standard Mode
13.30
34 / 35

Common Sense Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleBench
Standard Mode
27.20
79 / 94

Agent Level Benchmark

2 evaluations
Benchmark / mode
Score
Rank/total
Aider-Polyglot
Standard Mode
55.10
27 / 59
τ²-Bench
Standard ModeTools
38.80
39 / 44