DataLearner logo

DeepSeek-V3 Benchmark Details

DeepSeek-V3 currently shows benchmark results led by HumanEval (18 / 140, score 89), BBH (3 / 21, score 92.30), MMLU (18 / 124, score 88.50).

Benchmark Results

DeepSeek-V3

Benchmark Results

Thinking

General Knowledge

4 evaluations
Benchmark / mode
Score
Rank/total
BBH
Standard Mode
92.30
3 / 21
MMLU
Standard Mode
88.50
18 / 124
MMLU Pro
Standard Mode
75.90
84 / 134
GPQA
Standard Mode
59.10
8 / 17

Coding and Software Engineer

2 evaluations
Benchmark / mode
Score
Rank/total
HumanEval
Standard Mode
89
18 / 140
LiveCodeBench
Standard Mode
34.60
115 / 128

Math and Reasoning

5 evaluations
Benchmark / mode
Score
Rank/total
MATH
Standard Mode
87.80
7 / 42
MATH-500
Standard Mode
87.80
40 / 45
AIME 2024
Standard Mode
39
52 / 62
IMO-ProofBench Advanced
Thinking Enabled
4.30
16 / 19
FrontierMath
Standard Mode
1.70
49 / 60

Other

1 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Standard Mode
59.10
231 / 274

Common Sense

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleQA
Standard Mode
24.90
31 / 47

Common Sense Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleBench
Standard Mode
18.90
90 / 94

Agent Level Benchmark

1 evaluations
Benchmark / mode
Score
Rank/total
Aider-Polyglot
Standard Mode
48.40
34 / 59