DataLearner logo

Llama3.1-405B Instruct Benchmark Details

Llama3.1-405B Instruct currently shows benchmark results led by MBPP (1 / 70, score 88.60), HumanEval (9 / 101, score 89), MMLU (17 / 124, score 88.60).

Benchmark Results

Llama3.1-405B Instruct

Benchmark Results

Thinking

General Knowledge

3 evaluations
Benchmark / mode
Score
Rank/total
BBH
Standard Mode
89.20
6 / 21
MMLU
Standard Mode
88.60
17 / 124
MMLU Pro
Standard Mode
73.40
88 / 133

Math and Reasoning

2 evaluations
Benchmark / mode
Score
Rank/total
MATH
Standard Mode
73.90
17 / 42
GSM8K
Standard Mode
0
68 / 70

Coding and Software Engineer

3 evaluations
Benchmark / mode
Score
Rank/total
HumanEval
Standard Mode
89
9 / 101
MBPP
Standard Mode
88.60
1 / 70
LiveCodeBench
Standard Mode
30.20
122 / 127

Other

1 evaluations
Benchmark / mode
Score
Rank/total
TruthfulQA
Standard Mode
0
2 / 4

Other

1 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Standard Mode
49
244 / 270

Common Sense

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleQA
Standard Mode
17.10
36 / 47

Writing and Creative Capabilities

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
867.50
86 / 99