DataLearner logo

Llama3.1-8B-Instruct Benchmark Details

Llama3.1-8B-Instruct currently shows benchmark results led by MBPP (18 / 70, score 69.40), GSM8K (21 / 70, score 82.40), HumanEval (37 / 101, score 66.50).

Benchmark Results

Llama3.1-8B-Instruct

Benchmark Results

Thinking

General Knowledge

2 evaluations
Benchmark / mode
Score
Rank/total
MMLU
Standard Mode
68.10
78 / 124
MMLU Pro
Standard Mode
44
127 / 133

Math and Reasoning

2 evaluations
Benchmark / mode
Score
Rank/total
GSM8K
Standard Mode
82.40
21 / 70
MATH
Standard Mode
47.60
35 / 42

Coding and Software Engineer

2 evaluations
Benchmark / mode
Score
Rank/total
MBPP
Standard Mode
69.40
18 / 70
HumanEval
Standard Mode
66.50
37 / 101

Other

1 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Standard Mode
26.30
263 / 270

Writing and Creative Capabilities

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
710.70
94 / 99