DataLearner logo

Llama3.1-70B-Instruct Benchmark Details

Llama3.1-70B-Instruct currently shows benchmark results led by MBPP (9 / 96, score 86), MMLU (35 / 124, score 86), HumanEval (41 / 140, score 80.50).

Benchmark Results

Llama3.1-70B-Instruct

Benchmark Results

Thinking
Tool usage

General Knowledge

3 evaluations
Benchmark / mode
Score
Rank/total
MMLU
Standard Mode
86
35 / 124
MMLU-Pro
Standard Mode
66.40
111 / 176
MMLU-Pro
unknown
62.84
119 / 176

Coding and Software Engineer

4 evaluations
Benchmark / mode
Score
Rank/total
MBPP
Standard Mode
86
9 / 96
HumanEval
Standard Mode
80.50
41 / 140
LiveCodeBench
Standard Mode
33.30
207 / 251
WeirdML v2
Standard ModeTools
8.97
52 / 52

Math and Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
MATH
Standard Mode
67.80
26 / 42

Other

1 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Standard Mode
48
416 / 463

Writing and Creative Capabilities

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
781.90
96 / 106

Agent Level Benchmark

1 evaluations
Benchmark / mode
Score
Rank/total
BALROG
Standard ModeTools
27.90
11 / 12

Claw-style Agent Evaluation

1 evaluations
Benchmark / mode
Score
Rank/total
10.70
44 / 45
Llama3.1-70B-Instruct Benchmark Results & Rankings | DataLearnerAI