DataLearner logo

Llama 4 Maverick Benchmark Details

Llama 4 Maverick currently shows benchmark results led by MBPP (22 / 96, score 77.60), MMLU (40 / 124, score 85.50), MMMU (39 / 74, score 73.40). 1 source link is attached for reference.

Benchmark Results

Llama 4 Maverick

Benchmark Results

Thinking
Tool usage

General Knowledge

6 evaluations
Benchmark / mode
Score
Rank/total
MMLU
Standard Mode
85.50
40 / 124
MMLU Pro
Standard Mode
62.90
118 / 176
HLE
Standard Mode
4.90
486 / 563
HLE
unknown
5.68
465 / 563
ARC-AGI-1
Thinking Enabled
4.38
144 / 147
ARC-AGI-2
Thinking Enabled
0
133 / 136

Coding and Software Engineer

3 evaluations
Benchmark / mode
Score
Rank/total
MBPP
Standard Mode
77.60
22 / 96
LiveCodeBench
Standard Mode
39.70
187 / 250
SciCode
Standard Mode
31.70
122 / 130

Math and Reasoning

3 evaluations
Benchmark / mode
Score
Rank/total
MATH
Standard Mode
61.20
30 / 42
AIME2025
Standard Mode
19.30
199 / 215
FrontierMath
Standard Mode
0.70
55 / 60

Other

1 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Standard Mode
67.10
343 / 462

Multimodal Understanding

3 evaluations
Benchmark / mode
Score
Rank/total
MMMU
unknown
73.40
39 / 74
MMMU-Pro
Standard Mode
62.10
165 / 227
GDP.pdf
Standard Mode
0.80
115 / 118

Common Sense Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleBench
Standard Mode
27.70
76 / 93

Agent Level Benchmark

4 evaluations
Benchmark / mode
Score
Rank/total
τ²-Bench - Telecom
Standard ModeTools
17.80
253 / 264
Aider-Polyglot
Standard Mode
15.60
54 / 59
Terminal Bench Hard
Standard ModeTools
6.80
194 / 244
τ³-Banking
Standard ModeTools
3.70
159 / 164

Instruction Following

1 evaluations
Benchmark / mode
Score
Rank/total
IF Bench
Standard Mode
43
194 / 282

Claw-style Agent Evaluation

1 evaluations
Benchmark / mode
Score
Rank/total
Pinch Bench
Thinking EnabledTools
46.10
37 / 38

AI Agent - Tool Usage

1 evaluations
Benchmark / mode
Score
Rank/total
Terminal-Bench 2.1
Standard ModeTools
7.90
177 / 192

Sources