DataLearner logo

Claude 3.5 Sonnet Benchmark Details

Claude 3.5 Sonnet currently shows benchmark results led by HumanEval (9 / 140, score 92), MMLU (19 / 124, score 88.30), MATH (18 / 42, score 71.10).

Benchmark Results

Claude 3.5 Sonnet

Benchmark Results

Thinking

General Knowledge

5 evaluations
Benchmark / mode
Score
Rank/total
MMLU
Standard Mode
88.30
19 / 124
MMLU-Pro
Standard Mode
77.64
79 / 176
MMLU-Pro
unknown
77.64
79 / 176
HLE
Standard Mode
3.70
542 / 565
HLE
unknown
4.08
527 / 565

Coding and Software Engineer

2 evaluations
Benchmark / mode
Score
Rank/total
HumanEval
Standard Mode
92
9 / 140
LiveCodeBench
Standard Mode
38.10
193 / 251

Math and Reasoning

3 evaluations
Benchmark / mode
Score
Rank/total
MATH
Standard Mode
71.10
18 / 42
FrontierMath
Standard Mode
1
52 / 60
0
72 / 80

Other

1 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Standard Mode
59.40
382 / 463

Writing and Creative Capabilities

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
1448.90
56 / 106

Multimodal Understanding

2 evaluations
Benchmark / mode
Score
Rank/total
MMMU
unknown
68.30
49 / 74
MMMU-Pro
unknown
51.50
198 / 229

Common Sense Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleBench
Standard Mode
27.50
77 / 93