DataLearner logo

Claude Sonnet 4 Benchmark Details

Claude Sonnet 4 currently shows benchmark results led by SWE-bench Verified (14 / 113, score 80.20), Terminal-Bench (10 / 35, score 41.30), MMLU Pro (38 / 132, score 84). 1 source link is attached for reference.

Benchmark Results

Claude Sonnet 4

Benchmark Results

Thinking
Tool usage

General Knowledge

12 evaluations
Benchmark / mode
Score
Rank/total
84
38 / 132
83.80
62 / 188
75.40
98 / 188
68
129 / 188
LiveBench
Standard Mode
50.98
89 / 115
61.27
65 / 115
40
49 / 68
23.80
56 / 68
9.60
150 / 173
5.52
164 / 173
5.90
46 / 62
1.30
55 / 62

Coding and Software Engineer

6 evaluations
Benchmark / mode
Score
Rank/total
CodeClash
Standard ModeTools
1223
4 / 8
80.20
14 / 113
72.70
52 / 113
66
59 / 123
48.50
96 / 123

Math and Reasoning

12 evaluations
Benchmark / mode
Score
Rank/total
85
51 / 107
70.50
72 / 107
38
96 / 107
43.40
50 / 62
27.10
8 / 16
9.70
5 / 10
5.20
8 / 10
4.10
41 / 60
4
5 / 9
3.30
6 / 9
0
72 / 80

Writing and Creative Capabilities

1 evaluations
Benchmark / mode
Score
Rank/total
83.05
14 / 23

AI Agent - Tool Usage

4 evaluations
Benchmark / mode
Score
Rank/total
42.20
23 / 25
41.30
10 / 35
35.50
18 / 35

Multimodal Understanding

1 evaluations
Benchmark / mode
Score
Rank/total
76.50
17 / 29

Common Sense Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
Simple Bench
Thinking Enabled
45.50
34 / 63

Agent Level Benchmark

4 evaluations
Benchmark / mode
Score
Rank/total
Aider-Polyglot
Standard Mode
56.40
26 / 59
61.30
20 / 59
52
34 / 43

Instruction Following

1 evaluations
Benchmark / mode
Score
Rank/total
55
24 / 31

Productivity Knowledge

1 evaluations
Benchmark / mode
Score
Rank/total
33
19 / 21

Long Context

1 evaluations
Benchmark / mode
Score
Rank/total
65
13 / 16

Claw-style Agent Evaluation

2 evaluations
Benchmark / mode
Score
Rank/total
Pinch Bench
Thinking EnabledTools
80.50
22 / 37
Claw Bench
Thinking EnabledTools
77.80
23 / 29

Sources