DataLearner logo

Claude Sonnet 5 Benchmark Details

Claude Sonnet 5 currently shows benchmark results led by HLE (9 / 181, score 57.40), SWE-bench Verified (7 / 114, score 85.20), BrowseComp (7 / 54, score 84.70). 1 source link is attached for reference.

Benchmark Results

Claude Sonnet 5

Benchmark Results

Thinking
Tool usage
Internet

General Knowledge

2 evaluations
Benchmark / mode
Score
Rank/total
HLE
Thinking Level · Extra High
43.20
52 / 181
HLE
Thinking Level · Extra HighTools
57.40
9 / 181

Other

2 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Thinking Level · Max
80.30
117 / 224
GPQA Diamond
Thinking Level · Extra High
90.53
33 / 224

Coding and Software Engineer

4 evaluations
Benchmark / mode
Score
Rank/total
Text Arena (Coding)
Thinking Level · High
1544.16
10 / 35
SWE-bench Verified
Thinking Level · Extra HighTools
85.20
7 / 114
WeirdML v2
Thinking Level · HighTools
68.78
20 / 52
DeepSWE
DeepTools
54
17 / 28

Writing and Creative Capabilities

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
1787.60
18 / 99

Common Sense Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleBench
Standard Mode
57.90
23 / 67

AI Agent - Information Search

1 evaluations
Benchmark / mode
Score
Rank/total
BrowseComp
Thinking EnabledToolsInternet
84.70
7 / 54

AI Agent - Tool Usage

3 evaluations
Benchmark / mode
Score
Rank/total
OSWorld-Verified
Thinking Level · Extra HighTools
81.20
6 / 26
Terminal-Bench 2.1
Thinking Level · HighTools
74.60
27 / 45
Terminal-Bench 2.1
Thinking Level · Extra HighTools
80.40
19 / 45

Math and Reasoning

2 evaluations
Benchmark / mode
Score
Rank/total
FrontierMath v2
Thinking Level · Max
65.61
17 / 34
FrontierMath Tier 4 v2
Thinking Level · Max
29.27
16 / 34

Sources