DataLearner logo

Claude Sonnet 3.7 Benchmark Details

Claude Sonnet 3.7 currently shows benchmark results led by Aider-Polyglot (18 / 59, score 64.90), SimpleBench (35 / 67, score 46.40), SWE-bench Verified (61 / 114, score 70.30). 1 source link is attached for reference.

Benchmark Results

Claude Sonnet 3.7

Benchmark Results

Thinking
Tool usage

General Knowledge

1 evaluations
Benchmark / mode
Score
Rank/total
10.30
155 / 181

Other

2 evaluations
Benchmark / mode
Score
Rank/total
77
133 / 226
68
167 / 226

Coding and Software Engineer

3 evaluations
Benchmark / mode
Score
Rank/total
70.30
61 / 114
62.30
80 / 114
GSO
Standard ModeTools
3.80
19 / 21

Math and Reasoning

5 evaluations
Benchmark / mode
Score
Rank/total
82.20
41 / 44
54.80
85 / 107
23.30
58 / 62
4.10
41 / 60
3.10
46 / 60

Common Sense Reasoning

2 evaluations
Benchmark / mode
Score
Rank/total
SimpleBench
Standard Mode
44.90
39 / 67
SimpleBench
Thinking Enabled
46.40
35 / 67

Agent Level Benchmark

7 evaluations
Benchmark / mode
Score
Rank/total
Aider-Polyglot
Standard Mode
60.40
21 / 59
64.90
18 / 59
61.80
30 / 43
METR Time Horizons v1.1
Standard ModeTools
60.39
16 / 22
56.09
17 / 22

Productivity Knowledge

1 evaluations
Benchmark / mode
Score
Rank/total
28
20 / 21

Long Context

1 evaluations
Benchmark / mode
Score
Rank/total
61
18 / 18

AI Agent - Tool Usage

1 evaluations
Benchmark / mode
Score
Rank/total

Multimodal Understanding

3 evaluations
Benchmark / mode
Score
Rank/total
GeoBench ACW
Standard Mode
68
14 / 20
VPCT
Standard Mode
39
18 / 24
VPCT
64K
35
22 / 24

Other

1 evaluations
Benchmark / mode
Score
Rank/total
Fiction.liveBench
Standard Mode
50
16 / 16

Sources