DataLearner logo

Claude Sonnet 3.7 Benchmark Details

Claude Sonnet 3.7 currently shows benchmark results led by Aider-Polyglot (18 / 59, score 64.90), Simple Bench (31 / 63, score 46.40), GPQA Diamond (94 / 187, score 77). 1 source link is attached for reference.

Benchmark Results

Claude Sonnet 3.7

Benchmark Results

Thinking

General Knowledge

3 evaluations
Benchmark / mode
Score
Rank/total
77
94 / 187
68
128 / 187
10.30
146 / 172

Coding and Software Engineer

2 evaluations
Benchmark / mode
Score
Rank/total
70.30
59 / 112
62.30
78 / 112

Math and Reasoning

5 evaluations
Benchmark / mode
Score
Rank/total
82.20
41 / 44
54.80
85 / 107
23.30
58 / 62
4.10
41 / 60
3.10
46 / 60

Common Sense Reasoning

2 evaluations
Benchmark / mode
Score
Rank/total
Simple Bench
Standard Mode
44.90
35 / 63
Simple Bench
Thinking Enabled
46.40
31 / 63

Agent Level Benchmark

5 evaluations
Benchmark / mode
Score
Rank/total
Aider-Polyglot
Standard Mode
60.40
21 / 59
64.90
18 / 59
61.80
30 / 43

Productivity Knowledge

1 evaluations
Benchmark / mode
Score
Rank/total
28
20 / 21

Long Context

1 evaluations
Benchmark / mode
Score
Rank/total
61
15 / 15

AI Agent - Tool Usage

1 evaluations
Benchmark / mode
Score
Rank/total

Sources