DataLearner logo

GPT OSS 20B Benchmark Details

GPT OSS 20B currently shows benchmark results led by AIME 2024 (3 / 61, score 96), AIME2025 (15 / 110, score 98.70), MMLU (41 / 124, score 85.30).

Benchmark Results

GPT OSS 20B

Benchmark Results

Thinking
Tool usage

Knowledge Exams

5 evaluations
Benchmark / mode
Score
Rank/total
MMLU
Thinking Mode
85.30
41 / 124
MMLU-Pro
Thinking Level · Medium
73.14
92 / 175
MMLU-Pro
Thinking Mode
74
90 / 175
HLE
Thinking Mode
10.90
191 / 235
HLE
Thinking ModeTools
17.30
173 / 235

Scientific Reasoning

5 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Thinking Level · Low
53.16
225 / 253
GPQA Diamond
Thinking Level · Medium
60.80
209 / 253
GPQA Diamond
Thinking Mode
71.50
170 / 253
GPQA Diamond
Thinking Level · High
45.96
236 / 253
CritPt
Thinking Level · High
1.40
139 / 204

Repository Engineering

1 evaluations
Benchmark / mode
Score
Rank/total
SWE-bench Verified
Thinking Mode
34
110 / 116

Mathematics

4 evaluations
Benchmark / mode
Score
Rank/total
AIME2025
Thinking Mode
79
68 / 110
AIME2025
Thinking ModeTools
98.70
15 / 110
AIME2025
Thinking Level · HighTools
89.17
43 / 110
AIME 2024
Thinking ModeTools
96
3 / 61

Writing

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
665.60
108 / 110

Algorithmic Coding

2 evaluations
Benchmark / mode
Score
Rank/total
CodeForces
Thinking Mode
2230
13 / 21
CodeForces
Thinking ModeTools
2516
10 / 21

Service Workflows

2 evaluations
Benchmark / mode
Score
Rank/total
τ²-Bench
Thinking ModeTools
47.70
37 / 44
τ³-Banking
Thinking Level · HighTools
7
144 / 167

Fact Finding

1 evaluations
Benchmark / mode
Score
Rank/total
BrowseComp
Thinking ModeTools
28.30
55 / 58

Agentic Development

2 evaluations
Benchmark / mode
Score
Rank/total
Terminal Bench Hard
Thinking Level · LowTools
4.50
214 / 244
Terminal Bench Hard
Thinking Level · HighTools
10.60
182 / 244

Tool Orchestration

2 evaluations
Benchmark / mode
Score
Rank/total
Pinch Bench
Thinking ModeTools
66
35 / 38
36.34
42 / 45

Scientific Computing

1 evaluations
Benchmark / mode
Score
Rank/total
SciCode
Thinking Level · High
38.90
113 / 134

Documents & Charts

1 evaluations
Benchmark / mode
Score
Rank/total
GDP.pdf
Thinking Level · High
2
110 / 122

Capability Indices

1 evaluations
Benchmark / mode
Score
Rank/total
ECI
unknown
137.81
104 / 167