DataLearner logo

GPT OSS 120B Benchmark Details

GPT OSS 120B currently shows benchmark results led by AIME 2024 (2 / 62, score 96.60), MMLU (11 / 124, score 90), AIME2025 (17 / 107, score 97.90).

Benchmark Results

GPT OSS 120B

Benchmark Results

Thinking
Tool usage

General Knowledge

6 evaluations
Benchmark / mode
Score
Rank/total
MMLU
Thinking Mode
90
11 / 124
MMLU Pro
Thinking Mode
79
67 / 134
LiveBench
Standard Mode
46.09
104 / 117
24.13
27 / 28
HLE
Thinking Mode
14.90
158 / 197
HLE
Thinking ModeTools
19
145 / 197

Other

1 evaluations
Benchmark / mode
Score
Rank/total
GPQA Diamond
Thinking Mode
80.10
146 / 274

Coding and Software Engineer

4 evaluations
Benchmark / mode
Score
Rank/total
CodeForces
Thinking Mode
2463
11 / 21
CodeForces
Thinking ModeTools
2622
9 / 21
SWE-bench Verified
Thinking Mode
60.10
83 / 116
SciCode
Thinking Level · High
38.89
14 / 16

Math and Reasoning

3 evaluations
Benchmark / mode
Score
Rank/total
AIME2025
Thinking Mode
83
52 / 107
AIME2025
Thinking ModeTools
97.90
17 / 107
AIME 2024
Thinking ModeTools
96.60
2 / 62

Writing and Creative Capabilities

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
959.20
89 / 106

Common Sense Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleBench
Thinking Mode
22.10
87 / 94

Agent Level Benchmark

2 evaluations
Benchmark / mode
Score
Rank/total
Aider-Polyglot
Thinking Level · High
41.80
38 / 59
τ³-Banking
Thinking Level · HighTools
12.78
13 / 13

Instruction Following

1 evaluations
Benchmark / mode
Score
Rank/total
IF Bench
Standard Mode
69
21 / 36

Long Context

1 evaluations
Benchmark / mode
Score
Rank/total
AA-LCR
Thinking Level · High
51
29 / 29

Claw-style Agent Evaluation

1 evaluations
Benchmark / mode
Score
Rank/total
Pinch Bench
Thinking ModeTools
60.60
36 / 38

Productivity Knowledge

1 evaluations
Benchmark / mode
Score
Rank/total
GDPval-AA v2
Thinking Level · HighTools
802.63
26 / 27