DataLearner logo

GPT OSS 120B Benchmark Details

GPT OSS 120B currently shows benchmark results led by AIME 2024 (2 / 62, score 96.60), MMLU (10 / 66, score 90), AIME2025 (17 / 107, score 97.90).

Benchmark Results

GPT OSS 120B

Benchmark Results

Thinking
Tool usage

General Knowledge

6 evaluations
Benchmark / mode
Score
Rank/total
90
10 / 66
80.10
81 / 187
79
66 / 132
LiveBench
Standard Mode
46.09
102 / 115
19
122 / 172
14.90
135 / 172

Coding and Software Engineer

3 evaluations
Benchmark / mode
Score
Rank/total
2622
8 / 16
2463
10 / 16
60.10
81 / 112

Math and Reasoning

3 evaluations
Benchmark / mode
Score
Rank/total
97.90
17 / 107
83
52 / 107
96.60
2 / 62

Common Sense Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
Simple Bench
Thinking Enabled
22.10
57 / 63

Agent Level Benchmark

1 evaluations
Benchmark / mode
Score
Rank/total
Aider-Polyglot
Thinking Level · High
41.80
38 / 59

Instruction Following

1 evaluations
Benchmark / mode
Score
Rank/total
69
17 / 30

Claw-style Agent Evaluation

1 evaluations
Benchmark / mode
Score
Rank/total
Pinch Bench
Thinking EnabledTools
60.60
35 / 37