DataLearner logo

GPT-5.1 Codex Benchmark Details

GPT-5.1 Codex currently shows benchmark results led by Terminal-Bench (2 / 35, score 56.30), LiveCodeBench (19 / 125, score 85.50), Terminal Bench Hard (75 / 244, score 34.80).

Benchmark Results

GPT-5.1 Codex

Benchmark Results

Thinking
Tool usage

Repository Engineering

1 evaluations
Benchmark / mode
Score
Rank/total
SWE-bench Verified
Thinking Level · HighTools
70.40
60 / 116

Algorithmic Coding

1 evaluations
Benchmark / mode
Score
Rank/total
LiveCodeBench
Thinking Level · HighTools
85.50
19 / 125

Agentic Development

2 evaluations
Benchmark / mode
Score
Rank/total
Terminal-Bench
Thinking Level · HighTools
56.30
2 / 35
Terminal Bench Hard
Thinking Level · HighTools
34.80
75 / 244

Cross-capability Suites

1 evaluations
Benchmark / mode
Score
Rank/total
LiveBench
Standard Mode
68.61
45 / 117

Visual Understanding

1 evaluations
Benchmark / mode
Score
Rank/total
MMMU-Pro
Thinking Level · High
72.50
116 / 229

Scientific Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
CritPt
Thinking Level · High
5.70
92 / 204

Code Generation & Editing

1 evaluations
Benchmark / mode
Score
Rank/total
Vibe Code Bench v1.1
Thinking Level · HighTools
13.12
54 / 61