DataLearner logo

GPT-5.1 Codex Benchmark Details

GPT-5.1 Codex currently shows benchmark results led by Terminal-Bench (2 / 35, score 56.30), LiveCodeBench (15 / 123, score 85.50), LiveBench (45 / 115, score 68.61).

Benchmark Results

GPT-5.1 Codex

Benchmark Results

Thinking

Coding and Software Engineer

2 evaluations
Benchmark / mode
Score
Rank/total
85.50
15 / 123
70.40
58 / 112

AI Agent - Tool Usage

1 evaluations
Benchmark / mode
Score
Rank/total
56.30
2 / 35

General Knowledge

1 evaluations
Benchmark / mode
Score
Rank/total
LiveBench
Standard Mode
68.61
45 / 115