DataLearner logo

GPT-5.1-Codex-Max Benchmark Details

GPT-5.1-Codex-Max currently shows benchmark results led by Terminal-Bench (1 / 35, score 58.10), LiveBench (22 / 115, score 73.98), SWE-bench Verified (30 / 114, score 76.80).

Benchmark Results

GPT-5.1-Codex-Max

Benchmark Results

Thinking
Tool usage

Coding and Software Engineer

1 evaluations
Benchmark / mode
Score
Rank/total
76.80
30 / 114

AI Agent - Tool Usage

1 evaluations
Benchmark / mode
Score
Rank/total
58.10
1 / 35

General Knowledge

1 evaluations
Benchmark / mode
Score
Rank/total
LiveBench
Deep Thinking Mode
73.98
22 / 115

Agent Level Benchmark

1 evaluations
Benchmark / mode
Score
Rank/total
METR Time Horizons v1.1
Standard ModeTools
161.75
10 / 22