DataLearner logo
OP

OpenAI o3

Scheduled retirement · 2026-12-11Reasoning modelo-serieso3

OpenAI o3

Release date: 2025-04-16Updated: 2026-09-04Scheduled retirement: 2026-12-11Views: 1,430
Live demoGitHubHugging FaceCompare
Parameters
No data
Context length
200K
Multilingual
Supported
Reasoning ability
5/5

OpenAI o3 is an AI model published by OpenAI, released on 2025-04-16, for Reasoning model, and 200K context length, with a 1672.50 score on Creative Writing.

Data sourced primarily from official releases (GitHub, Hugging Face, papers), then benchmark leaderboards, then third-party evaluators. Learn about our data methodology

OpenAI o3

Model basics

Reasoning traces
Supported
Thinking modes
Thinking Level · Deep
Context length
200K tokens
Max output length
100K tokens
Model type
Reasoning model
Modality (in / out)
Text, Image → Text
Release date
2025-04-16
Model file size
No data
MoE architecture
No
Total params / Active params
No data / Not applicable
Knowledge cutoff
No data
OpenAI o3

Open source & experience

Code license
Proprietary
Weights license
Proprietary
GitHub repo
N/A
Hugging Face
N/A
Live demo
OpenAI o3

Official resources

Paper
DataLearnerAI blog
N/A
OpenAI o3

API details

API speed
1/5
💡Default unit: $/1M tokens. If vendors use other units, follow their published pricing.
Standard
TypeConditionInputOutput
Text-$10.00/ 1M tokens$40.00/ 1M tokens
Image-$10.00/ 1M tokens

“—” means the modality is not billed in that direction, or the vendor has not published a price for it.

OpenAI o3

Benchmark Results

OpenAI o3 currently shows benchmark results led by Aider-Polyglot (5 / 59, score 81.30), MATH-500 (5 / 45, score 98.10), MMLU Pro (24 / 134, score 85.60). This page also consolidates core specs, context limits, and API pricing so you can evaluate the model from benchmark results and deployment constraints together.

Thinking
Tool usage

General Knowledge

4 evaluations
Benchmark / mode
Score
Rank/total
MMLU Pro
Standard Mode
85.60
24 / 134
ARC-AGI-1
Thinking Mode
60.80
61 / 92
HLE
Thinking Mode
20.32
139 / 197
ARC-AGI-2
Thinking Mode
6.50
67 / 85

General Evaluation

3 evaluations
Benchmark / mode
Score
Rank/total
79.80
149 / 274
80.81
142 / 274
GPQA Diamond
Thinking Mode
83.30
121 / 274

Common Sense

1 evaluations
Benchmark / mode
Score
Rank/total
SimpleQA
Standard Mode
49.40
14 / 47

Coding and Software Engineer

5 evaluations
Benchmark / mode
Score
Rank/total
CodeClash
Standard ModeTools
1343
3 / 8
LiveCodeBench
Standard Mode
75.80
42 / 128
SWE-bench Verified
Thinking Mode
69.10
67 / 116
WeirdML v2
HighTools
52.42
36 / 52
GSO
HighTools
8.80
15 / 21

Math and Reasoning

12 evaluations
Benchmark / mode
Score
Rank/total
MATH-500
Standard Mode
98.10
5 / 45
AIME 2024
Standard Mode
91.60
12 / 62
AIME2025
Thinking Mode
88.90
41 / 107
19.30
51 / 58
29.82
44 / 58
33.33
42 / 58
IMO-ProofBench
Thinking Mode
20.50
11 / 16
20.50
8 / 19
10.30
25 / 60
10
28 / 60
10.30
25 / 60
2.10
56 / 80

Writing and Creative Capabilities

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
1672.50
31 / 106

AI Agent - Tool Usage

1 evaluations
Benchmark / mode
Score
Rank/total
Terminal-Bench
Thinking Mode
30.20
21 / 35

Multimodal Understanding

5 evaluations
Benchmark / mode
Score
Rank/total
MMMU
Standard Mode
82.90
6 / 28
MMMU
Thinking Mode
82.90
6 / 28
74
11 / 20
60
19 / 20
VPCT
Medium
52
10 / 24

Common Sense Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
53.10
47 / 94

Agent Level Benchmark

3 evaluations
Benchmark / mode
Score
Rank/total
91.27
14 / 22
Aider-Polyglot
Standard Mode
76.90
9 / 59
81.30
5 / 59

Long Context

1 evaluations
Benchmark / mode
Score
Rank/total
88.90
6 / 16
OpenAI o3

Publisher

OpenAI o3

Model Overview

OpenAI o3 is an AI model published by OpenAI, released on 2025-04-16, for Reasoning model, and 200K context length, with a 1672.50 score on Creative Writing.

DataLearner on WeChat

Follow DataLearner on WeChat for AI model updates and research notes.

DataLearner WeChat QR code