DataLearner logo

GPT-4o(2024-11-20)vsGPT-4o

Across 8 shared benchmarks, GPT-4o leads overall: GPT-4o(2024-11-20) wins 2, GPT-4o wins 3, with 3 ties and an average score difference of -1.81.

OpenAI
GPT-4o(2024-11-20)

OpenAI · 2024-11-20 · Chat model

OpenAI
GPT-4o

OpenAI · 2024-05-13 · Multimodal model

GPT-4o(2024-11-20)2 wins(25%)Ties3(38%)3 winsGPT-4o

Benchmark scores

Grouped by capability, sorted by largest gap within each. 8 shared benchmarks.

Coding and Software Engineer

GPT-4o(2024-11-20) 1/2
BenchmarkGPT-4o(2024-11-20)GPT-4oDiff
HumanEval90.207 / 39908 / 39+0.20
SWE-bench Verified31107 / 112Normal (No Tools)31107 / 112

General Knowledge

GPT-4o 1/2
BenchmarkGPT-4o(2024-11-20)GPT-4oDiff
MMLU85.7037 / 6688.7015 / 66-3
MMLU Pro77.9075 / 13277.9075 / 132

Math and Reasoning

GPT-4o 1/2
BenchmarkGPT-4o(2024-11-20)GPT-4oDiff
MATH68.5024 / 4275.9016 / 42-7.40
FrontierMath0.3057 / 600.3057 / 60

Agent Level Benchmark

GPT-4o 1/1
BenchmarkGPT-4o(2024-11-20)GPT-4oDiff
Aider-Polyglot18.2050 / 59Normal (No Tools)23.1047 / 59Normal (No Tools)-4.90

Common Sense

GPT-4o(2024-11-20) 1/1
BenchmarkGPT-4o(2024-11-20)GPT-4oDiff
SimpleQA38.8021 / 4738.2022 / 47+0.60

Specs

FieldGPT-4o(2024-11-20)GPT-4o
PublisherOpenAIOpenAI
Release date2024-11-202024-05-13
Model typeChat modelMultimodal model
ArchitectureDenseDense
ParametersNot availableNot available
Context length128K128K
Max outputNot available16K

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemGPT-4o(2024-11-20)GPT-4o
Text inputNot public$2.5 / 1M tokens
Text outputNot public$10 / 1M tokens

One or both models have incomplete public pricing.

Summary

  • GPT-4o(2024-11-20)leads in:Coding and Software Engineer (1/2), Common Sense (1/1)
  • GPT-4oleads in:General Knowledge (1/2), Math and Reasoning (1/2), Agent Level Benchmark (1/1)

On average across the 8 shared benchmarks, GPT-4o scores 1.81 higher.

Largest single-benchmark gap: MATH — GPT-4o(2024-11-20) 68.50 vs GPT-4o 75.90 (-7.40).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.