DataLearner logo

GPT-4vsGPT-3.5

Across 4 shared benchmarks, GPT-4 leads overall: GPT-4 wins 4, GPT-3.5 wins 0, with 0 ties and an average score difference of +21.13.

OpenAI
GPT-4

OpenAI · 2023-03-14 · Foundation model

OpenAI
GPT-3.5

OpenAI · 2022-11-30 · Chat model

GPT-44 wins(100%)(0%)0 winsGPT-3.5

Benchmark scores

Grouped by capability, sorted by largest gap within each. 4 shared benchmarks.

General Knowledge

GPT-4 2/2
BenchmarkGPT-4GPT-3.5Diff
MMLU86.4032 / 124Normal (No Tools)7073 / 124Historical report (mode unspecified)+16.40
C-Eval68.7021 / 48Historical report (mode unspecified)54.4027 / 48Historical report (mode unspecified)+14.30

Coding and Software Engineer

GPT-4 1/1
BenchmarkGPT-4GPT-3.5Diff
HumanEval6736 / 101Normal (No Tools)48.1054 / 101Historical report (mode unspecified)+18.90

Math and Reasoning

GPT-4 1/1
BenchmarkGPT-4GPT-3.5Diff
GSM8K9211 / 70Historical report (mode unspecified)57.1037 / 70Historical report (mode unspecified)+34.90

Specs

FieldGPT-4GPT-3.5
PublisherOpenAIOpenAI
Release date2023-03-142022-11-30
Model typeFoundation modelChat model
ArchitectureDenseDense
Parameters175B175B
Context length128K4K
Max outputNot availableNot available

Summary

  • GPT-4leads in:General Knowledge (2/2), Coding and Software Engineer (1/1), Math and Reasoning (1/1)

On average across the 4 shared benchmarks, GPT-4 scores 21.13 higher.

Largest single-benchmark gap: GSM8K — GPT-4 92 vs GPT-3.5 57.10 (+34.90).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.