GPT-4vsGPT-3.5
Across 4 shared benchmarks, GPT-4 leads overall: GPT-4 wins 4, GPT-3.5 wins 0, with 0 ties and an average score difference of +21.13.
GPT-44 wins(100%)(0%)0 winsGPT-3.5
Benchmark scores
Grouped by capability, sorted by largest gap within each. 4 shared benchmarks.
General Knowledge
GPT-4 2/2| Benchmark | GPT-4 | GPT-3.5 | Diff |
|---|---|---|---|
| MMLU | 86.4032 / 124Normal (No Tools) | 7073 / 124Historical report (mode unspecified) | +16.40 |
| C-Eval | 68.7021 / 48Historical report (mode unspecified) | 54.4027 / 48Historical report (mode unspecified) | +14.30 |
Coding and Software Engineer
GPT-4 1/1| Benchmark | GPT-4 | GPT-3.5 | Diff |
|---|---|---|---|
| HumanEval | 6736 / 101Normal (No Tools) | 48.1054 / 101Historical report (mode unspecified) | +18.90 |
Math and Reasoning
GPT-4 1/1| Benchmark | GPT-4 | GPT-3.5 | Diff |
|---|---|---|---|
| GSM8K | 9211 / 70Historical report (mode unspecified) | 57.1037 / 70Historical report (mode unspecified) | +34.90 |
Specs
| Field | GPT-4 | GPT-3.5 |
|---|---|---|
| Publisher | OpenAI | OpenAI |
| Release date | 2023-03-14 | 2022-11-30 |
| Model type | Foundation model | Chat model |
| Architecture | Dense | Dense |
| Parameters | 175B | 175B |
| Context length | 128K | 4K |
| Max output | Not available | Not available |
Summary
- GPT-4leads in:General Knowledge (2/2), Coding and Software Engineer (1/1), Math and Reasoning (1/1)
On average across the 4 shared benchmarks, GPT-4 scores 21.13 higher.
Largest single-benchmark gap: GSM8K — GPT-4 92 vs GPT-3.5 57.10 (+34.90).
Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.