GPT-4vsGPT-3.5
在 4 个共同 benchmark 中,GPT-4 整体领先:GPT-4 领先 4 项,GPT-3.5 领先 0 项,持平 0 项,平均分差 +21.13。
GPT-44 项(100%)(0%)0 项GPT-3.5
评测分数
按能力类目分组,每组内按分差大小排列;共 4 项。
General Knowledge
GPT-4 领先 2/2| 评测项 | GPT-4 | GPT-3.5 | 分差 |
|---|---|---|---|
| MMLU | 86.4032 / 124Normal (No Tools) | 7073 / 124Historical report (mode unspecified) | +16.40 |
| C-Eval | 68.7021 / 48Historical report (mode unspecified) | 54.4027 / 48Historical report (mode unspecified) | +14.30 |
Coding and Software Engineer
GPT-4 领先 1/1| 评测项 | GPT-4 | GPT-3.5 | 分差 |
|---|---|---|---|
| HumanEval | 6736 / 101Normal (No Tools) | 48.1054 / 101Historical report (mode unspecified) | +18.90 |
Math and Reasoning
GPT-4 领先 1/1| 评测项 | GPT-4 | GPT-3.5 | 分差 |
|---|---|---|---|
| GSM8K | 9211 / 70Historical report (mode unspecified) | 57.1037 / 70Historical report (mode unspecified) | +34.90 |
规格对比
| 字段 | GPT-4 | GPT-3.5 |
|---|---|---|
| 发布机构 | OpenAI | OpenAI |
| 发布时间 | 2023-03-14 | 2022-11-30 |
| 模型类型 | 基础大模型 | 聊天大模型 |
| 架构 | 稠密模型 | 稠密模型 |
| 参数规模 | 1750亿 | 1750亿 |
| 上下文长度 | 128K | 4K |
| 最大输出 | 暂无数据 | 暂无数据 |
小结
- GPT-4在以下类目领先:General Knowledge (2/2)、Coding and Software Engineer (1/1)、Math and Reasoning (1/1)
4 个共同 benchmark 上,GPT-4 平均高出 21.13 分。
单项差距最大的 benchmark:GSM8K — GPT-4 92,GPT-3.5 57.10(分差 +34.90)。
本页正文由结构化模型、价格与 benchmark 数据生成,不使用实时 LLM 撰写。