DataLearner logo

GPT-4 Benchmark Details

GPT-4 currently shows benchmark results led by GSM8K (11 / 70, score 92), MMLU (32 / 124, score 86.40), HumanEval (39 / 140, score 82). This page also compares it with 1 competitor models and 2 predecessor or same-series models, including performance and pricing views when available. 1 source link is attached for reference.

Benchmark Results

GPT-4

Benchmark Results

Thinking

General Knowledge

3 evaluations
Benchmark / mode
Score
Rank/total
MMLU
Standard Mode
86.40
32 / 124
MMLU
Standard Mode
86.40
32 / 124
C-Eval
Standard Mode
68.70
21 / 48

Math and Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
GSM8K
Standard Mode
92
11 / 70

Coding and Software Engineer

2 evaluations
Benchmark / mode
Score
Rank/total
HumanEval
Standard Mode
82
39 / 140
HumanEval
Standard Mode
67
63 / 140

Reading Comprehension

1 evaluations
Benchmark / mode
Score
Rank/total
DROP
Standard Mode
80.90
7 / 9

Writing and Creative Capabilities

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
749.90
99 / 106

Competitor Comparison

Benchmark scores for GPT-4 compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

4 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkGPT-4CurrentClaude3-Opus
MMLU
Accuracy
综合评估
86.40Standard Mode
86.80Standard Mode
GSM8K
Accuracy
数学推理
92.00Standard Mode
95.00Standard Mode
HumanEval
Accuracy
编程与软件工程
82.00Standard Mode
84.90Standard Mode
DROP
F1
阅读理解
80.90Standard Mode
83.10Standard Mode

Standard API Pricing: GPT-4 vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
Claude3-Opus
Anthropic$15 / 1M tokens$75 / 1M tokens

Version History

How each version of the GPT-4 series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

4 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkGPT-4CurrentGPT-3.5GPT-3
C-Eval
Accuracy
综合评估
68.70Standard Mode
54.40Standard Mode
--
MMLU
Accuracy
综合评估
86.40Standard Mode
70.00Standard Mode
53.90Standard Mode
GSM8K
Accuracy
数学推理
92.00Standard Mode
57.10Standard Mode
--
HumanEval
Accuracy
编程与软件工程
82.00Standard Mode
48.10Standard Mode
--

Single-Benchmark Version Trend

Viewing: C-Eval · 综合评估

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the GPT-4 Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier.

Comparable standard text pricing is not available for these models.

Sources