DataLearner logo

GPT-4 Benchmark Details

GPT-4 currently shows benchmark results led by GSM8K (11 / 70, score 92), HumanEval (22 / 101, score 82), MMLU (32 / 124, score 86.40). This page also compares it with 1 competitor models and 2 predecessor or same-series models, including performance and pricing views when available. 1 source link is attached for reference.

Benchmark Results

GPT-4

Benchmark Results

Thinking

General Knowledge

3 evaluations
Benchmark / mode
Score
Rank/total
MMLU
Standard Mode
86.40
32 / 124
MMLU
Standard Mode
86.40
32 / 124
C-Eval
Standard Mode
68.70
21 / 48

Math and Reasoning

1 evaluations
Benchmark / mode
Score
Rank/total
GSM8K
Standard Mode
92
11 / 70

Coding and Software Engineer

2 evaluations
Benchmark / mode
Score
Rank/total
HumanEval
Standard Mode
82
22 / 101
HumanEval
Standard Mode
67
36 / 101

Reading Comprehension

1 evaluations
Benchmark / mode
Score
Rank/total
DROP
Standard Mode
80.90
7 / 9

Writing and Creative Capabilities

1 evaluations
Benchmark / mode
Score
Rank/total
Creative Writing
Standard Mode
749.30
92 / 99

Competitor Comparison

Benchmark scores for GPT-4 compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

4 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkGPT-4CurrentClaude3-Opus
MMLU
综合评估
86.40Standard Mode
86.80Standard Mode
GSM8K
数学推理
92.00Standard Mode
95.00Standard Mode
HumanEval
编程与软件工程
82.00Standard Mode
84.90Standard Mode
DROP
阅读理解
80.90Standard Mode
83.10Standard Mode

Standard API Pricing: GPT-4 vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
Claude3-Opus
Anthropic$15 / 1M tokens$75 / 1M tokens

Version History

How each version of the GPT-4 series stacks up on benchmark tests

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

4 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.

BenchmarkGPT-4CurrentGPT-3.5GPT-3
C-Eval
综合评估
68.70Standard Mode
54.40Standard Mode
--
MMLU
综合评估
86.40Standard Mode
70.00Standard Mode
53.90Standard Mode
GSM8K
数学推理
92.00Standard Mode
57.10Standard Mode
--
HumanEval
编程与软件工程
82.00Standard Mode
48.10Standard Mode
--

Single-Benchmark Version Trend

Viewing: C-Eval · 综合评估

Benchmark
NormalNormal + ToolsThinkingThinking + ToolsDeepDeep + Tools

X-axis shows model and release date, Y-axis shows score; solid lines connect the same mode across versions, while dotted guides align modes within the same generation.

Standard API Pricing Across the GPT-4 Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier.

Comparable standard text pricing is not available for these models.

Sources