DataLearner logo

GPT-4 Benchmark Details

GPT-4 currently shows benchmark results led by MMLU (31 / 66, score 86.40), HumanEval (27 / 39, score 67), DROP (7 / 9, score 80.90). This page also compares it with 1 competitor models and 2 predecessor or same-series models, including performance and pricing views when available. 1 source link is attached for reference.

Benchmark Results

GPT-4

Benchmark Results

Thinking

General Knowledge

1 evaluations
Benchmark / mode
Score
Rank/total
MMLU
Standard Mode
86.40
31 / 66

Coding and Software Engineer

1 evaluations
Benchmark / mode
Score
Rank/total
HumanEval
Standard Mode
67
27 / 39

Reading Comprehension

1 evaluations
Benchmark / mode
Score
Rank/total
DROP
Standard Mode
80.90
7 / 9

Competitor Comparison

Benchmark scores for GPT-4 compared against top models in its class

Benchmark categories:
The chart shows each model’s highest score per benchmark within the current filter. Out-of-100 benchmarks use raw heights; out-of-range benchmarks are scaled within that benchmark while labels keep the original scores.

3 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.

BenchmarkGPT-4CurrentClaude3-Opus
MMLU
综合评估
86.40Standard Mode
86.80Standard Mode
HumanEval
编程与软件工程
67.00Standard Mode
84.90Standard Mode
DROP
阅读理解
80.90Standard Mode
83.10Standard Mode

Standard API Pricing: GPT-4 vs. Peer Models

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens

ModelSupplierStandard inputStandard outputBase price applies to
Claude3-Opus
Anthropic$15 / 1M tokens$75 / 1M tokens

Version History

How each version of the GPT-4 series stacks up on benchmark tests

No benchmark data matches the selected filters.

Standard API Pricing Across the GPT-4 Series

Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.

Source: DataLearnerAI. Standard text prices shown here use the default supplier.

Comparable standard text pricing is not available for these models.

Sources