GPT-4 Benchmark Details
GPT-4 currently shows benchmark results led by MMLU (31 / 66, score 86.40), HumanEval (27 / 39, score 67), DROP (7 / 9, score 80.90). This page also compares it with 1 competitor models and 2 predecessor or same-series models, including performance and pricing views when available. 1 source link is attached for reference.
Benchmark Results
Benchmark Results
Competitor Comparison
Benchmark scores for GPT-4 compared against top models in its class
3 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.
Standard API Pricing: GPT-4 vs. Peer Models
Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.
Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens
| Model | Supplier | Standard input | Standard output | Base price applies to |
|---|---|---|---|---|
Claude3-Opus | Anthropic | $15 / 1M tokens | $75 / 1M tokens | — |
Version History
How each version of the GPT-4 series stacks up on benchmark tests
Standard API Pricing Across the GPT-4 Series
Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.
Source: DataLearnerAI. Standard text prices shown here use the default supplier.