GPT-4 Benchmark Details
GPT-4 currently shows benchmark results led by GSM8K (11 / 70, score 92), MMLU (32 / 124, score 86.40), HumanEval (39 / 140, score 82). This page also compares it with 1 competitor models and 2 predecessor or same-series models, including performance and pricing views when available. 1 source link is attached for reference.
Benchmark Results
Benchmark Results
General Knowledge
3 evaluationsCoding and Software Engineer
2 evaluationsWriting and Creative Capabilities
1 evaluationsCompetitor Comparison
Benchmark scores for GPT-4 compared against top models in its class
4 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.
Standard API Pricing: GPT-4 vs. Peer Models
Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.
Source: DataLearnerAI. Standard text prices shown here use the default supplier. · USD / 1M tokens
| Model | Supplier | Standard input | Standard output | Base price applies to |
|---|---|---|---|---|
Claude3-Opus | Anthropic | $15 / 1M tokens | $75 / 1M tokens | — |
Version History
How each version of the GPT-4 series stacks up on benchmark tests
4 benchmarks with comparable scores. Each model shows its best score; mode label is displayed below.· Click a row to view its trend chart.
Single-Benchmark Version Trend
Viewing: C-Eval · 综合评估
Standard API Pricing Across the GPT-4 Series
Shows standard text input and output pricing side by side for each model. If extended-context pricing exists, the chart keeps the base rate and explains the threshold below.
Source: DataLearnerAI. Standard text prices shown here use the default supplier.