DataLearner logo

Claude3-OpusvsGPT-4

Across 4 shared benchmarks, Claude3-Opus leads overall: Claude3-Opus wins 4, GPT-4 wins 0, with 0 ties and an average score difference of +9.

Anthropic
Claude3-Opus

Anthropic · 2024-03-04 · Multimodal model

OpenAI
GPT-4

OpenAI · 2023-03-14 · Foundation model

Claude3-Opus4 wins(100%)(0%)0 winsGPT-4

Benchmark scores

Grouped by capability, sorted by largest gap within each. 4 shared benchmarks.

Coding and Software Engineer

Claude3-Opus 1/1
BenchmarkClaude3-OpusGPT-4Diff
HumanEval84.9037 / 140Normal (No Tools)6763 / 140Normal (No Tools)+17.90

General Evaluation

Claude3-Opus 1/1
BenchmarkClaude3-OpusGPT-4Diff
GPQA Diamond50.40408 / 462Normal (No Tools)34.90444 / 462Normal (No Tools)+15.50

General Knowledge

Claude3-Opus 1/1
BenchmarkClaude3-OpusGPT-4Diff
MMLU86.8028 / 124Normal (No Tools)86.4032 / 124Normal (No Tools)+0.40

Reading Comprehension

Claude3-Opus 1/1
BenchmarkClaude3-OpusGPT-4Diff
DROP83.106 / 9Normal (No Tools)80.907 / 9Normal (No Tools)+2.20

Specs

FieldClaude3-OpusGPT-4
PublisherAnthropicOpenAI
Release date2024-03-042023-03-14
Model typeMultimodal modelFoundation model
ArchitectureDenseDense
ParametersNot available175B
Context length200K128K
Max outputNot availableNot available

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemClaude3-OpusGPT-4
Text input$15 / 1M tokensNot public
Text output$75 / 1M tokensNot public
Cache read$1.5 / 1M tokensNot public
Cache write$18.75 / 1M tokensNot public

One or both models have incomplete public pricing.

Summary

  • Claude3-Opusleads in:Coding and Software Engineer (1/1), General Evaluation (1/1), General Knowledge (1/1), Reading Comprehension (1/1)

On average across the 4 shared benchmarks, Claude3-Opus scores 9 higher.

Largest single-benchmark gap: HumanEval — Claude3-Opus 84.90 vs GPT-4 67 (+17.90).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.