DataLearner logo

DeepSeek-V3vsGPT-4o(2024-11-20)

Across 7 shared benchmarks, DeepSeek-V3 leads overall: DeepSeek-V3 wins 4, GPT-4o(2024-11-20) wins 3, with 0 ties and an average score difference of +5.23.

DeepSeek-AI
DeepSeek-V3

DeepSeek-AI · 2024-12-26 · Chat model

OpenAI
GPT-4o(2024-11-20)

OpenAI · 2024-11-20 · Chat model

DeepSeek-V34 wins(57%)(43%)3 winsGPT-4o(2024-11-20)

Benchmark scores

Grouped by capability, sorted by largest gap within each. 7 shared benchmarks.

General Knowledge

Even 2/2
BenchmarkDeepSeek-V3GPT-4o(2024-11-20)Diff
MMLU88.5017 / 6685.7037 / 66+2.80
MMLU Pro75.9083 / 13277.9075 / 132-2

Math and Reasoning

DeepSeek-V3 2/2
BenchmarkDeepSeek-V3GPT-4o(2024-11-20)Diff
MATH87.807 / 4268.5024 / 42+19.30
FrontierMath1.7049 / 600.3057 / 60+1.40

Agent Level Benchmark

DeepSeek-V3 1/1
BenchmarkDeepSeek-V3GPT-4o(2024-11-20)Diff
Aider-Polyglot48.4034 / 59Normal (No Tools)18.2050 / 59Normal (No Tools)+30.20

Coding and Software Engineer

GPT-4o(2024-11-20) 1/1
BenchmarkDeepSeek-V3GPT-4o(2024-11-20)Diff
HumanEval899 / 3990.207 / 39-1.20

Common Sense

GPT-4o(2024-11-20) 1/1
BenchmarkDeepSeek-V3GPT-4o(2024-11-20)Diff
SimpleQA24.9031 / 4738.8021 / 47-13.90

Specs

FieldDeepSeek-V3GPT-4o(2024-11-20)
PublisherDeepSeek-AIOpenAI
Release date2024-12-262024-11-20
Model typeChat modelChat model
ArchitectureDenseDense
Parameters681BNot available
Context length128K128K
Max outputNot availableNot available

API pricing

Prices use DataLearner records when available; missing fields are not inferred.

ItemDeepSeek-V3GPT-4o(2024-11-20)
Text input$0.27 / 1M tokensNot public
Text output$1.1 / 1M tokensNot public
Cache read$0.07 / 1M tokensNot public
Cache write$0.27 / 1M tokensNot public

One or both models have incomplete public pricing.

Summary

  • DeepSeek-V3leads in:Math and Reasoning (2/2), Agent Level Benchmark (1/1)
  • GPT-4o(2024-11-20)leads in:Coding and Software Engineer (1/1), Common Sense (1/1)
  • Tied in:General Knowledge

On average across the 7 shared benchmarks, DeepSeek-V3 scores 5.23 higher.

Largest single-benchmark gap: Aider-Polyglot — DeepSeek-V3 48.40 vs GPT-4o(2024-11-20) 18.20 (+30.20).

Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.