DeepSeek-V3vsGPT-4o(2024-11-20)
Across 7 shared benchmarks, DeepSeek-V3 leads overall: DeepSeek-V3 wins 4, GPT-4o(2024-11-20) wins 3, with 0 ties and an average score difference of +5.23.
DeepSeek-V3
DeepSeek-AI · 2024-12-26 · Chat model
GPT-4o(2024-11-20)
OpenAI · 2024-11-20 · Chat model
DeepSeek-V34 wins(57%)(43%)3 winsGPT-4o(2024-11-20)
Benchmark scores
Grouped by capability, sorted by largest gap within each. 7 shared benchmarks.
General Knowledge
Even 2/2| Benchmark | DeepSeek-V3 | GPT-4o(2024-11-20) | Diff |
|---|---|---|---|
| MMLU | 88.5017 / 66 | 85.7037 / 66 | +2.80 |
| MMLU Pro | 75.9083 / 132 | 77.9075 / 132 | -2 |
Math and Reasoning
DeepSeek-V3 2/2| Benchmark | DeepSeek-V3 | GPT-4o(2024-11-20) | Diff |
|---|---|---|---|
| MATH | 87.807 / 42 | 68.5024 / 42 | +19.30 |
| FrontierMath | 1.7049 / 60 | 0.3057 / 60 | +1.40 |
Agent Level Benchmark
DeepSeek-V3 1/1| Benchmark | DeepSeek-V3 | GPT-4o(2024-11-20) | Diff |
|---|---|---|---|
| Aider-Polyglot | 48.4034 / 59Normal (No Tools) | 18.2050 / 59Normal (No Tools) | +30.20 |
Coding and Software Engineer
GPT-4o(2024-11-20) 1/1| Benchmark | DeepSeek-V3 | GPT-4o(2024-11-20) | Diff |
|---|---|---|---|
| HumanEval | 899 / 39 | 90.207 / 39 | -1.20 |
Common Sense
GPT-4o(2024-11-20) 1/1| Benchmark | DeepSeek-V3 | GPT-4o(2024-11-20) | Diff |
|---|---|---|---|
| SimpleQA | 24.9031 / 47 | 38.8021 / 47 | -13.90 |
Specs
| Field | DeepSeek-V3 | GPT-4o(2024-11-20) |
|---|---|---|
| Publisher | DeepSeek-AI | OpenAI |
| Release date | 2024-12-26 | 2024-11-20 |
| Model type | Chat model | Chat model |
| Architecture | Dense | Dense |
| Parameters | 681B | Not available |
| Context length | 128K | 128K |
| Max output | Not available | Not available |
API pricing
Prices use DataLearner records when available; missing fields are not inferred.
| Item | DeepSeek-V3 | GPT-4o(2024-11-20) |
|---|---|---|
| Text input | $0.27 / 1M tokens | Not public |
| Text output | $1.1 / 1M tokens | Not public |
| Cache read | $0.07 / 1M tokens | Not public |
| Cache write | $0.27 / 1M tokens | Not public |
One or both models have incomplete public pricing.
Summary
- DeepSeek-V3leads in:Math and Reasoning (2/2), Agent Level Benchmark (1/1)
- GPT-4o(2024-11-20)leads in:Coding and Software Engineer (1/1), Common Sense (1/1)
- Tied in:General Knowledge
On average across the 7 shared benchmarks, DeepSeek-V3 scores 5.23 higher.
Largest single-benchmark gap: Aider-Polyglot — DeepSeek-V3 48.40 vs GPT-4o(2024-11-20) 18.20 (+30.20).
Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.