DeepSeek-V4-FlashvsHy3 Pre
DeepSeek-V4-Flash and Hy3 Pre are tied across 5 shared benchmarks: DeepSeek-V4-Flash leads on 2, Hy3 Pre leads on 2, with 1 ties and an average score difference of +5.28.
DeepSeek-AI · 2026-04-24 · Reasoning model
腾讯AI实验室 · 2026-04-23 · Reasoning model
Benchmark scores
Grouped by capability, sorted by largest gap within each. 5 shared benchmarks.
Agent Level Benchmark
DeepSeek-V4-Flash 2/2| Benchmark | DeepSeek-V4-Flash | Hy3 Pre | Diff |
|---|---|---|---|
| τ²-Bench - Telecom | 94.4035 / 264Normal (With Tools) | 67.50143 / 264Normal (With Tools) | +26.90 |
| Terminal Bench Hard | 34.1082 / 244Normal (With Tools) | 31.80100 / 244Normal (With Tools) | +2.30 |
General Evaluation
Hy3 Pre 1/1| Benchmark | DeepSeek-V4-Flash | Hy3 Pre | Diff |
|---|---|---|---|
| GPQA Diamond | 71.20311 / 462Normal (No Tools) | 73.20299 / 462Normal (No Tools) | -2 |
General Knowledge
Even 1/1| Benchmark | DeepSeek-V4-Flash | Hy3 Pre | Diff |
|---|---|---|---|
| CritPt | 0.30182 / 200Normal (No Tools) | 0.30182 / 200Normal (No Tools) | — |
Instruction Following
Hy3 Pre 1/1| Benchmark | DeepSeek-V4-Flash | Hy3 Pre | Diff |
|---|---|---|---|
| IF Bench | 47.20170 / 282Normal (No Tools) | 48166 / 282Normal (No Tools) | -0.80 |
Specs
| Field | DeepSeek-V4-Flash | Hy3 Pre |
|---|---|---|
| Publisher | DeepSeek-AI | 腾讯AI实验室 |
| Release date | 2026-04-24 | 2026-04-23 |
| Model type | Reasoning model | Reasoning model |
| Architecture | MoE | MoE |
| Parameters | 284B | 295B |
| Context length | 1M | 256K |
| Max output | 384K | Not available |
API pricing
Prices use DataLearner records when available; missing fields are not inferred.
| Item | DeepSeek-V4-Flash | Hy3 Pre |
|---|---|---|
| Text input | $0.14 / 1M tokens | Not public |
| Text output | $0.28 / 1M tokens | Not public |
| Cache read | $0.0028 / 1M tokens | Not public |
One or both models have incomplete public pricing.
Summary
- DeepSeek-V4-Flashleads in:Agent Level Benchmark (2/2)
- Hy3 Preleads in:General Evaluation (1/1), Instruction Following (1/1)
- Tied in:General Knowledge
On average across the 5 shared benchmarks, DeepSeek-V4-Flash scores 5.28 higher.
Largest single-benchmark gap: τ²-Bench - Telecom — DeepSeek-V4-Flash 94.40 vs Hy3 Pre 67.50 (+26.90).
Page generated from structured model, pricing and benchmark records. No real-time LLM is used to write the prose.