Was the V4 Flash model record retired or removed?
No. This update preserves its lifecycle, independent record, and original weights while updating information and scores.
DeepSeek V4 Flash
V4 Flash remains a distinct 284B/13B MoE with its existing lifecycle and original weights. Updated 0731 max-effort scores come from the September 10 release table. The deepseek-v4-flash API alias temporarily routes to V4.1 Flash; this does not change the original model or weights. USD price rows on this page are historical, not current quotes for routed requests. Current routed CNY rates per million tokens: off-peak cache-hit/cache-miss/output 0.02/1/4; peak 0.04/2/8 (weekdays 09:00-12:00, 14:00-18:00 Asia/Shanghai).
Data sourced primarily from official releases (GitHub, Hugging Face, papers), then benchmark leaderboards, then third-party evaluators. Learn about our data methodology
| Type | Condition | Input | Output |
|---|---|---|---|
| Text | - | $0.140/ 1M tokens | $0.280/ 1M tokens |
| Type | TTL | Write | Read |
|---|---|---|---|
| Text | - | — | $0.0028/ 1M tokens |
“—” means the modality is not billed in that direction, or the vendor has not published a price for it.
DeepSeek-V4-Flash currently shows benchmark results led by LiveCodeBench (6 / 250, score 91.60), IF Bench (11 / 282, score 79.20), HLE (43 / 563, score 51.50). This page also consolidates core specs, context limits, and API pricing so you can evaluate the model from benchmark results and deployment constraints together.
This page preserves DeepSeek V4 Flash as a distinct model with its existing lifecycle status. No retirement, deletion, or model-page redirect is applied. The V4 Flash 0731 column in the supplied September 10 release table updates its comparison results; it must not be confused with V4.1 Flash.
GPQA Diamond 89.9; HLE pure-text subset 37.8; Codeforces rating 3289; MathArena Apex 58.6; Terminal-Bench 2.1/3.0/4.0: 82.7/7.6/7.0; DeepSWE v1.1 54.4; NL2Repo-Bench 54.2; CyberGym 76.7; SEC-Bench Pro 30.9; ExploitGym 1.8 (budget unspecified); HLE with tools 51.5; Automation-Bench 37.7; Agents’ Last Exam 25.2.
The supplied clarification identifies max effort, with tool use tracked separately and internet access not inferred. Pure-text HLE is kept separate from full HLE, and budget-unspecified ExploitGym is separate from 2h/6h/unlimited protocols. Dashes for ProgramBench and visual evaluations mean no result supplied, not zero. The original 0731 Toolathlon Verified score of 70.3 and other historical modes and third-party results are retained. The new Automation-Bench value 37.7 replaces 25.1 in the max-with-tools entry; the old release value remains documented in the local update record.
Existing specifications remain: 284B total and 13B active parameters, MoE, text-only input and output, 1M context and 384K maximum output. The V4 architecture combines Compressed Sparse Attention, Heavily Compressed Attention, mHC connections and the Muon optimizer. The original V4 Flash weights and their existing MIT license remain distinct from the new API route. The 0731 release updated API post-training, not the previously published weights.
According to the supplied September 10 materials, deepseek-v4-flash temporarily routes to V4.1 Flash; the recommended API name is deepseek-flash. Calls through the compatibility name execute the new model and use its peak/off-peak prices. This does not alter the original model’s lifecycle, weights, or historical benchmarks.
The price rows retained for the original model are historical: USD 0.0028 cache-hit input, 0.14 cache-miss input, and 0.28 output per million tokens. They are not current quotes for requests routed through the compatibility alias. V4.1 Flash separately lists off-peak CNY 0.02/1/4 and peak CNY 0.04/2/8 for those three categories. Peak hours are Monday–Friday 09:00–12:00 and 14:00–18:00 Asia/Shanghai.
Original off/high/max thinking modes and the existing default are preserved. Original API support includes OpenAI/Anthropic compatibility, Responses API, JSON Output, Tool Calls, prefix completion, and non-thinking FIM. V4 Flash remains a text model; routing its old API name to a vision model does not give its own weights image support.
Updated scores and routing information come from the supplied DeepSeek September 10, 2026 release materials without independent source verification in this update. Original 0731 public Code Agent notes specified DeepSeek Harness minimal mode, max effort, temperature=1.0 and top_p=0.95; this does not establish identical settings for every historical benchmark. Reference: DeepSeek change log.
No. This update preserves its lifecycle, independent record, and original weights while updating information and scores.
The supplied materials say the compatibility name temporarily routes to V4.1 Flash at its prices. Original V4 Flash USD price rows are historical.
Follow DataLearner on WeChat for AI model updates and research notes.
