How many parameters are active?
Approximately 8B on input and 16B on output, out of 552B total. The single active-parameter field records the output-side value.
DeepSeek V4.1 Flash
Also known as: deepseek-flash / deepseek-v4-flash / deepseek-v4-flash-vision-exp
552B native-vision MoE for multilingual tasks; ~8B active on input and ~16B on output (the displayed scalar). Public MIT-licensed weights in 48 safetensors shards totaling about 510.3 GB; 1,000,000-token context and 384K output. Catalog reasoning rating: 4.5/5. Thinking defaults to high; reported evaluations use max. Concurrency is 2,500 with CNY peak/off-peak pricing.
Data sourced primarily from official releases (GitHub, Hugging Face, papers), then benchmark leaderboards, then third-party evaluators. Learn about our data methodology
| Type | Condition | Input | Output |
|---|---|---|---|
| Text | - | ¥0.150/ 1M | ¥0.600/ 1M |
| Type | Condition | Input | Output |
|---|---|---|---|
| Text | - | ¥0.300/ 1M | ¥1.20/ 1M |
| Type | TTL | Write | Read |
|---|---|---|---|
| Text | - | — | ¥0.020/ 1M |
“—” means the modality is not billed in that direction, or the vendor has not published a price for it.
DeepSeek-V4.1-Flash currently shows benchmark results led by HLE (3 / 197, score 63.90), Terminal-Bench 2.1 (1 / 53, score 90.60), CodeForces (1 / 21, score 3471). This page also consolidates core specs, context limits, and API pricing so you can evaluate the model from benchmark results and deployment constraints together.
Want a custom combination? Open the compare tool
Released September 10, 2026, V4.1 Flash is the smallest model in DeepSeek’s new architecture family. Its MoE + Causal-Encoder-Decoder design is asymmetric: 552B total parameters, approximately 8B active on input and 16B active on output. The single active-parameter field on this site records the output-side 16B, not the input-side count. Vision understanding is native rather than an attached vision module. It accepts text and images and produces text.
The model supports multilingual tasks, but the official model card does not publish a complete language count or BCP-47 language list; this catalog therefore leaves the exact count and codes unspecified.
The Hugging Face repository contains 48 safetensors shards totaling about 510.3 GB (475.3 GiB).
Context is 1,000,000 tokens; maximum output is 384,000 tokens. Compared with the previous generation, KV-cache HBM demand is reduced to 1/4 and SSD demand to 1/8, lowering cache costs for long-context and multi-turn agent workloads. Public weights and the technical report are available. The Hugging Face model card states that the repository and model weights are licensed under MIT; API use remains subject to the provider documentation.
The catalog rates its reasoning capability at 4.5/5; the API default remains high effort and the official evaluations use max effort.
Thinking is enabled by default at high effort. Supported levels are low/high/max, plus off. Request mappings: minimal/low → low; medium/high/xhigh → high; max/ultra → max. Low suits simple tasks, high everyday agent tasks, and max complex work.
For the OpenAI format, use reasoning_effort and pass {"thinking":{"type":"enabled"}} in extra_body; disabled turns thinking off. The Anthropic format uses {"reasoning":{"effort":"none/low/high/max"}}, with none disabling thinking. In thinking mode, temperature, presence_penalty and frequency_penalty have no effect; top_p is floored at 0.95. In non-thinking mode, top_p is fixed at 1.0.
Recommended model name: deepseek-flash. OpenAI base URL: https://api.deepseek.com; Anthropic base URL: https://api.deepseek.com/anthropic. Concurrency limit: 2,500. Supports JSON Output, Tool Calls, Responses API, Anthropic API, Chat Prefix Completion (Beta), and FIM Completion (Beta) in non-thinking mode only.
Per million tokens, off-peak cache-hit input/cache-miss input/output cost CNY 0.02 / 1 / 4; peak rates are CNY 0.04 / 2 / 8. Peak hours are Monday–Friday 09:00–12:00 and 14:00–18:00 Asia/Shanghai (UTC+8); all other times are off-peak. These are original CNY rates, not inferred USD equivalents.
GPQA Diamond 90.9; HLE 36.8 (pure-text subset 39.1); Codeforces rating 3471; MathArena Apex 65.6; Terminal-Bench 2.1/3.0/4.0: 90.6/30.0/31.2; DeepSWE v1.1 74.2; ProgramBench 20.3; NL2Repo-Bench 65.4; CyberGym 88.1; SEC-Bench Pro 62.8; ExploitGym 15.3; HLE with tools 63.9; Automation-Bench 54.8; Agents’ Last Exam 31.8; Chartography with tools 78.9; BabyVision with tools 89.6; ZeroBench-main with tools 49.0.
The supplied clarification identifies max effort for these results. Knowledge, mathematics, and competitive-programming results are recorded without tools; agent tasks and explicit w/tools results use tools. Internet access is not inferred. The pure-text HLE subset is separate from full HLE. ExploitGym’s time budget was not supplied, so its results are kept separate from 2h, 6h, and unlimited protocols.
The supplied materials cite older DeepSeek Harness minimal-mode notes with temperature=1.0 and top_p=0.95. Reuse of that entire configuration for V4.1 is not confirmed; only max effort is explicitly confirmed here. Temperature is ineffective in the current thinking API.
The release materials describe an overall advantage over V4 Pro in capability, price, speed, and completion time, not a win on every benchmark: GPQA is 90.9 versus 92.4, and pure-text HLE is 39.1 versus 42.7. Agent gains include Terminal-Bench 3.0 at 30.0 versus 11.8 and DeepSWE at 74.2 versus 62.7. V4.1 Flash also supports native image understanding.
Temporary compatibility names deepseek-v4-flash and deepseek-v4-flash-vision-exp route to V4.1 Flash; old model weights, specifications, and historical scores remain distinct. Requests for deepseek-v4-pro are scheduled to route to V4.1 Flash at its prices from September 14, 2026 at 12:00 Asia/Shanghai until V4.1 Pro launches. This describes a future routing plan, not a lifecycle change applied to the older models.
Supplied launch-day community reports cite approximately 280–500 tokens/s, peaks above 500 at light concurrency, and one comparison of 355 versus 63 tokens/s (about 5.7×), with first-token latency of 178 versus 766 ms. A simple greeting test reported 159.3 tokens/s and 0.8 seconds end-to-end. These are attributed collectively to OrcaRouter and community testing; reproducible environments and individual links were not supplied. They are not standardized official results, fixed throughput specifications, or service guarantees. KV-cache compression may reduce bandwidth pressure but does not establish context-independent speed.
Updated from the supplied September 10, 2026 DeepSeek release materials, without independent source verification in this update. Reference entries: official change log and API pricing. The Hugging Face model card states that the repository and model weights use MIT License; the API documentation is the source for service pricing and routing.
Approximately 8B on input and 16B on output, out of 552B total. The single active-parameter field records the output-side value.
High is the API default. The supplied official evaluations use max effort, with tool use tracked separately.
Per million tokens, off-peak cache-hit input, cache-miss input and output cost CNY 0.02, 1 and 4. Peak rates are CNY 0.04, 2 and 8.
Follow DataLearner on WeChat for AI model updates and research notes.
