DataLearner logo
DE

DeepSeek-V4-Flash-Vision-Exp

PreviewMultimodal modelTool useDeepSeek FlashDeepSeek V4

DeepSeek V4 Flash Vision Experimental

Release date: 2026-08-21Updated: 2026-08-21 21:16:57.596299
Live demoGitHubHugging FaceCompare
Parameters
Not disclosed
Context length
1M
Chinese support
Not supported
Reasoning ability

DeepSeek released DeepSeek-V4-Flash-Vision-Exp as an experimental API vision model on August 21, 2026. It supports text and image input, a 1M-token context window, 384K maximum output, tool calls, and multiple compatible APIs. Each image uses at most 384 input tokens. DeepSeek also published six peak/off-peak price rows and nine agent/vision benchmark scores. Parameter count, downloadable weights, license, and knowledge cutoff remain undisclosed.

Data sourced primarily from official releases (GitHub, Hugging Face, papers), then benchmark leaderboards, then third-party evaluators. Learn about our data methodology

DeepSeek-V4-Flash-Vision-Exp

Model basics

Reasoning traces
Supported
Thinking modes
Thinking Level · High (Default)Standard ModeThinking Level · LowThinking Level · Max
Context length
1M tokens
Max output length
384K tokens
Model type
Multimodal model
Modality (in / out)
Text, Image → Text
Release date
2026-08-21
Model file size
No data
MoE architecture
No
Total params / Active params
No data / N/A
Knowledge cutoff
No data
DeepSeek-V4-Flash-Vision-Exp

Open source & experience

Code license
No data
Weights license
No data
GitHub repo
GitHub link unavailable
Hugging Face
Hugging Face link unavailable
Live demo
No live demo
DeepSeek-V4-Flash-Vision-Exp

Official resources

Paper
DataLearnerAI blog
No blog post yet
DeepSeek-V4-Flash-Vision-Exp

API details

API speed
No data
💡Default unit: $/1M tokens. If vendors use other units, follow their published pricing.
off_peak
TypeConditionInputOutput
Text-$0.220/ 1M$0.660/ 1M
peak
TypeConditionInputOutput
Text-$0.440/ 1M$1.32/ 1M
Cache PricingPrompt Cache
TypeTTLWriteRead
Text--$0.0070/ 1M
DeepSeek-V4-Flash-Vision-Exp

Benchmark Results

DeepSeek-V4-Flash-Vision-Exp currently shows benchmark results led by Terminal-Bench 2.1 (11 / 44, score 83.90), NL2Repo-Bench (3 / 8, score 57.70), DeepSWE (12 / 27, score 59.30). This page also consolidates core specs, context limits, and API pricing so you can evaluate the model from benchmark results and deployment constraints together.

Thinking

AI Agent - Tool Usage

2 evaluations
Benchmark / mode
Score
Rank/total
83.90
11 / 44
25.70
7 / 8

Coding and Software Engineer

3 evaluations
Benchmark / mode
Score
Rank/total
DSBench-Hard
MaxTools
63.60
1 / 1
DeepSWE
MaxTools
59.30
12 / 27
57.70
3 / 8

Agent Level Benchmark

2 evaluations
Benchmark / mode
Score
Rank/total
ApexBench
MaxTools
36.50
1 / 1
27.30
6 / 11

Multimodal Understanding

2 evaluations
Benchmark / mode
Score
Rank/total
Chartography
MaxTools
64.30
1 / 1
35
2 / 3

Compare with other models

No curated comparisons for this model yet.

Want a custom combination? Open the compare tool

DeepSeek-V4-Flash-Vision-Exp

Publisher

DeepSeek V4 Flash Vision Experimental

Model Overview

Official API release

DeepSeek-V4-Flash-Vision-Exp is an experimental multimodal vision-understanding model released on the DeepSeek API platform on August 21, 2026. The API model ID is deepseek-v4-flash-vision-exp. It accepts text and images and produces text for image description, screenshot reading, chart analysis, and vision-dependent agent tasks. DeepSeek says its pure-text capabilities are on par with the official DeepSeek-V4-Flash model, but does not state that the two share the same architecture, parameter count, or weights.


Specifications and API capabilities

The official context window is 1,000,000 tokens and maximum output is 384,000 tokens. Both non-thinking and thinking modes are supported, with thinking enabled by default. The API supports JSON output, tool calls, the Responses API, an Anthropic-compatible API, and beta Chat Prefix Completion; FIM is not supported. The published concurrency limit is 2,500.


Vision inputs

JPEG, PNG, GIF, and WebP images can be supplied as Base64 data URLs, external URLs, or Files API file_id/file_data references. Low detail downsizes to 512×512, while high and original preserve source detail; auto currently behaves like original. Images are automatically resized and each image consumes at most 384 input tokens. Multiple images are counted independently, and image tokens use ordinary input-token pricing.

External image URLs are limited to 8,192 characters and the HTTP request body to 48 MiB. A Base64 or external image can be up to 32 MiB, while a Files API image can be up to 64 MiB. A request can contain up to 600 images. The total image limit is 64 MiB without file_id, or up to 200 MiB including file_id. Each side is normally limited to 8,192 pixels, reduced to 4,096 pixels when 15 or more images are used.


Official evaluations

DeepSeek reports: Terminal Bench 2.1 83.9, NL2Repo 57.7, DeepSWE 59.3, DSBench-Hard 63.6, AutomationBench (Public) 25.7, ApexBench (Pass@1) 36.5, Agents' Last Exam 27.3, Chartography 64.3, and ZeroBench (Pass@5) 35.0. Public text-only Code Agent tasks used DeepSeek Harness minimal, max effort, top_p=0.95, and temperature=1.0. DeepSeek describes a substantial improvement on vision-dependent agent benchmarks, approaching Opus-4.8.


API pricing

All prices are USD per one million tokens. Off-peak prices are $0.007 for cache-hit input, $0.22 for cache-miss input, and $0.66 for output. Peak prices are $0.014, $0.44, and $1.32, respectively. Peak periods are 01:00–04:00 UTC and 06:00–10:00 UTC each day; all other times are off-peak. Image tokens are charged as input tokens.


Disclosure boundary

DeepSeek has not published this experimental model's total or active parameter count, standalone architecture details, knowledge cutoff, downloadable weights, or license. DataLearner leaves those fields unknown instead of inheriting values from another V4 model.


Official sources

DeepSeek-V4-Flash-Vision-Exp

FAQ

Is DeepSeek-V4-Flash-Vision-Exp officially available?

Yes. DeepSeek released it on the API platform on August 21, 2026 under the exact model ID deepseek-v4-flash-vision-exp. It is labeled experimental rather than generally available.

What image formats and input methods are supported?

It supports JPEG, PNG, GIF, and WebP through Base64 data URLs, external URLs, or Files API file references. Chat Completions, Responses API, and Anthropic-compatible request formats are documented.

What are the context window and maximum output?

The official context window is 1,000,000 tokens and the maximum output is 384,000 tokens.

How are images tokenized and billed?

Images are automatically resized and each image consumes at most 384 input tokens. Multiple images are counted independently, and those tokens are billed at the applicable input-token rate.

How much does the API cost?

Off-peak prices per 1M tokens are $0.007 cache-hit input, $0.22 cache-miss input, and $0.66 output. Peak prices are $0.014, $0.44, and $1.32. Peak periods are 01:00–04:00 UTC and 06:00–10:00 UTC daily.

Are model weights or parameter counts available?

No. DeepSeek has not published the model's parameter counts, standalone architecture details, downloadable weights, license, or knowledge cutoff.

DataLearner on WeChat

Follow DataLearner on WeChat for AI model updates and research notes.

DataLearner WeChat QR code