DeepSeek-V4-Flash-Vision-Exp
Released Aug 21, 2026 · knowledge cutoff unpublished
- Status
- Deprecated
- Location
- China
- Modality
- Multimodal
- Context window
- 1M
- Max output
- 393K
- Speed (max effort)
- 109 tok/s · 20s to first answer token
- Price ($/MTok in / out)
- $0.22 / $0.66
- Cost per task
- —
- Open weights
- Yes
Benchmarks1 of 10 reported
Terminal-Bench 2.1
MMLU-Proretired
GPQA Diamond
SWE-bench Verified
SWE-bench Pro
AIME (latest)retired
Humanity's Last Exam
LMArena Elo
GDPval-AA v2
ARC-AGI-2
Strengths & weaknesses
Strengths
- Adds image input to V4-Flash without giving up its text behaviour — agent, reasoning and knowledge tasks land where the text-only 0731 build does
- Every image is capped at 384 tokens, so screenshot-heavy agent loops stay cheap even at high volume
- Competent on layout, charts and screenshots inside a tool loop — the multimodal agent gains are where DeepSeek concentrated the continued training
- MIT weights followed the API by ten days, with vLLM serving configs and a reference inference implementation published alongside
Weaknesses
- Every image is downscaled to roughly 800x800 with no high-resolution mode — small glyphs such as receipts, stack traces and dense spreadsheets degrade badly
- Misreads exact chart values and fails caption-checking far more often than rivals — anything legal, medical or financial needs verification against the source
- Labelled experimental by DeepSeek itself — an Exp model whose behaviour may shift before a production vision release
DeepSeek's first multimodal V4 model: a vision encoder and aligner on the 284B-total/13B-active V4-Flash MoE with continued training, live on the API from 2026-08-21 and open-sourced under MIT on 2026-08-31 (FP8 with FP4 expert weights, roughly 168 GB). Terminal-Bench 2.1 83.9 is DeepSeek's own figure from its eleven-benchmark launch table (no logs, sample sizes or confidence intervals published); the table also reports DeepSWE 59.3, Agents' Last Exam 27.3, ApexBench 36.5, ZeroBench pass@5 35.0 and Chartography 64.3 against Opus 4.8, none tracked here. Artificial Analysis scores it 51 at max effort (one above V4-Flash-0731) and folds GPQA Diamond and HLE into that index, but the per-eval figures were not readable, so those cells stay null. Speed (109 tok/s, 19.54 s to first answer token) is AA's max-effort measurement; AA's per-task cost was not visible, only the $235.89 whole-index total, so costPerTask is null rather than derived. Pricing is the off-peak rate that has applied to the whole V4-Flash line since 2026-08-16 — peak hours (01:00-04:00 and 06:00-10:00 UTC) bill double at $0.44/$1.32, the rate AA lists; images bill as tokens, capped at 384 per image; cache hits $0.007/MTok. Max output follows the 384K DeepSeek recommends for V4-Flash at high and max effort. predecessorId is null: DeepSeek frames it as an experimental sibling of V4-Flash-0731, not a replacement. api-docs.deepseek.com and the HF card were not fetchable this session; figures are cross-checked across DeepSeek's X post, OpenRouter, AA and launch coverage. Retired 2026-09-10: DeepSeek's change log states "the previous-generation models V4 Flash and V4 Flash Vision Exp have been retired" with the launch of DeepSeek-V4.1-Flash, whose vision is native rather than bolted on; the `deepseek-v4-flash-vision-exp` id is still accepted but temporarily routed to V4.1-Flash and billed at Flash rates. Weights remain on Hugging Face.
For developers
API model strings
- deepseek
deepseek-v4-flash-vision-exp
Licence
MIT License · Permissive · commercial use permitted
Retirement
Sep 10, 2026
Lineage
No recorded predecessor