Wait Which Model?
← Back to directory
DeepSeek

DeepSeek-V4-Flash-Vision-Exp

Released Aug 21, 2026 · knowledge cutoff unpublished

Status
Deprecated
Location
China
Modality
Multimodal
Context window
1M
Max output
393K
Speed (max effort)
109 tok/s · 20s to first answer token
Price ($/MTok in / out)
$0.22 / $0.66
Cost per task
Open weights
Yes
Benchmarks1 of 10 reported

Terminal-Bench 2.1

83.9%

MMLU-Proretired

GPQA Diamond

SWE-bench Verified

SWE-bench Pro

AIME (latest)retired

Humanity's Last Exam

LMArena Elo

GDPval-AA v2

ARC-AGI-2

Strengths & weaknesses

Strengths

  • Adds image input to V4-Flash without giving up its text behaviour — agent, reasoning and knowledge tasks land where the text-only 0731 build does
  • Every image is capped at 384 tokens, so screenshot-heavy agent loops stay cheap even at high volume
  • Competent on layout, charts and screenshots inside a tool loop — the multimodal agent gains are where DeepSeek concentrated the continued training
  • MIT weights followed the API by ten days, with vLLM serving configs and a reference inference implementation published alongside

Weaknesses

  • Every image is downscaled to roughly 800x800 with no high-resolution mode — small glyphs such as receipts, stack traces and dense spreadsheets degrade badly
  • Misreads exact chart values and fails caption-checking far more often than rivals — anything legal, medical or financial needs verification against the source
  • Labelled experimental by DeepSeek itself — an Exp model whose behaviour may shift before a production vision release

DeepSeek's first multimodal V4 model: a vision encoder and aligner on the 284B-total/13B-active V4-Flash MoE with continued training, live on the API from 2026-08-21 and open-sourced under MIT on 2026-08-31 (FP8 with FP4 expert weights, roughly 168 GB). Terminal-Bench 2.1 83.9 is DeepSeek's own figure from its eleven-benchmark launch table (no logs, sample sizes or confidence intervals published); the table also reports DeepSWE 59.3, Agents' Last Exam 27.3, ApexBench 36.5, ZeroBench pass@5 35.0 and Chartography 64.3 against Opus 4.8, none tracked here. Artificial Analysis scores it 51 at max effort (one above V4-Flash-0731) and folds GPQA Diamond and HLE into that index, but the per-eval figures were not readable, so those cells stay null. Speed (109 tok/s, 19.54 s to first answer token) is AA's max-effort measurement; AA's per-task cost was not visible, only the $235.89 whole-index total, so costPerTask is null rather than derived. Pricing is the off-peak rate that has applied to the whole V4-Flash line since 2026-08-16 — peak hours (01:00-04:00 and 06:00-10:00 UTC) bill double at $0.44/$1.32, the rate AA lists; images bill as tokens, capped at 384 per image; cache hits $0.007/MTok. Max output follows the 384K DeepSeek recommends for V4-Flash at high and max effort. predecessorId is null: DeepSeek frames it as an experimental sibling of V4-Flash-0731, not a replacement. api-docs.deepseek.com and the HF card were not fetchable this session; figures are cross-checked across DeepSeek's X post, OpenRouter, AA and launch coverage. Retired 2026-09-10: DeepSeek's change log states "the previous-generation models V4 Flash and V4 Flash Vision Exp have been retired" with the launch of DeepSeek-V4.1-Flash, whose vision is native rather than bolted on; the `deepseek-v4-flash-vision-exp` id is still accepted but temporarily routed to V4.1-Flash and billed at Flash rates. Weights remain on Hugging Face.

For developers

API model strings

  • deepseekdeepseek-v4-flash-vision-exp

Licence

MIT License · Permissive · commercial use permitted

Retirement

Sep 10, 2026

Lineage

No recorded predecessor

News