DeepSeek-V4-Flash-0731
Released Jul 31, 2026 · knowledge cutoff unpublished
- Status
- Superseded
- Location
- China
- Modality
- Text
- Context window
- 1M
- Max output
- 393K
- Speed (max effort)
- 123 tok/s · 1.34s to first answer token
- Price ($/MTok in / out)
- $0.14 / $0.28
- Cost per task (max effort)
- $0.027
- Open weights
- Yes
Benchmarks3 of 8 reported
GPQA Diamond
Terminal-Bench 2.1
Humanity's Last Exam
MMLU-Pro
SWE-bench Verified
AIME (latest)
LMArena Elo
ARC-AGI-2
Strengths & weaknesses
Strengths
- Same weights and architecture as the April preview — every gain comes from a rebuilt post-training recipe, not new pre-training
- Clears DeepSeek's own larger V4-Pro preview on every agentic benchmark the lab published
- Three reasoning-effort levels and a speculative-decoding module carried inside the one checkpoint
- MIT-licensed and ungated, so on-premise commercial deployment needs no negotiation
Weaknesses
- Thinks at length before acting on max effort — long agent runs burn far more tokens and wall-clock than the per-call latency suggests
- Fabricates heavily on knowledge probes, and factual coverage has not moved with the agentic gains
- Ships without a Jinja chat template — messages go through DeepSeek's own Python encoding scripts instead
The production release of DeepSeek-V4-Flash, superseding April's preview: same 284B/13B MoE and DSpark speculative-decoding module, with all gains from re-post-training. Terminal-Bench 2.1 82.7 is DeepSeek's own figure at max reasoning effort in its minimal-mode harness; Artificial Analysis measures 78.7 on the same benchmark. GPQA Diamond 90.8 and HLE 36.8% (no tools) are Artificial Analysis' measurements — DeepSeek published neither. Also reported by DeepSeek: NL2Repo 54.2, Cybergym 76.7, DeepSWE 54.4, Toolathlon-Verified 70.3, Agents' Last Exam 25.2. MMLU-Pro, SWE-bench Verified, AIME, LMArena and ARC-AGI-2 are unpublished. Artificial Analysis scores it 50 on its Intelligence Index, 10 points above the April preview’s 40. Max output is the 384K DeepSeek recommends for high and max effort; the OpenRouter endpoint caps completions at 65,536.
For developers
API model strings
- deepseek
deepseek-v4-flash
Licence
MIT License · Permissive · commercial use permitted
Retirement
No retirement announced
Lineage
No recorded predecessor