Wait Which Model?
← Back to directory
DeepSeek

DeepSeek-V4-Flash-0731

Released Jul 31, 2026 · knowledge cutoff unpublished

Status
Superseded
Location
China
Modality
Text
Context window
1M
Max output
393K
Speed (max effort)
123 tok/s · 1.34s to first answer token
Price ($/MTok in / out)
$0.14 / $0.28
Cost per task (max effort)
$0.027
Open weights
Yes
Benchmarks3 of 8 reported

GPQA Diamond

90.8%

Terminal-Bench 2.1

82.7%

Humanity's Last Exam

36.8%

MMLU-Pro

SWE-bench Verified

AIME (latest)

LMArena Elo

ARC-AGI-2

Strengths & weaknesses

Strengths

  • Same weights and architecture as the April preview — every gain comes from a rebuilt post-training recipe, not new pre-training
  • Clears DeepSeek's own larger V4-Pro preview on every agentic benchmark the lab published
  • Three reasoning-effort levels and a speculative-decoding module carried inside the one checkpoint
  • MIT-licensed and ungated, so on-premise commercial deployment needs no negotiation

Weaknesses

  • Thinks at length before acting on max effort — long agent runs burn far more tokens and wall-clock than the per-call latency suggests
  • Fabricates heavily on knowledge probes, and factual coverage has not moved with the agentic gains
  • Ships without a Jinja chat template — messages go through DeepSeek's own Python encoding scripts instead

The production release of DeepSeek-V4-Flash, superseding April's preview: same 284B/13B MoE and DSpark speculative-decoding module, with all gains from re-post-training. Terminal-Bench 2.1 82.7 is DeepSeek's own figure at max reasoning effort in its minimal-mode harness; Artificial Analysis measures 78.7 on the same benchmark. GPQA Diamond 90.8 and HLE 36.8% (no tools) are Artificial Analysis' measurements — DeepSeek published neither. Also reported by DeepSeek: NL2Repo 54.2, Cybergym 76.7, DeepSWE 54.4, Toolathlon-Verified 70.3, Agents' Last Exam 25.2. MMLU-Pro, SWE-bench Verified, AIME, LMArena and ARC-AGI-2 are unpublished. Artificial Analysis scores it 50 on its Intelligence Index, 10 points above the April preview’s 40. Max output is the 384K DeepSeek recommends for high and max effort; the OpenRouter endpoint caps completions at 65,536.

For developers

API model strings

  • deepseekdeepseek-v4-flash

Licence

MIT License · Permissive · commercial use permitted

Retirement

No retirement announced

Lineage

No recorded predecessor

News