Qwen3.7 Flash
Released Jul 25, 2026 · knowledge cutoff unpublished
- Status
- Unknown
- Location
- China
- Modality
- Multimodal
- Context window
- 1M
- Max output
- 66K
- Speed
- —
- Price ($/MTok in / out)
- $0.03 / $0.13
- Cost per task
- —
- Open weights
- No
Benchmarks0 of 8 reported
MMLU-Pro
GPQA Diamond
SWE-bench Verified
Terminal-Bench 2.1
AIME (latest)
Humanity's Last Exam
LMArena Elo
ARC-AGI-2
Strengths & weaknesses
Strengths
- Takes video as well as images — the only model in the Qwen3.7 line that does
- Thinking is switchable rather than always-on, with a large dedicated thinking budget
- Cheap enough at the base tier to run as a subagent you never bother metering
Weaknesses
- Latency has a long tail — the slowest percentile runs to well over a minute, which cascades through tool-calling chains
- Roughly one tool call in eleven comes back malformed, worse than the Plus tier in the same family
- Transcribes text out of images poorly, so document and screenshot work needs a different model
Announced only as a one-paragraph QwenCloud changelog entry on 2026-07-25, with no technical report, no benchmark table and no parameter count; the OpenRouter listing went live 2026-07-27 under a snapshot named qwen3.7-flash-2026-07-15. None of the site's tracked benchmarks has been published or independently measured — neither Artificial Analysis nor LMArena has run it, so cost per task is unset. Pricing is the sub-32K prompt tier; 32K-256K prompts cost $0.10/$0.40 and 256K-1M cost $0.20/$0.80. Usable input caps at 991,808 tokens in standard mode. Fine-tuning is unsupported and batch inference is Beijing-region only.
For developers
API model strings
Not researched
Licence
Proprietary — weights not released
Retirement
No retirement announced
Lineage
No recorded predecessor