Wait Which Model?
← Back to directory
Alibaba (Qwen)

Qwen3.7 Flash

Released Jul 25, 2026 · knowledge cutoff unpublished

Status
Unknown
Location
China
Modality
Multimodal
Context window
1M
Max output
66K
Speed
Price ($/MTok in / out)
$0.03 / $0.13
Cost per task
Open weights
No
Benchmarks0 of 8 reported

MMLU-Pro

GPQA Diamond

SWE-bench Verified

Terminal-Bench 2.1

AIME (latest)

Humanity's Last Exam

LMArena Elo

ARC-AGI-2

Strengths & weaknesses

Strengths

  • Takes video as well as images — the only model in the Qwen3.7 line that does
  • Thinking is switchable rather than always-on, with a large dedicated thinking budget
  • Cheap enough at the base tier to run as a subagent you never bother metering

Weaknesses

  • Latency has a long tail — the slowest percentile runs to well over a minute, which cascades through tool-calling chains
  • Roughly one tool call in eleven comes back malformed, worse than the Plus tier in the same family
  • Transcribes text out of images poorly, so document and screenshot work needs a different model

Announced only as a one-paragraph QwenCloud changelog entry on 2026-07-25, with no technical report, no benchmark table and no parameter count; the OpenRouter listing went live 2026-07-27 under a snapshot named qwen3.7-flash-2026-07-15. None of the site's tracked benchmarks has been published or independently measured — neither Artificial Analysis nor LMArena has run it, so cost per task is unset. Pricing is the sub-32K prompt tier; 32K-256K prompts cost $0.10/$0.40 and 256K-1M cost $0.20/$0.80. Usable input caps at 991,808 tokens in standard mode. Fine-tuning is unsupported and batch inference is Beijing-region only.

For developers

API model strings

Not researched

Licence

Proprietary — weights not released

Retirement

No retirement announced

Lineage

No recorded predecessor

News