Qwen3.8-Omni-Flash
Released Sep 18, 2026 · knowledge cutoff unpublished
- Status
- Unknown
- Location
- China
- Modality
- Multimodal
- Context window
- 1M
- Max output
- 131K
- Speed
- —
- Price ($/MTok in / out)
- $0.11 / $0.38
- Cost per task
- —
- Open weights
- No
Benchmarks2 of 10 reported
GPQA Diamond
SWE-bench Pro
MMLU-Proretired
SWE-bench Verified
Terminal-Bench 2.1
AIME (latest)retired
Humanity's Last Exam
LMArena Elo
GDPval-AA v2
ARC-AGI-2
Strengths & weaknesses
Strengths
- Reasons over audio and video with an agentic, coarse-to-fine 'selective seeking' pass — decides what to watch or listen to first and spends tokens only on the segments that matter, rather than transcoding a whole stream
- Alibaba's own eval of real meeting audio shows diarization errors falling from 88% to 3% and transcription errors from near 90% to 17% versus the prior Omni line
- Sharply cheaper on audio/video workloads than its predecessor — Alibaba reports a 98% drop in per-hour audio cost and 93% on combined audio-video
Weaknesses
- Text-only output — no generated speech; Alibaba's own Model Studio docs point developers back to Qwen3.5-Omni when a spoken response is needed
- Selective seeking trims tokens at the cost of more tool round-trips, decode time and storage reads, so latency doesn't fall as cleanly as the token count does
- Hosted-only despite sitting on the open-weight Qwen3.8-Flash-Next architecture — Alibaba has published no downloadable weights for the Omni-Flash checkpoint itself
Alibaba's first omni-modal model built around agentic tool use (text, image, audio and video in; text out), launched 2026-09-18 on Qwen Chat, QwenCloud and Model Studio's OpenAI-compatible endpoint. GPQA Diamond 91.0 and SWE-bench Pro 63.3 are Alibaba's own reported figures via launch coverage (official qwen.ai/alibabacloud.com pages were unreachable this session); MMLU-Pro, Terminal-Bench, AIME, HLE, LMArena Elo, GDPval-AA and ARC-AGI-2 were not found reported for this specific checkpoint and are logged to stats-gaps.md. Pricing (~$0.11/$0.38 per MTok) is converted from Alibaba Cloud Model Studio's official CNY rate (0.8/2.7 per million tokens) and matches one reseller's quote, but a second reseller lists $0.15/$0.47 — flagging the conflict rather than picking one; per-hour audio/video surcharge pricing is not broken out here. Context window is Alibaba's marketed 1M figure; QwenCloud's own listing shows a 991K/131K input/output split, used for maxOutput. Coverage repeatedly calls Qwen3.5-Omni-Plus this model's 'predecessor' by name, but Qwen3.5-Omni-Plus is not itself a tracked model in this dataset, so predecessorId stays null (logged to spec-gaps.md). No Artificial Analysis coverage found, hence costPerTask and speed are null; no official retirement page exists yet.
For developers
API model strings
- Alibaba Cloud Model Studio
qwen3.8-omni-flash
Licence
Proprietary — weights not released
Retirement
No retirement announced
Lineage
No recorded predecessor