Qwen3.8-Max-0902
Released Sep 2, 2026 · knowledge cutoff unpublished
- Status
- Unknown
- Location
- China
- Modality
- Multimodal
- Context window
- 1M
- Max output
- 131K
- Speed
- —
- Price ($/MTok in / out)
- $2 / $6
- Cost per task
- —
- Open weights
- No
Benchmarks0 of 10 reported
MMLU-Proretired
GPQA Diamond
SWE-bench Verified
SWE-bench Pro
Terminal-Bench 2.1
AIME (latest)retired
Humanity's Last Exam
LMArena Elo
GDPval-AA v2
ARC-AGI-2
Strengths & weaknesses
Strengths
- Long-horizon autonomous coding is the headline change — engineering-scale projects the July build lost track of now run end to end
- Steadier multi-tool orchestration on office-style cowork tasks, with fewer deliverables abandoned mid-run
- Chart reasoning and document parsing improved, so the same endpoint covers vision-heavy office work
- Same price and API surface as Qwen3.8-Max — the upgrade is a model-string change
Weaknesses
- Still trails Claude Opus 5 on most of the coding and office rows in Qwen's own table — the gap narrowed rather than closed
- Alibaba published no standard reasoning or knowledge benchmarks for this snapshot, so any change outside coding and agent work is uncharacterised
- Served from Alibaba Cloud under Chinese jurisdiction, which some enterprise buyers rule out regardless of capability
Dated snapshot of Qwen3.8-Max (Model Studio alias qwen3.8-max-2026-09-02): same 2.4T-parameter base and 1M window, further post-trained on coding and cowork. Announced on Qwen's X account on 2026-09-02 China time (it appeared on QwenCloud late on 2026-09-01 US time). Qwen's own table, relayed by CellCog, Lookonchain and wccftech, reports Terminal-Bench 3.0 29.0% (Qwen3.8-Max: 11.3%), DeepSWE 1.1 69.3% (56.6%) and NL2Repo-Bench 64.9% (55.9%), plus ProgramBench, SWE-Marathon, CoWorkBench, JobBench and Toolathlon rows on which Opus 5 stays ahead; none of these are tracked keys and Terminal-Bench 3.0 is not comparable with the 2.1 cell, so every benchmark cell is null. The GPQA Diamond 92.6 / HLE 43.6 / Terminal-Bench 2.1 86.6 figures circulating for it are Alibaba's August Qwen3.8-Max numbers, not re-runs on this snapshot. Arena's Code Arena: WebDev debut was #1 at 1691 (22 above the previous Qwen3.8-Max); no Text Arena score yet. Context is Qwen's own 1M claim — the sibling qwen3-8-max entry records Model Studio's 983,616 usable figure, which was not re-verified for this snapshot. Pricing unchanged at $2/$6 per MTok with cache reads $0.25 (implicit) / $0.17 (explicit); Artificial Analysis had not published a separate 0902 listing when checked. predecessorId is null — Qwen's just-got-upgraded framing is lineage, and the earlier build is still served under its own id. Qwen's X post and the Model Studio / QwenCloud pages were not fetchable this session.
For developers
API model strings
- Alibaba Cloud Model Studio
qwen3.8-max-0902
Licence
Proprietary — weights not released
Retirement
No retirement announced
Lineage
No recorded predecessor