Wait Which Model?
← Back to directory
Alibaba (Qwen)

Qwen3.8-Max-0902

Released Sep 2, 2026 · knowledge cutoff unpublished

Status
Unknown
Location
China
Modality
Multimodal
Context window
1M
Max output
131K
Speed
Price ($/MTok in / out)
$2 / $6
Cost per task
Open weights
No
Benchmarks0 of 10 reported

MMLU-Proretired

GPQA Diamond

SWE-bench Verified

SWE-bench Pro

Terminal-Bench 2.1

AIME (latest)retired

Humanity's Last Exam

LMArena Elo

GDPval-AA v2

ARC-AGI-2

Strengths & weaknesses

Strengths

  • Long-horizon autonomous coding is the headline change — engineering-scale projects the July build lost track of now run end to end
  • Steadier multi-tool orchestration on office-style cowork tasks, with fewer deliverables abandoned mid-run
  • Chart reasoning and document parsing improved, so the same endpoint covers vision-heavy office work
  • Same price and API surface as Qwen3.8-Max — the upgrade is a model-string change

Weaknesses

  • Still trails Claude Opus 5 on most of the coding and office rows in Qwen's own table — the gap narrowed rather than closed
  • Alibaba published no standard reasoning or knowledge benchmarks for this snapshot, so any change outside coding and agent work is uncharacterised
  • Served from Alibaba Cloud under Chinese jurisdiction, which some enterprise buyers rule out regardless of capability

Dated snapshot of Qwen3.8-Max (Model Studio alias qwen3.8-max-2026-09-02): same 2.4T-parameter base and 1M window, further post-trained on coding and cowork. Announced on Qwen's X account on 2026-09-02 China time (it appeared on QwenCloud late on 2026-09-01 US time). Qwen's own table, relayed by CellCog, Lookonchain and wccftech, reports Terminal-Bench 3.0 29.0% (Qwen3.8-Max: 11.3%), DeepSWE 1.1 69.3% (56.6%) and NL2Repo-Bench 64.9% (55.9%), plus ProgramBench, SWE-Marathon, CoWorkBench, JobBench and Toolathlon rows on which Opus 5 stays ahead; none of these are tracked keys and Terminal-Bench 3.0 is not comparable with the 2.1 cell, so every benchmark cell is null. The GPQA Diamond 92.6 / HLE 43.6 / Terminal-Bench 2.1 86.6 figures circulating for it are Alibaba's August Qwen3.8-Max numbers, not re-runs on this snapshot. Arena's Code Arena: WebDev debut was #1 at 1691 (22 above the previous Qwen3.8-Max); no Text Arena score yet. Context is Qwen's own 1M claim — the sibling qwen3-8-max entry records Model Studio's 983,616 usable figure, which was not re-verified for this snapshot. Pricing unchanged at $2/$6 per MTok with cache reads $0.25 (implicit) / $0.17 (explicit); Artificial Analysis had not published a separate 0902 listing when checked. predecessorId is null — Qwen's just-got-upgraded framing is lineage, and the earlier build is still served under its own id. Qwen's X post and the Model Studio / QwenCloud pages were not fetchable this session.

For developers

API model strings

  • Alibaba Cloud Model Studioqwen3.8-max-0902

Licence

Proprietary — weights not released

Retirement

No retirement announced

Lineage

No recorded predecessor

News