Wait Which Model?
← Back to directory
Alibaba (Qwen)

Qwen3.8-2.4T-A95B

Released Aug 13, 2026 · knowledge cutoff unpublished

Status
Superseded
Location
China
Modality
Text
Context window
262K
Max output
131K
Speed
Price ($/MTok in / out)
$2 / $6
Cost per task
Open weights
Yes
Benchmarks3 of 8 reported

GPQA Diamond

92.6%

Terminal-Bench 2.1

86.6%

Humanity's Last Exam

43.6%

MMLU-Pro

SWE-bench Verified

AIME (latest)

LMArena Elo

ARC-AGI-2

Strengths & weaknesses

Strengths

  • The first Max-class Qwen ever released with downloadable weights — the flagship tier itself, not a distilled sibling
  • Loads straight into vLLM and SGLang from the release, with FP8 weights published alongside the full checkpoint
  • Reasoning runs deep before it answers, which is what carries it through long research and agent tasks

Weaknesses

  • Text-only, and the native window is a quarter of what the closed Max serves — vision and the 1M context stay behind the API
  • Thinking mode cannot be switched off, so a one-line factual question still pays the full reasoning cost
  • Licence pulls large deployments back into a negotiated agreement once revenue or user thresholds are crossed

Qwen's own model card describes this as the post-trained base that Qwen3.8-Max is built on, with Max adding vision and a non-thinking mode — so it is a distinct, reduced artefact rather than the tracked qwen3-8-max entry going open. 2.4T total / 95B active MoE. Context is 262,144 native, extensible to ~1,010,000; max output is the 131,072 final-response cap (reasoning content may run to 262,144). Benchmarks are Alibaba's own self-reported card figures (GPQA Diamond 92.6, Terminal-Bench 2.1 86.6, HLE 43.6) — note Artificial Analysis independently measured the closely-related Max at 92.7 / 81.3 / 41.4, so the Terminal-Bench and HLE claims here should be read as unverified. SWE-bench Pro 67.7 was also reported but is not the Verified variant this site tracks. MMLU-Pro, AIME, LMArena Elo and ARC-AGI-2 are unpublished, and AA had not indexed this variant separately at the time of this check, so costPerTask and speed are null. Pricing is OpenRouter's hosted rate; Alibaba publishes no first-party price for the open-weight release, which is also served by Together AI and Fireworks AI. The licence permits commercial use but requires prominent model-name attribution above 100M MAU or $20M monthly revenue, and a separate agreement for Model-as-a-Service or AI-assistant products above $50M annual revenue.

For developers

API model strings

Not researched

Licence

Qwen3.8-Max License · Restricted · commercial use permitted

Retirement

No retirement announced

Lineage

No recorded predecessor

News