Qwen3.8-2.4T-A95B
Released Aug 13, 2026 · knowledge cutoff unpublished
- Status
- Superseded
- Location
- China
- Modality
- Text
- Context window
- 262K
- Max output
- 131K
- Speed
- —
- Price ($/MTok in / out)
- $2 / $6
- Cost per task
- —
- Open weights
- Yes
Benchmarks3 of 8 reported
GPQA Diamond
Terminal-Bench 2.1
Humanity's Last Exam
MMLU-Pro
SWE-bench Verified
AIME (latest)
LMArena Elo
ARC-AGI-2
Strengths & weaknesses
Strengths
- The first Max-class Qwen ever released with downloadable weights — the flagship tier itself, not a distilled sibling
- Loads straight into vLLM and SGLang from the release, with FP8 weights published alongside the full checkpoint
- Reasoning runs deep before it answers, which is what carries it through long research and agent tasks
Weaknesses
- Text-only, and the native window is a quarter of what the closed Max serves — vision and the 1M context stay behind the API
- Thinking mode cannot be switched off, so a one-line factual question still pays the full reasoning cost
- Licence pulls large deployments back into a negotiated agreement once revenue or user thresholds are crossed
Qwen's own model card describes this as the post-trained base that Qwen3.8-Max is built on, with Max adding vision and a non-thinking mode — so it is a distinct, reduced artefact rather than the tracked qwen3-8-max entry going open. 2.4T total / 95B active MoE. Context is 262,144 native, extensible to ~1,010,000; max output is the 131,072 final-response cap (reasoning content may run to 262,144). Benchmarks are Alibaba's own self-reported card figures (GPQA Diamond 92.6, Terminal-Bench 2.1 86.6, HLE 43.6) — note Artificial Analysis independently measured the closely-related Max at 92.7 / 81.3 / 41.4, so the Terminal-Bench and HLE claims here should be read as unverified. SWE-bench Pro 67.7 was also reported but is not the Verified variant this site tracks. MMLU-Pro, AIME, LMArena Elo and ARC-AGI-2 are unpublished, and AA had not indexed this variant separately at the time of this check, so costPerTask and speed are null. Pricing is OpenRouter's hosted rate; Alibaba publishes no first-party price for the open-weight release, which is also served by Together AI and Fireworks AI. The licence permits commercial use but requires prominent model-name attribution above 100M MAU or $20M monthly revenue, and a separate agreement for Model-as-a-Service or AI-assistant products above $50M annual revenue.
For developers
API model strings
Not researched
Licence
Qwen3.8-Max License · Restricted · commercial use permitted
Retirement
No retirement announced
Lineage
No recorded predecessor