← Back to directory
Alibaba (Qwen)
Qwen3.7 Max
Released May 19, 2026 · knowledge cutoff unpublished
- Status
- Superseded
- Location
- China
- Modality
- Text
- Context window
- 1M
- Max output
- 66K
- Speed
- —
- Price ($/MTok in / out)
- $2.50 / $7.50
- Cost per task
- $1.03
- Open weights
- No
Benchmarks4 of 8 reported
MMLU-Pro
89.6%
GPQA Diamond
92.4%
SWE-bench Verified
80.4%
Humanity's Last Exam
41.4%
Terminal-Bench 2.1
—
AIME (latest)
—
LMArena Elo
—
ARC-AGI-2
—
Strengths & weaknesses
Strengths
- Long-horizon autonomy — runs for tens of hours and thousands of tool calls without hand-holding
- Precise instruction following, including multilingual and formatting constraints
- Heavy caching discounts make repetitive agent loops cheap in practice
Weaknesses
- Very verbose — a real agentic session costs far more than the per-token rate suggests
- Prose is over-engineered and reads machine-written
- Closed weights, a break from the rest of the Qwen line
SWE-bench Pro 60.6; Terminal-Bench 2.0 69.7. SWE-bench Verified 80.4 and HLE 41.4 lab/W&B-reported; MMLU-Pro 89.6 is third-party (LLM-Stats). Alibaba reports HMMT 97.1 instead of AIME.
For developers
API model strings
Not researched
Licence
Proprietary — weights not released
Retirement
No retirement announced
Lineage
No recorded predecessor