GLM-5.3-Flash
Released Aug 26, 2026 · knowledge cutoff unpublished
- Status
- Unknown
- Location
- China
- Modality
- Multimodal
- Context window
- 1M
- Max output
- 131K
- Speed (max effort)
- 50 tok/s · 1.49s to first answer token
- Price ($/MTok in / out)
- $0.15 / $0.50
- Cost per task (max effort)
- $0.090
- Open weights
- Yes
Benchmarks2 of 10 reported
Terminal-Bench 2.1
LMArena Elo
MMLU-Proretired
GPQA Diamond
SWE-bench Verified
SWE-bench Pro
AIME (latest)retired
Humanity's Last Exam
GDPval-AA v2
ARC-AGI-2
Strengths & weaknesses
Strengths
- Combines sparse and linear attention in one open checkpoint, cutting attention compute and KV-cache cost several-fold while keeping the million-token window usable rather than degrading toward the edges
- Handles native multimodal tool calls — text, image, video and file input in the same turn — without a vision adapter bolted onto a text-only base
- Ran anonymously on OpenRouter as "Ox Alpha" for a week under heavy public load before Zhipu claimed it, serving at real production volume before anyone knew whose model it was
Weaknesses
- Streams roughly half as fast as its own flagship sibling GLM-5.3 — the Flash name doesn't guarantee low latency, so time-sensitive use still needs its own benchmarking
- Default responses aren't automation-ready for strict structured output — needs explicit response-format constraints where some rivals return clean JSON out of the box
- Reasoning is always-on with only a low/high/max budget dial, not an off switch — every call carries at least some thinking overhead
320B-total/18B-active MoE, Zhipu's first natively multimodal model in the GLM-5 line. Terminal-Bench 2.1 (84.3%) is Z.ai's own reported figure; GPQA Diamond, HLE, SWE-bench Verified, AIME and ARC-AGI-2 are left null (LMArena Elo was listed later, below) — Z.ai's own table reports non-standard benchmarks instead (DeepSWE v1.1 63.4%, AutomationBench 48.8%, Toolathlon Verified 78.4%), and a single-source claim of an Artificial Analysis GPQA Diamond (91.2%) and HLE (55.3%) score could not be corroborated across independent sources on re-check. costPerTask ($0.09) and speed (49.8 tok/s, 1.49s to first answer token) are Artificial Analysis' max-effort measurements — the model has no medium-effort variant (only low/high/max). Pricing is Z.ai's standard $0.15/$0.50 rate; a 50%-off launch promotion runs through 2026-09-09. predecessorId is left null since Zhipu's own materials frame this as a new capability tier rather than an explicit replacement for GLM-5.2 or GLM-5.3. LMArena Elo 1475 is the "glm-5.3-flash" text-arena listing (arena.ai, 2026-09-13 update, rank 29). On 2026-09-18 Z.ai added GLM-5.3-FlashX (`glm-5.3-flashx`), a hosted high-speed tier of the same weights at up to 200 tok/s for $0.37/$1.25 — no separate benchmarks or checkpoint, so it is recorded here rather than as its own entry; FlashX is not on the GLM Coding Plan.
For developers
API model strings
- zhipu
glm-5.3-flash
Licence
MIT License · Permissive · commercial use permitted
Retirement
No retirement announced
Lineage
No recorded predecessor