Grok 4.5
Released Jul 9, 2026 · knowledge cutoff 2026-02
- Status
- Superseded
- Location
- United States
- Modality
- Multimodal
- Context window
- 500K
- Max output
- —
- Speed
- —
- Price ($/MTok in / out)
- $2 / $6
- Cost per task (high effort)
- $0.35
- Open weights
- No
Benchmarks4 of 8 reported
GPQA Diamond
Humanity's Last Exam
LMArena Elo
ARC-AGI-2
MMLU-Pro
SWE-bench Verified
Terminal-Bench 2.1
AIME (latest)
Strengths & weaknesses
Strengths
- Finishes tasks in fewer steps, so agent runs are cheap in practice rather than just per token
- Fast, decisive answers with very little hedging
- Trained partly on real developer sessions, which shows in routine coding-agent work
Weaknesses
- Hallucinates far more than the model before it — confident, wrong, and in need of grounding for factual work
- Drops off on the hardest and most novel problems, where 'close enough' isn't
- Less public detail on safety and governance than rival labs published
Announced July 8, public release July 9, 2026. GPQA Diamond 93.1%, HLE 40.3%, ARC-AGI-2 52.6%, and LMArena Elo 1468 are third-party (BenchLM/arena.ai) — xAI did not publish these itself, leading with SWE-bench Pro (64.7%) and SWE Marathon (29.0%) instead, so sweBench (Verified) is left null. Max output tokens unconfirmed — xAI's own model page documents the 500K context window and tiered pricing but no output cap; a ~30K figure appears only on secondary blogs, so left null. Knowledge cutoff confirmed as February 1, 2026 in xAI's official docs (docs.x.ai).
For developers
API model strings
Not researched
Licence
Proprietary — weights not released
Retirement
No retirement announced
Lineage
No recorded predecessor