Wait Which Model?
← Back to directory
xAI

Grok 4.5

Released Jul 9, 2026 · knowledge cutoff 2026-02

Status
Superseded
Location
United States
Modality
Multimodal
Context window
500K
Max output
Speed
Price ($/MTok in / out)
$2 / $6
Cost per task (high effort)
$0.35
Open weights
No
Benchmarks4 of 8 reported

GPQA Diamond

93.1%

Humanity's Last Exam

40.3%

LMArena Elo

1468

ARC-AGI-2

52.6%

MMLU-Pro

SWE-bench Verified

Terminal-Bench 2.1

AIME (latest)

Strengths & weaknesses

Strengths

  • Finishes tasks in fewer steps, so agent runs are cheap in practice rather than just per token
  • Fast, decisive answers with very little hedging
  • Trained partly on real developer sessions, which shows in routine coding-agent work

Weaknesses

  • Hallucinates far more than the model before it — confident, wrong, and in need of grounding for factual work
  • Drops off on the hardest and most novel problems, where 'close enough' isn't
  • Less public detail on safety and governance than rival labs published

Announced July 8, public release July 9, 2026. GPQA Diamond 93.1%, HLE 40.3%, ARC-AGI-2 52.6%, and LMArena Elo 1468 are third-party (BenchLM/arena.ai) — xAI did not publish these itself, leading with SWE-bench Pro (64.7%) and SWE Marathon (29.0%) instead, so sweBench (Verified) is left null. Max output tokens unconfirmed — xAI's own model page documents the 500K context window and tiered pricing but no output cap; a ~30K figure appears only on secondary blogs, so left null. Knowledge cutoff confirmed as February 1, 2026 in xAI's official docs (docs.x.ai).

For developers

API model strings

Not researched

Licence

Proprietary — weights not released

Retirement

No retirement announced

Lineage

No recorded predecessor

News