Grok 4.7
Released Sep 21, 2026 · knowledge cutoff 2026-05
- Status
- Unknown
- Location
- United States
- Modality
- Multimodal
- Context window
- 500K
- Max output
- —
- Speed
- —
- Price ($/MTok in / out)
- $2 / $6
- Cost per task
- —
- Open weights
- No
Benchmarks0 of 10 reported
MMLU-Proretired
GPQA Diamond
SWE-bench Verified
SWE-bench Pro
Terminal-Bench 2.1
AIME (latest)retired
Humanity's Last Exam
LMArena Elo
GDPval-AA v2
ARC-AGI-2
Strengths & weaknesses
Strengths
- Works longer on hard tasks and re-checks its own output before declaring done — the RL run was weighted toward problems that take hours, with self-verification as the explicit training target
- Far easier to talk through architecture and in-depth explanations with than 4.6, which early Cursor users had found unusable for that kind of design work
- Understands the Grok Bot harness natively, so agent runs need less scaffolding to get well-formed tool use
- Serves at 4.6's speed despite the larger base model, so the capability step costs nothing in latency
Weaknesses
- Burns roughly twice the tokens of Grok 4.6 on the same job in early Cursor use, so the unchanged per-token price understates what a task costs — several developers called it a sidegrade on value
- Has picked up a verbose, presumptuous house style — the concise structured findings of 4.5 gave way to forced cadences, per early Cursor feedback, and Artificial Analysis flags it as very verbose on its index run
- Crossing the 200K-token prompt threshold reprices the whole request at double rate, and the fast variant doubles it again — long-context or low-latency runs cost far more than the headline
Announced 2026-09-21 after Musk's 2026-09-02 'ten days' promise slipped twice; xAI says it uses a new, larger base model than Grok 4.6 (the 2.1T-parameter figure is Musk's pre-release claim, not on the card) with a longer RL run on multi-hour tasks. xAI's card reports only CursorBench 4.0 46.3%, DeepSWE v1.1 71.0% (high effort), EEBench 64.0%, AA Briefcase v1.1 1,657, Terminal-Bench 4.0 38.0%, Harvey Legal Agent 19.6% and HealthBench Professional 56.7% — none map to a tracked key, and Terminal-Bench 4.0 is not comparable with the 2.1 this site records, so terminalBench is null; Artificial Analysis independently measures Terminal-Bench 4.0 at 26%. AA scores it 46 on its Intelligence Index at both high (the API default) and xhigh effort, but had published no cost-per-task or speed figures as of 2026-09-22, and rates it 1694 (high) / 1695 (xhigh) on GDPval-AA v2.1 — a rescaled successor to the v2 this site's gdpvalAA key records (Grok 4.6 is 1632 on v2.1 against its recorded 1753 on v2), so that cell stays null. GPQA Diamond, HLE, SWE-bench Verified/Pro, AIME, MMLU-Pro and ARC-AGI-2 are unpublished by xAI and AA; arena.ai's board (last updated 2026-09-13) has no listing. Knowledge cutoff May 2026 per docs.x.ai's models page. Pricing $2/$6 with cached input $0.50 applies below a 200K-token prompt; at or above it every token in the request bills at $4/$12 ($1.00 cached), and the fast variant costs double for twice the output speed. Effort levels low/medium/high/xhigh; batch API unsupported; xAI publishes no max-output figure. predecessorId is null: the card says only that it is 'served at the same price and speed as Grok 4.6', and grok-4.6 remains listed in the API. xAI also claims it is its strongest model on refusals and jailbreak resistance. The company rebranded from xAI to SpaceXAI on 2026-07-06; companies.json still carries the old name.
For developers
API model strings
- xai
grok-4.7
Licence
Proprietary — weights not released
Retirement
No retirement announced
Lineage
No recorded predecessor