Wait Which Model?
← Back to directory
xAI

Grok 4.7

Released Sep 21, 2026 · knowledge cutoff 2026-05

Status
Unknown
Location
United States
Modality
Multimodal
Context window
500K
Max output
Speed
Price ($/MTok in / out)
$2 / $6
Cost per task
Open weights
No
Benchmarks0 of 10 reported

MMLU-Proretired

GPQA Diamond

SWE-bench Verified

SWE-bench Pro

Terminal-Bench 2.1

AIME (latest)retired

Humanity's Last Exam

LMArena Elo

GDPval-AA v2

ARC-AGI-2

Strengths & weaknesses

Strengths

  • Works longer on hard tasks and re-checks its own output before declaring done — the RL run was weighted toward problems that take hours, with self-verification as the explicit training target
  • Far easier to talk through architecture and in-depth explanations with than 4.6, which early Cursor users had found unusable for that kind of design work
  • Understands the Grok Bot harness natively, so agent runs need less scaffolding to get well-formed tool use
  • Serves at 4.6's speed despite the larger base model, so the capability step costs nothing in latency

Weaknesses

  • Burns roughly twice the tokens of Grok 4.6 on the same job in early Cursor use, so the unchanged per-token price understates what a task costs — several developers called it a sidegrade on value
  • Has picked up a verbose, presumptuous house style — the concise structured findings of 4.5 gave way to forced cadences, per early Cursor feedback, and Artificial Analysis flags it as very verbose on its index run
  • Crossing the 200K-token prompt threshold reprices the whole request at double rate, and the fast variant doubles it again — long-context or low-latency runs cost far more than the headline

Announced 2026-09-21 after Musk's 2026-09-02 'ten days' promise slipped twice; xAI says it uses a new, larger base model than Grok 4.6 (the 2.1T-parameter figure is Musk's pre-release claim, not on the card) with a longer RL run on multi-hour tasks. xAI's card reports only CursorBench 4.0 46.3%, DeepSWE v1.1 71.0% (high effort), EEBench 64.0%, AA Briefcase v1.1 1,657, Terminal-Bench 4.0 38.0%, Harvey Legal Agent 19.6% and HealthBench Professional 56.7% — none map to a tracked key, and Terminal-Bench 4.0 is not comparable with the 2.1 this site records, so terminalBench is null; Artificial Analysis independently measures Terminal-Bench 4.0 at 26%. AA scores it 46 on its Intelligence Index at both high (the API default) and xhigh effort, but had published no cost-per-task or speed figures as of 2026-09-22, and rates it 1694 (high) / 1695 (xhigh) on GDPval-AA v2.1 — a rescaled successor to the v2 this site's gdpvalAA key records (Grok 4.6 is 1632 on v2.1 against its recorded 1753 on v2), so that cell stays null. GPQA Diamond, HLE, SWE-bench Verified/Pro, AIME, MMLU-Pro and ARC-AGI-2 are unpublished by xAI and AA; arena.ai's board (last updated 2026-09-13) has no listing. Knowledge cutoff May 2026 per docs.x.ai's models page. Pricing $2/$6 with cached input $0.50 applies below a 200K-token prompt; at or above it every token in the request bills at $4/$12 ($1.00 cached), and the fast variant costs double for twice the output speed. Effort levels low/medium/high/xhigh; batch API unsupported; xAI publishes no max-output figure. predecessorId is null: the card says only that it is 'served at the same price and speed as Grok 4.6', and grok-4.6 remains listed in the API. xAI also claims it is its strongest model on refusals and jailbreak resistance. The company rebranded from xAI to SpaceXAI on 2026-07-06; companies.json still carries the old name.

For developers

API model strings

  • xaigrok-4.7

Licence

Proprietary — weights not released

Retirement

No retirement announced

Lineage

No recorded predecessor

News