GPT-6 Astra
Released Sep 3, 2026 · knowledge cutoff 2026-04
- Status
- Frontier
- Location
- United States
- Modality
- Multimodal
- Context window
- 1.1M
- Max output
- 128K
- Speed
- —
- Price ($/MTok in / out)
- $10 / $50
- Cost per task (medium effort)
- $0.75
- Open weights
- No
Benchmarks3 of 10 reported
GPQA Diamond
LMArena Elo
ARC-AGI-2
MMLU-Proretired
SWE-bench Verified
SWE-bench Pro
Terminal-Bench 2.1
AIME (latest)retired
Humanity's Last Exam
GDPval-AA v2
Strengths & weaknesses
Strengths
- Drives a desktop end to end — locating UI targets, filling spreadsheets, building sites — and finishes computer-use tasks markedly faster than Sol did rather than just more accurately
- Holds recall across the full million-token window instead of degrading past a few hundred thousand, so long-document and whole-repo work doesn't need chunking
- Asynchronous tool calls let it fire a call and keep reasoning on independent work while it waits, so long agent runs stall less between steps
- Far harder to hijack through content it reads: indirect prompt-injection attempts land at roughly a third of Sol's rate
Weaknesses
- Its written reasoning is less monitorable than Sol's — OpenAI says so explicitly, so chain-of-thought is a weaker debugging and oversight handle exactly where the model is most autonomous
- The API still defaults to low reasoning effort while the marketing describes the top end, so the model most people call is not the model in the launch table until they set reasoning.effort
- Gated at every layer — trusted-access first, enterprise workspaces off until an admin opts in, cybersecurity capability behind the Daybreak programme — so it can't simply be picked up and trialled
- Behavior is a moving target this early: no hands-on reviews of voice, verbosity or agentic stamina exist yet, and independent leaderboards haven't placed it
OpenAI's first model designated Critical for cybersecurity under its Preparedness Framework, launched via staged rollout (trusted-access enterprises first, then ChatGPT Plus/Pro/Business/Enterprise and the wider API), hence availability "restricted". Context 1,050,000 tokens (922K max input) and 128K output, cutoff 2026-04-30, and $10/$50 per MTok ($1 cached input, $12.50 cache writes) are from OpenAI's own model docs; prompts over 272K input tokens bill at 2x input/cache and 1.5x output for the whole request, and a fast tier costs 2x ($20/$100). GPQA Diamond 96.0 is from OpenAI's launch table; the rest of that table uses benchmarks this site doesn't track (Terminal-Bench 4.0 57.7% — not the tracked 2.1 metric — DeepSWE v1.1 74.1%, FrontierMath Tier 4 v2 97.6%, OSWorld 2.0 72.6%, ScreenSpot-Pro 92.7%, BrowseComp 91.5%, ExploitBench 100%, MRCR v2 8-needle 100% at 256K-512K and 96.3% at 512K-1M). HLE is left null because the only published figure, 57.2%, is with tools and so isn't comparable to the no-tools scores recorded here. OpenAI has not published SWE-bench Verified since February 2026 and Astra is absent from the SWE-bench Pro leaderboard; the arena.ai listing and ARC Prize's ARC-AGI-2 result arrived after launch (below). ARC Prize measured ARC-AGI-3 (62.7% on its standard harness, 99.9% via a provider-adapter harness ARC Prize warns is not comparable to human testing) but published no ARC-AGI-2 figure. Artificial Analysis scores it 61 on the Intelligence Index at max effort (level with GPT-5.6 Sol, 5 below Claude Fable 5.1) at $1.67 per task, and $0.75 per task at medium effort, the figure recorded here; AA reports no speed measurements yet, and flags a ~80 Elo regression on GDPval-AA v2 against Sol without publishing the absolute rating. ARC-AGI-2 95.0% is ARC Prize-verified at max effort (xhigh 93.3%, high/medium 92.1%, low 85.4%), published 2026-09-02 at arcprize.org/results/openai-gpt-6-astra; cost per task not stated on that page. LMArena Elo 1480 (±12) is the "gpt-6-astra-max" text-arena listing (arena.ai, 2026-09-13 update, rank 24).
For developers
API model strings
- openai
gpt-6-astra
Licence
Proprietary — weights not released
Retirement
No retirement announced
Lineage
No recorded predecessor