GPT-5.6 Terra
Released Jul 9, 2026 · knowledge cutoff unpublished
- Status
- Unknown
- Location
- United States
- Modality
- Multimodal
- Context window
- 1M
- Max output
- 128K
- Speed
- —
- Price ($/MTok in / out)
- $2 / $12
- Cost per task (medium effort)
- $0.12
- Open weights
- No
Benchmarks1 of 8 reported
Terminal-Bench 2.1
MMLU-Pro
GPQA Diamond
SWE-bench Verified
AIME (latest)
Humanity's Last Exam
LMArena Elo
ARC-AGI-2
Strengths & weaknesses
Strengths
- Behaves like the flagship on ordinary work — the gap only opens on the hardest long-horizon agentic tasks
- Long-context recall holds up close to Sol's, where the budget Luna tier falls away sharply
- Full feature set rather than a cut-down one: image input, tool calling, extended reasoning, prompt caching and structured outputs
Weaknesses
- OpenAI published no per-tier academic benchmarks, so where it really sits against rivals is hard to pin down
- Trails Sol most on exactly the long-horizon agentic work that justifies reaching for a frontier model
Middle tier of the three-model GPT-5.6 family. OpenAI published only agentic suites per tier — Terminal-Bench 2.1 87.4%, Coding Agent Index 77.4, Agents' Last Exam 50.4, MRCR long-context recall 89.6% — and none of the site's tracked academic benchmarks, so those stay null. Artificial Analysis Intelligence Index 55 (vs Sol 59, Luna 51) at 140.4 output tokens/sec. Pricing is OpenAI's list rate, cut from $2.50/$15.00 to $2.00/$12.00 on 2026-07-30 (Luna was cut 80% in the same move). Context window per Artificial Analysis and OpenRouter; max output is tracker-reported, not from OpenAI's docs. Knowledge cutoff not published per tier.
For developers
API model strings
Not researched
Licence
Proprietary — weights not released
Retirement
No retirement announced
Lineage
No recorded predecessor