GPT-5.6 Luna
Released Jul 9, 2026 · knowledge cutoff 2026-02
- Status
- Superseded
- Location
- United States
- Modality
- Multimodal
- Context window
- 1.1M
- Max output
- 128K
- Speed
- —
- Price ($/MTok in / out)
- $0.20 / $1.20
- Cost per task (medium effort)
- $0.011
- Open weights
- No
Benchmarks6 of 8 reported
GPQA Diamond
SWE-bench Verified
Terminal-Bench 2.1
Humanity's Last Exam
LMArena Elo
ARC-AGI-2
MMLU-Pro
AIME (latest)
Strengths & weaknesses
Strengths
- Answers land almost immediately, so it works for interactive chat and not just overnight batch jobs
- Holds its output format across thousands of calls — long structured-extraction runs come back parseable without repair passes
- Carries the family's full toolkit rather than a cut-down one: image input, tool calling, prompt caching and tunable reasoning effort
Weaknesses
- Quietly switches from quoting sources to summarising them once a document set gets deep, so retrieval work needs chunking rather than one full-window prompt
- Raising reasoning effort stops helping on genuinely unfamiliar problems — the top settings behave much like the low ones there
- Character below the top effort settings is a different model in practice, which makes hands-on reports of it hard to reconcile
Budget tier of the three-model GPT-5.6 family (Sol / Terra / Luna), all released 2026-07-09. Pricing is the post-cut rate: OpenAI cut Luna 80% from $1.00/$6.00 on 2026-07-30. Requests over 272K input tokens bill at 2x input and 1.5x output. Context window, max output and knowledge cutoff are from OpenAI's developer docs. GPQA Diamond 92.3 and Terminal-Bench 2.1 84.7 are from OpenAI's launch table, which also gave SWE-Bench Pro 62.7, Agents' Last Exam 50.3 and MRCR v2 long-context recall 41.3; HLE 37.2 (Artificial Analysis) and ARC-AGI-2 59.5 (ARC Prize) are max-effort figures, matching how Sol is recorded, and SWE-bench Verified 93.0 is third-party (Vals AI). Thinking Machines' Inkling-Small table, run at effort 0.99, reports Terminal-Bench 2.1 82.5 against OpenAI's 84.7 — official preferred — and its GPQA 89.5, HLE 35.6 and ARC-AGI-2 47.6 are the xhigh-effort variant rather than a conflict; its AIME 2026 97.6 is uncorroborated, so AIME stays null. LMArena Elo 1450 is the "gpt-5.6-luna-xhigh" text-arena listing. Artificial Analysis Intelligence Index 51 at max effort, 38 at the medium effort the costPerTask figure comes from.
For developers
API model strings
Not researched
Licence
Proprietary — weights not released
Retirement
No retirement announced
Lineage
No recorded predecessor