Wait Which Model?
← Back to directory
OpenAI

GPT-5.6 Luna

Released Jul 9, 2026 · knowledge cutoff 2026-02

Status
Superseded
Location
United States
Modality
Multimodal
Context window
1.1M
Max output
128K
Speed
Price ($/MTok in / out)
$0.20 / $1.20
Cost per task (medium effort)
$0.011
Open weights
No
Benchmarks6 of 8 reported

GPQA Diamond

92.3%

SWE-bench Verified

93%

Terminal-Bench 2.1

84.7%

Humanity's Last Exam

37.2%

LMArena Elo

1450

ARC-AGI-2

59.5%

MMLU-Pro

AIME (latest)

Strengths & weaknesses

Strengths

  • Answers land almost immediately, so it works for interactive chat and not just overnight batch jobs
  • Holds its output format across thousands of calls — long structured-extraction runs come back parseable without repair passes
  • Carries the family's full toolkit rather than a cut-down one: image input, tool calling, prompt caching and tunable reasoning effort

Weaknesses

  • Quietly switches from quoting sources to summarising them once a document set gets deep, so retrieval work needs chunking rather than one full-window prompt
  • Raising reasoning effort stops helping on genuinely unfamiliar problems — the top settings behave much like the low ones there
  • Character below the top effort settings is a different model in practice, which makes hands-on reports of it hard to reconcile

Budget tier of the three-model GPT-5.6 family (Sol / Terra / Luna), all released 2026-07-09. Pricing is the post-cut rate: OpenAI cut Luna 80% from $1.00/$6.00 on 2026-07-30. Requests over 272K input tokens bill at 2x input and 1.5x output. Context window, max output and knowledge cutoff are from OpenAI's developer docs. GPQA Diamond 92.3 and Terminal-Bench 2.1 84.7 are from OpenAI's launch table, which also gave SWE-Bench Pro 62.7, Agents' Last Exam 50.3 and MRCR v2 long-context recall 41.3; HLE 37.2 (Artificial Analysis) and ARC-AGI-2 59.5 (ARC Prize) are max-effort figures, matching how Sol is recorded, and SWE-bench Verified 93.0 is third-party (Vals AI). Thinking Machines' Inkling-Small table, run at effort 0.99, reports Terminal-Bench 2.1 82.5 against OpenAI's 84.7 — official preferred — and its GPQA 89.5, HLE 35.6 and ARC-AGI-2 47.6 are the xhigh-effort variant rather than a conflict; its AIME 2026 97.6 is uncorroborated, so AIME stays null. LMArena Elo 1450 is the "gpt-5.6-luna-xhigh" text-arena listing. Artificial Analysis Intelligence Index 51 at max effort, 38 at the medium effort the costPerTask figure comes from.

For developers

API model strings

Not researched

Licence

Proprietary — weights not released

Retirement

No retirement announced

Lineage

No recorded predecessor

News