Wait Which Model?
← Back to directory
Anthropic

Claude Opus 4.8

Released May 28, 2026 · knowledge cutoff 2026-01

Status
Superseded
Location
United States
Modality
Multimodal
Context window
1M
Max output
128K
Speed
Price ($/MTok in / out)
$5 / $25
Cost per task (max effort)
$1.80
Open weights
No
Benchmarks5 of 8 reported

GPQA Diamond

93.6%

SWE-bench Verified

88.6%

Humanity's Last Exam

49.8%

LMArena Elo

1484

ARC-AGI-2

72.1%

MMLU-Pro

Terminal-Bench 2.1

AIME (latest)

Strengths & weaknesses

Strengths

  • Flags its own uncertainty and asks rather than guessing — far less confident nonsense
  • The effort setting genuinely changes token spend, so quick and deep work share one model
  • More expressive and varied on creative work than the 4.x releases before it

Weaknesses

  • The extra hedging is friction when you want a confident answer to a low-stakes question
  • Can get caught in self-correction loops, re-checking work that was already right
  • Clarifying questions interrupt fully autonomous runs that were meant to need no input

HLE 49.8% no tools / 57.9% with tools. LMArena Elo corrected to 1484 (current Opus 4.8 Thinking listing, July 2026); an earlier 1510 figure did not match the live leaderboard. ARC-AGI-2 72.1% added July 2026 once ARC Prize published verified results: 72.08% at high effort (71.67% medium, 62.22% low).

For developers

API model strings

Not researched

Licence

Proprietary — weights not released

Retirement

No retirement announced

Lineage

No recorded predecessor

News