Claude Opus 5
Released Jul 24, 2026 · knowledge cutoff 2026-05
- Status
- Frontier
- Location
- United States
- Modality
- Multimodal
- Context window
- 1M
- Max output
- 128K
- Speed (medium effort)
- 54 tok/s · 6.69s to first answer token
- Price ($/MTok in / out)
- $5 / $25
- Cost per task (medium effort)
- $0.62
- Open weights
- No
Benchmarks4 of 8 reported
GPQA Diamond
SWE-bench Verified
Humanity's Last Exam
ARC-AGI-2
MMLU-Pro
Terminal-Bench 2.1
AIME (latest)
LMArena Elo
Strengths & weaknesses
Strengths
- Verifies its own work unprompted, delegates readily, and will widen a task's scope on its own judgement
- Holds the thread across long multi-step work and pins down vague requirements before building
- Reads a large unfamiliar codebase carefully before it changes anything
- Solves novel-pattern problems that no earlier model could, rather than recalling them
Weaknesses
- Verbose and pushy — it narrates constantly, and default output length responds to prompting more than to the effort setting
- Early users describe it as neurotic and apologetic, and less willing to follow a prescribed workflow than Opus 4.8
- Easy to misconfigure: effort and skill settings change its behavior more than the version bump does
HLE 56.3% no tools / 64.7% with tools. SWE-bench Verified 96.0% is Anthropic's own system-card figure (five-trial average); a third-party aggregator (vals.ai) cites 97.0%. GPQA Diamond varies 93.4-93.7% across trackers; 93.7% used. ARC-AGI-2 90.4% (max effort) and ARC-AGI-3 30.2% (high effort, a separate untracked scale) per ARC Prize (arcprize.org). Max output is 128K on the synchronous API, up to 300K via the Message Batches API. AIME not published — Anthropic appears to be moving to newer evals (IMO 2026, Frontier-Bench v0.1) instead.
For developers
API model strings
- anthropic
claude-opus-5 - aws-bedrock
anthropic.claude-opus-5
Licence
Proprietary — weights not released
Retirement
No retirement announced
Lineage
Replaces Claude Opus 4.8