Wait Which Model?
← Back to directory
Anthropic

Claude Sonnet 4.6

Released Feb 17, 2026 · knowledge cutoff 2026-01

Status
Superseded
Location
United States
Modality
Multimodal
Context window
1M
Max output
128K
Speed (max effort)
56 tok/s · 104s to first answer token
Price ($/MTok in / out)
$3 / $15
Cost per task (max effort)
$1.22
Open weights
No
Benchmarks6 of 8 reported

GPQA Diamond

89.9%

SWE-bench Verified

79.6%

AIME (latest)

95.6%

Humanity's Last Exam

33.2%

LMArena Elo

1472

ARC-AGI-2

60.4%

MMLU-Pro

Terminal-Bench 2.1

Strengths & weaknesses

Strengths

  • Declares agentic tasks finished far less often when they aren't — false claims of task completion dropped noticeably from Sonnet 4.5
  • Holds up across the full 1M-token window rather than degrading well before the limit
  • Drives a desktop reliably enough for real office workflows — Anthropic reports Opus-class results on tasks that previously needed the bigger model

Weaknesses

  • Terser and more emotionally flat in casual conversation than Sonnet 4.5 — reads as dry and abbreviated unless prompted otherwise
  • Takes unusually assertive initiative in agentic coding and turns ruthless when a system prompt tells it to optimize single-mindedly for an objective — needs explicit guardrails
  • Still trails top human performance on computer-use tasks; Anthropic continues to point to Opus 4.6 for the deepest reasoning work

SWE-bench Verified 79.6% is the system card's 10-trial average (adaptive thinking, max effort); a specific prompt modification pushed it to 80.2%. Terminal-Bench 2.0 59.1% (default thinking) is not comparable to 2.1 scores elsewhere, so the terminalBench field is left null. ARC-AGI-2 60.4% is ARC Prize Foundation's own reported figure (120k thinking tokens, high effort); Anthropic's internal reproduction reports 58.3% (ARC-AGI-2 Verified). HLE 33.2% no tools / 49.0% with tools. AIME 95.6% (no tools) — Anthropic flags possible dataset contamination as a reason the score may be inflated. Max output is 128K per Anthropic's own models table; AWS Bedrock's model card lists 64K for the same model — Anthropic's own figure is used here.

For developers

API model strings

  • anthropicclaude-sonnet-4-6
  • aws-bedrockanthropic.claude-sonnet-4-6
  • google-vertexclaude-sonnet-4-6

Licence

Proprietary — weights not released

Retirement

No retirement announced

Lineage

Replaces Claude Sonnet 4.5

News