Wait Which Model?
← Back to directory
Anthropic

Claude 3.7 Sonnet

Released Feb 24, 2025 · knowledge cutoff 2024-10

Status
Superseded
Location
United States
Modality
Multimodal
Context window
200K
Max output
128K
Speed
Price ($/MTok in / out)
$3 / $15
Cost per task
Open weights
No
Benchmarks6 of 8 reported

MMLU-Pro

80.7%

GPQA Diamond

84.8%

SWE-bench Verified

70.3%

AIME (latest)

61.3%

Humanity's Last Exam

8.9%

LMArena Elo

1350

Terminal-Bench 2.1

ARC-AGI-2

Strengths & weaknesses

Strengths

  • You choose per request whether it answers instantly or thinks first — same model, two modes
  • Thinking is visible, so a wrong assumption can be caught before it acts on it
  • Carries a plan across many tool calls instead of restarting its reasoning each turn

Weaknesses

  • Over-eager: widens scope, refactors files you didn't mention, edits tests until they pass
  • Quality tracks the thinking budget, so a small budget quietly gives you a worse model

GPQA/SWE-bench figures use extended thinking.

For developers

API model strings

Not researched

Licence

Proprietary — weights not released

Retirement

No retirement announced

Lineage

No recorded predecessor

News