← Back to directory
Anthropic
Claude 3.7 Sonnet
Released Feb 24, 2025 · knowledge cutoff 2024-10
- Status
- Superseded
- Location
- United States
- Modality
- Multimodal
- Context window
- 200K
- Max output
- 128K
- Speed
- —
- Price ($/MTok in / out)
- $3 / $15
- Cost per task
- —
- Open weights
- No
Benchmarks6 of 8 reported
MMLU-Pro
80.7%
GPQA Diamond
84.8%
SWE-bench Verified
70.3%
AIME (latest)
61.3%
Humanity's Last Exam
8.9%
LMArena Elo
1350
Terminal-Bench 2.1
—
ARC-AGI-2
—
Strengths & weaknesses
Strengths
- You choose per request whether it answers instantly or thinks first — same model, two modes
- Thinking is visible, so a wrong assumption can be caught before it acts on it
- Carries a plan across many tool calls instead of restarting its reasoning each turn
Weaknesses
- Over-eager: widens scope, refactors files you didn't mention, edits tests until they pass
- Quality tracks the thinking budget, so a small budget quietly gives you a worse model
GPQA/SWE-bench figures use extended thinking.
For developers
API model strings
Not researched
Licence
Proprietary — weights not released
Retirement
No retirement announced
Lineage
No recorded predecessor