Nova 2 Pro
Released Dec 2, 2025 · knowledge cutoff unpublished
- Status
- Superseded
- Location
- United States
- Modality
- Multimodal
- Context window
- 1M
- Max output
- 66K
- Speed (medium effort)
- 119 tok/s · 13s to first answer token
- Price ($/MTok in / out)
- $1.25 / $10
- Cost per task
- —
- Open weights
- No
Benchmarks4 of 8 reported
MMLU-Pro
GPQA Diamond
SWE-bench Verified
AIME (latest)
Terminal-Bench 2.1
Humanity's Last Exam
LMArena Elo
ARC-AGI-2
Strengths & weaknesses
Strengths
- Native web grounding and a built-in code interpreter mean it can pull live information and validate its own output without external tool wiring
- Configurable extended-thinking depth (low/medium/high) lets one model span quick replies and systematic multi-step reasoning rather than shipping separate fast/deep variants
- Strong multi-document and video reasoning — Amazon's own comparison table shows it equal-or-better on 15 of 19 benchmarks against Gemini 2.5 Pro
Weaknesses
- Still preview-only eight months after announcement — early access is limited to Amazon Nova Forge customers, with no general Bedrock console/API listing or official model-card page as of this check
- Trails GPT-5.1 on roughly half its benchmark comparisons (8 of 16) despite leading Gemini 2.5 Pro on most of the same set
- SWE-bench Verified score depends heavily on scaffolding — 61.5% unassisted vs. 70.0% with Amazon's own inference-time scaling and agentic scaffolding, an 8.5-point gap that makes head-to-head reads of the headline number unreliable without knowing the setup
Announced 2025-12-02 at AWS re:Invent alongside Nova 2 Lite, Nova 2 Sonic and Nova 2 Omni. Benchmark figures (MMLU-Pro 81.6, GPQA-Diamond 81.4, AIME 2025 92.3, SWE-Bench Verified 61.5/70.0) are Amazon's own, from the official Nova 2 Family technical report (assets.amazon.science/.../nova-2-0-technical-report2.pdf, Tables 1 and 4); the 70.0% SWE-bench figure uses 'inference time scaling and internal Agentic scaffolding' per the report's own footnote, so the unassisted 61.5% is used as the tracked value with the scaffolded figure noted here rather than in the field. Terminal-Bench in the same report (41.3%) is version 1.0, not the tracked 2.1 metric, so left null. HLE, ARC-AGI-2 and lmarenaElo were not found on any official report or independent leaderboard. Context window (1M) and max output (65,536) are from AWS's official Nova 2 user guide ('all models support up to 1 million tokens of context... generate up to 65,536 tokens'); Artificial Analysis separately lists a narrower 256K context window and $1.25/$10.00 pricing for its tracked preview build, which is the source used for pricing and speed here since AWS's own Bedrock pricing page and model-card system (which returns 404 for Nova 2 Pro specifically, unlike Nova 2 Lite/Premier) have not yet published either for this model — flagging the context-window conflict rather than picking one. availability is 'restricted': as of this check Nova 2 Pro has no Bedrock model-card page and remains preview/early-access-only for Nova Forge customers, absent from the official Nova 2 model list that does cover Lite, Sonic and Multimodal Embeddings. No official per-provider apiIds string exists yet, so left []. predecessorId left null: the Nova 2 technical report calls Nova Premier 'our former flagship' while comparing it to Nova 2 Lite specifically, not Pro, so it doesn't rise to an explicit statement that Nova 2 Pro replaces Nova Premier — logged to spec-gaps.md. Closed weights; no license record; no published knowledge cutoff.
For developers
API model strings
Not researched
Licence
Proprietary — weights not released
Retirement
No retirement announced
Lineage
No recorded predecessor