Gemini 3.7 Flash
Released Aug 13, 2026 · knowledge cutoff 2026-03
- Status
- Frontier
- Location
- United States / United Kingdom
- Modality
- Multimodal
- Context window
- 1M
- Max output
- 66K
- Speed (medium effort)
- 274 tok/s · 4.22s to first answer token
- Price ($/MTok in / out)
- $0.75 / $3.75
- Cost per task (medium effort)
- $0.26
- Open weights
- No
Benchmarks3 of 8 reported
GPQA Diamond
Terminal-Bench 2.1
Humanity's Last Exam
MMLU-Pro
SWE-bench Verified
AIME (latest)
LMArena Elo
ARC-AGI-2
Strengths & weaknesses
Strengths
- Adapts to roadblocks mid-task instead of stalling — clarifies intent and follows multi-step instructions more reliably than 3.6 Flash
- Low-effort mode lands close to the previous model's high-effort quality, so routine tasks run at a fraction of the cost and wait
- Cleaner multi-step tool calling and planning means fewer retries and less manual oversight in agent pipelines
Weaknesses
- Time to first answer token stretches to roughly 10 seconds median at high effort, with worst-case waits near a minute — the effort dial needs deliberate tuning, not the default
- Launch coverage is narrowly focused on coding and agent benchmarks — general chat and creative behavior are largely uncharacterized so far
Shipped three weeks after 3.6 Flash as an algorithmic refinement of the same base model (Google's own framing: 'the core reasoning foundation'), not a new pretraining run; Google replaced 3.6 Flash with 3.7 Flash in the Gemini app's model picker on launch day (9to5Google), which is the explicit replacement evidence behind predecessorId here. Google's own headline coding-benchmark table reports DeepSWE v1.1 65.3% (was 49.0%), FrontierCode 1.1 Main 43.6% (was 34.4%), AutomationBench 30.4% (was 17.0%) and WebDev Arena Elo 1588 (was 1538) — none of these map to this site's tracked keys. GPQA Diamond 94.5% and HLE (Verified) 53.6% are Artificial Analysis' high-effort measurements; Terminal-Bench 2.1 85.8% is Google's own reported figure (vs 78.0% for 3.6 Flash). costPerTask and speed use AA's medium-effort listing per this site's convention ($0.26/task, 273.6 tok/s, 4.22s to first answer token); AA's high-effort listing instead shows $0.40/task at the introductory price. Pricing is the introductory rate ($0.75/$3.75 per 1M input/output tokens) in effect through 2026-12-31; it rises to $1.50/$7.50 on 2027-01-01.
For developers
API model strings
- google
gemini-3.7-flash
Licence
Proprietary — weights not released
Retirement
No retirement announced
Lineage
Replaces Gemini 3.6 Flash