Wait Which Model?
← Back to directory
Google DeepMind

Gemini 3.7 Flash

Released Aug 13, 2026 · knowledge cutoff 2026-03

Status
Frontier
Location
United States / United Kingdom
Modality
Multimodal
Context window
1M
Max output
66K
Speed (medium effort)
274 tok/s · 4.22s to first answer token
Price ($/MTok in / out)
$0.75 / $3.75
Cost per task (medium effort)
$0.26
Open weights
No
Benchmarks3 of 8 reported

GPQA Diamond

94.5%

Terminal-Bench 2.1

85.8%

Humanity's Last Exam

53.6%

MMLU-Pro

SWE-bench Verified

AIME (latest)

LMArena Elo

ARC-AGI-2

Strengths & weaknesses

Strengths

  • Adapts to roadblocks mid-task instead of stalling — clarifies intent and follows multi-step instructions more reliably than 3.6 Flash
  • Low-effort mode lands close to the previous model's high-effort quality, so routine tasks run at a fraction of the cost and wait
  • Cleaner multi-step tool calling and planning means fewer retries and less manual oversight in agent pipelines

Weaknesses

  • Time to first answer token stretches to roughly 10 seconds median at high effort, with worst-case waits near a minute — the effort dial needs deliberate tuning, not the default
  • Launch coverage is narrowly focused on coding and agent benchmarks — general chat and creative behavior are largely uncharacterized so far

Shipped three weeks after 3.6 Flash as an algorithmic refinement of the same base model (Google's own framing: 'the core reasoning foundation'), not a new pretraining run; Google replaced 3.6 Flash with 3.7 Flash in the Gemini app's model picker on launch day (9to5Google), which is the explicit replacement evidence behind predecessorId here. Google's own headline coding-benchmark table reports DeepSWE v1.1 65.3% (was 49.0%), FrontierCode 1.1 Main 43.6% (was 34.4%), AutomationBench 30.4% (was 17.0%) and WebDev Arena Elo 1588 (was 1538) — none of these map to this site's tracked keys. GPQA Diamond 94.5% and HLE (Verified) 53.6% are Artificial Analysis' high-effort measurements; Terminal-Bench 2.1 85.8% is Google's own reported figure (vs 78.0% for 3.6 Flash). costPerTask and speed use AA's medium-effort listing per this site's convention ($0.26/task, 273.6 tok/s, 4.22s to first answer token); AA's high-effort listing instead shows $0.40/task at the introductory price. Pricing is the introductory rate ($0.75/$3.75 per 1M input/output tokens) in effect through 2026-12-31; it rises to $1.50/$7.50 on 2027-01-01.

For developers

API model strings

  • googlegemini-3.7-flash

Licence

Proprietary — weights not released

Retirement

No retirement announced

Lineage

Replaces Gemini 3.6 Flash

News