Wait Which Model?
← Back to directory
Google DeepMind

Gemini 3.8 Flash

Released Sep 2, 2026 · knowledge cutoff 2026-03

Status
Frontier
Location
United States / United Kingdom
Modality
Multimodal
Context window
1M
Max output
66K
Speed (medium effort)
312 tok/s · 6.44s to first answer token
Price ($/MTok in / out)
$0.75 / $3.75
Cost per task (medium effort)
$0.41
Open weights
No
Benchmarks5 of 10 reported

GPQA Diamond

95.3%

Terminal-Bench 2.1

89.4%

Humanity's Last Exam

54.9%

LMArena Elo

1494

GDPval-AA v2

1545

MMLU-Proretired

SWE-bench Verified

SWE-bench Pro

AIME (latest)retired

ARC-AGI-2

Strengths & weaknesses

Strengths

  • Takes smaller reasoning steps, calls tools iteratively and re-checks its own work — long-horizon coding runs go further before it declares victory than 3.7 Flash did
  • Low effort answers in well under a second while high effort deliberates for a quarter of a minute — one endpoint spans quick chat and patient agent work
  • Markedly harder to prompt-inject than the 3.x Flash releases before it on Gray Swan's robustness suite, so it can be pointed at untrusted web content with less babysitting

Weaknesses

  • Burns roughly 40% more output tokens than 3.7 Flash on the same work by design — Google itself points efficiency-first workloads back to 3.7 Flash, which stays on sale
  • Time to first answer token climbs to about 13 seconds at high effort, and Google's own model card warns of occasional slowness and timeouts
  • Rejects the minimal thinking level outright, and at low effort long code answers can hit the generation ceiling before the visible answer completes

Third Flash release in six weeks; the model card says it is 'based on Gemini 3.7 Flash' and Google's docs say 3.7 Flash 'remains fully supported' with no shutdown date on the deprecations page, so predecessorId stays null (lineage, not replacement). Terminal-Bench 2.1 89.4% (Terminus 2 harness) and HLE-Verified 54.9% (the 1,811-item verified set, no tools) are Google's own self-computed figures from the evaluation-methodology PDF, which lists 3.7 Flash at 85.8% and 53.6%; GDPval-AA v2 1545 is Artificial Analysis' high-effort listing as cited in that table; GPQA Diamond 95.3% is AA's high-effort measurement, not Google-reported; LMArena Elo 1494 is the 'gemini-3.8-flash-high' text-arena listing (rank 8, preliminary) on launch day. Google published no SWE-bench Verified, SWE-bench Pro, MMLU-Pro, AIME or ARC-AGI-2 figures — the SWE-bench Pro 61.6% and Terminal-Bench 90.8% quoted by DataCamp appear nowhere in Google's material and are not used. Google's table also reports DeepSWE v1.1 73.7% (was 65.3%), Terminal-Bench 4.0 19.1%, Vals Finance Agent v2 61.4%, Harvey Legal 10.0%, OSWorld 2.0 59.0% partial and LABBench2 86.2%, none of which map to tracked keys. costPerTask and speed are AA's medium-effort variant per site convention ($0.41/task, 312.3 tok/s, 6.44 s to first answer token); high effort scores AA Intelligence Index 59 (up 3 on 3.7 Flash) at $0.58/task, 302.1 tok/s and 13.3 s, low effort 52 at $0.24/task and 0.70 s. Pricing is the introductory rate through 2026-12-31, rising to $1.50/$7.50 on 2027-01-01 (batch half price, context caching $0.075). Model card notes some domains' knowledge is limited to January 2025 despite the March 2026 cutoff. Released alongside Gemini 3.8 Flash Cyber, the same base model with looser cyber mitigations, gated to the Fairwind Program.

For developers

API model strings

  • googlegemini-3.8-flash

Licence

Proprietary — weights not released

Retirement

No retirement announced

Lineage

No recorded predecessor

News