Wait Which Model?
← Back to directory
Google DeepMind

Gemini 3.5 Flash

Released May 19, 2026 · knowledge cutoff 2025-01

Status
Superseded
Location
United States / United Kingdom
Modality
Multimodal
Context window
1M
Max output
66K
Speed
Price ($/MTok in / out)
$1.50 / $9
Cost per task (high effort)
$0.59
Open weights
No
Benchmarks5 of 8 reported

SWE-bench Verified

81%

Terminal-Bench 2.1

76.2%

Humanity's Last Exam

40.2%

LMArena Elo

1481

ARC-AGI-2

72.1%

MMLU-Pro

GPQA Diamond

AIME (latest)

Strengths & weaknesses

Strengths

  • The fastest capable option of its moment — iteration speed beats waiting for a better answer
  • Near-Pro coding quality at Flash latency
  • Reasoning context carries across turns, so multi-turn agent loops don't re-derive everything

Weaknesses

  • Instruction-following is its weakest axis; tightly constrained output formats slip
  • Output cap low enough to truncate large generated files
  • Default effort level shifted at launch, quietly changing behavior for existing prompts

Announced at Google I/O 2026; Terminal-Bench 2.1 76.2%. HLE 40.2 and ARC-AGI-2 72.1 are Google's own figures on the official Gemini 3.5 Flash model card (verified 2026-07-29). SWE-bench Verified 81.0 came from I/O reporting and is NOT on that model card, which reports SWE-Bench Pro (Public) 55.1% instead — treat 81.0 as unconfirmed. LMArena Elo 1481 at launch window.

For developers

API model strings

Not researched

Licence

Proprietary — weights not released

Retirement

No retirement announced

Lineage

No recorded predecessor

News