Wait Which Model?
← Back to directory
Google DeepMind

Gemini 3.5 Flash-Lite

Released Jul 21, 2026 · knowledge cutoff 2026-03

Status
Superseded
Location
United States / United Kingdom
Modality
Multimodal
Context window
1M
Max output
66K
Speed
Price ($/MTok in / out)
$0.30 / $2.50
Cost per task
$0.086
Open weights
No
Benchmarks4 of 8 reported

GPQA Diamond

83.8%

Terminal-Bench 2.1

54%

Humanity's Last Exam

17.5%

LMArena Elo

1459

MMLU-Pro

SWE-bench Verified

AIME (latest)

ARC-AGI-2

Strengths & weaknesses

Strengths

  • Very high sustained token throughput — built for real-time, high-volume work
  • Holds a persona and multi-turn instructions far better than earlier Lite models
  • Drives a desktop and works over long inputs better than its size implies

Weaknesses

  • Time to first token can feel unresponsive in chat even though it streams quickly once started
  • Small-model ceiling: it executes well but won't reason its way out of an unfamiliar problem

Model card reports SWE-Bench Pro 54.2%, OSWorld-Verified 74.0%, MLE-Bench 39.2%, CharXiv 74.5% no tools / 76.5% with tools; Google published none of the tracked benchmark keys at launch. GPQA Diamond 83.8% and HLE 17.5% are Artificial Analysis-measured figures relayed by BenchLM — AA's own model page publishes only the composite Intelligence Index (36) — so treat both as third-party. LMArena Elo 1459 is the 'gemini-3.5-flash-lite' text-arena listing (rank 43, arena.ai). Also rolling out inside Google Search.

For developers

API model strings

Not researched

Licence

Proprietary — weights not released

Retirement

No retirement announced

Lineage

No recorded predecessor

News