Wait Which Model?
← Back to directory
Google DeepMind

Gemma 4

Released Apr 2, 2026 · knowledge cutoff 2025-01

Status
Superseded
Location
United States / United Kingdom
Modality
Multimodal
Context window
262K
Max output
Speed
Price ($/MTok in / out)
$0.10 / $0.34
Cost per task
$0.034
Open weights
Yes
Benchmarks5 of 8 reported

MMLU-Pro

85.2%

GPQA Diamond

84.3%

AIME (latest)

89.2%

Humanity's Last Exam

19.5%

LMArena Elo

1452

SWE-bench Verified

Terminal-Bench 2.1

ARC-AGI-2

Strengths & weaknesses

Strengths

  • Dense rather than sparse — predictable latency and memory, no routing variance to debug
  • Fits on a single high-end consumer GPU once quantized
  • Native system-role support and a switchable thinking mode, both unusual at this size
  • Reliable instruction following, which makes it a good retrieval and pipeline component

Weaknesses

  • Full-precision serving still wants datacenter-class memory
  • No first-party hosted endpoint — you run it yourself or trust a third-party host
  • A component rather than an agent; it isn't built for long autonomous runs

Benchmarks shown are for the 31B Dense variant. LMArena Elo (~1452) and the MMLU-Pro/GPQA/AIME/HLE figures are third-party community-reported, not an official Google table (BenchLM lists GPQA Diamond as 85.7% and HLE as 26.5%, so treat these as approximate). Pricing $0.10/$0.34 per MTok is OpenRouter's listing for google/gemma-4-31b-it — there is no first-party Google API price (Gemma 4 is not on Vertex AI), and other hosts differ (CoreWeave $0.12/$0.35, DeepInfra $0.13/$0.38 per Artificial Analysis). Weights were quietly published March 31, 2026; the formal Google blog announcement followed April 2, 2026. Knowledge cutoff (January 2025) confirmed via Google's official Gemma 4 model card (ai.google.dev).

For developers

API model strings

Not researched

Licence

Not researched

Retirement

No retirement announced

Lineage

No recorded predecessor

News