Gemma 4
Released Apr 2, 2026 · knowledge cutoff 2025-01
- Status
- Superseded
- Location
- United States / United Kingdom
- Modality
- Multimodal
- Context window
- 262K
- Max output
- —
- Speed
- —
- Price ($/MTok in / out)
- $0.10 / $0.34
- Cost per task
- $0.034
- Open weights
- Yes
Benchmarks5 of 8 reported
MMLU-Pro
GPQA Diamond
AIME (latest)
Humanity's Last Exam
LMArena Elo
SWE-bench Verified
Terminal-Bench 2.1
ARC-AGI-2
Strengths & weaknesses
Strengths
- Dense rather than sparse — predictable latency and memory, no routing variance to debug
- Fits on a single high-end consumer GPU once quantized
- Native system-role support and a switchable thinking mode, both unusual at this size
- Reliable instruction following, which makes it a good retrieval and pipeline component
Weaknesses
- Full-precision serving still wants datacenter-class memory
- No first-party hosted endpoint — you run it yourself or trust a third-party host
- A component rather than an agent; it isn't built for long autonomous runs
Benchmarks shown are for the 31B Dense variant. LMArena Elo (~1452) and the MMLU-Pro/GPQA/AIME/HLE figures are third-party community-reported, not an official Google table (BenchLM lists GPQA Diamond as 85.7% and HLE as 26.5%, so treat these as approximate). Pricing $0.10/$0.34 per MTok is OpenRouter's listing for google/gemma-4-31b-it — there is no first-party Google API price (Gemma 4 is not on Vertex AI), and other hosts differ (CoreWeave $0.12/$0.35, DeepInfra $0.13/$0.38 per Artificial Analysis). Weights were quietly published March 31, 2026; the formal Google blog announcement followed April 2, 2026. Knowledge cutoff (January 2025) confirmed via Google's official Gemma 4 model card (ai.google.dev).
For developers
API model strings
Not researched
Licence
Not researched
Retirement
No retirement announced
Lineage
No recorded predecessor