Gemini 3.8 Flash Cyber
Released Sep 2, 2026 · knowledge cutoff unpublished
- Status
- Unknown
- Location
- United States / United Kingdom
- Modality
- Multimodal
- Context window
- —
- Max output
- —
- Speed
- —
- Price ($/MTok in / out)
- — / —
- Cost per task
- —
- Open weights
- No
Benchmarks0 of 10 reported
MMLU-Proretired
GPQA Diamond
SWE-bench Verified
SWE-bench Pro
Terminal-Bench 2.1
AIME (latest)retired
Humanity's Last Exam
LMArena Elo
GDPval-AA v2
ARC-AGI-2
Strengths & weaknesses
Strengths
- Closes the loop from finding a vulnerability to a verified, deployable patch — Chrome's security team got more correct fixes out of it than from far larger commercial models
- Ships with looser cyber-offence mitigations than the public 3.8 Flash, so vulnerability reproduction and pentest-recall work the general model refuses goes through
- Hunts beyond C/C++ across some twenty languages at Flash-class cost — cheap enough to run as a continuous scanning loop rather than a one-off audit
Weaknesses
- Reachable only through the Fairwind Program — background-checked organisations, access limited to internal security teams, no redistribution — so it cannot be trialled casually
- Tuned toward patching rather than exploitation by design, so it is a weaker fit for red-team exploit development than for defensive triage
- No published general-capability scores, specs or pricing — behaviour outside vulnerability work is undocumented
Cybersecurity variant of Gemini 3.8 Flash — Google says both releases are 'powered by the same foundational intelligence', with Cyber shipping 'a more permissive set of mitigations for cybersecurity' — available only through the new Fairwind Program (trusted governments and national cyber authorities, critical-infrastructure operators and core software maintainers; background-checked, 650+ partner organisations at launch), hence availability 'restricted'. Served as a managed model on Gemini Enterprise Agent Platform (zero data retention) and inside Google's CodeMender agent; no public API id, pricing, context window, output limit or knowledge cutoff has been published for it, so those cells are null rather than copied from 3.8 Flash. Google reports only cyber-specific evals: CWE-Bench pass@1 47.2% against 47.8% for the leading frontier model, a success rate above 70% on an internal 20-language vulnerability-discovery benchmark, CyberGym above 3.5 Flash Cyber and larger frontier models (no figure given), 2.6x more correct Chrome vulnerability patches than the best commercial models (Chrome Security team), and 7.5–9.7 points higher recall at 2.3–5.2x lower cost on Wiz's penetration-testing benchmarks — none of the site's tracked benchmarks, and none of 3.8 Flash's scores are stated to apply. Artificial Analysis has not indexed it and the gating makes that unlikely. Gemini 3.5 Flash Cyber (July 2026) is not tracked here and Google states no supersession, so predecessorId stays null.
For developers
API model strings
Not researched
Licence
Proprietary — weights not released
Retirement
No retirement announced
Lineage
No recorded predecessor