Google adds three Gemini Flash models while Pro delay continues
Google released lower-cost Gemini Flash models and a restricted cyber variant, but its next top-tier Pro model remains in partner testing.
By Renata Fuchs · Policy Reporter
· 4 min read
Google announced three new Gemini Flash models, Gemini 3.6 Flash, Gemini 3.5 Flash-Lite and Gemini 3.5 Flash Cyber, widening its lower-cost AI lineup while leaving Gemini 3.5 Pro out of public release. The move keeps Google focused on price, speed and specialized security use cases as OpenAI, Anthropic, Meta and Chinese labs press ahead with more capable public models.
Google did not disclose revenue, customer adoption or infrastructure capacity tied to the launch. It said Gemini 3.5 Pro is still being tested with partners and will be released when ready. The company also said pretraining for Gemini 4 has begun, calling it its most ambitious training run, but gave no launch date or benchmark data for that system.
Flash update targets cost and token efficiency
Gemini 3.6 Flash is positioned as a cheaper, more efficient model rather than Google’s highest-performing system. Google priced it at $1.50 per million input tokens and $7.50 per million output tokens. According to Artificial Analysis, the model is expected to use about 17% fewer output tokens than Gemini 3.5 Flash, while Google said savings can reach 65% on selected benchmarks such as DeepSWE.
Google reported benchmark gains over Gemini 3.5 Flash, including DeepSWE rising from 37% to 49%, MLE-Bench from 49.7% to 63.9%, and OSWorld-Verified from 78.4% to 83%. The GDPval-AA v2 benchmark increased from 1,349 to 1,421 points, according to the company. Gemini 3.6 Flash also has a one million token context window.
Google added Computer Use as a built-in client-side tool in the Gemini API and Gemini Enterprise, letting the model operate across browser, desktop and mobile workflows. The company also said it strengthened Frontier Safety protections related to cyberattacks and CBRN risks, meaning chemical, biological, radiological and nuclear threats.
Flash-Lite is built for throughput
Gemini 3.5 Flash-Lite is the smaller release, aimed at low-latency and high-volume workloads. Artificial Analysis lists it at 350 output tokens per second. Google priced the model at $0.30 per million input tokens and $2.50 per million output tokens.
Google said Flash-Lite beats the older Gemini 3 Flash on several agentic and coding benchmarks, including SWE-Bench Pro and OSWorld-Verified. Against Gemini 3.1 Flash-Lite, Google reported Terminal-Bench 2.1 improving from 31% to 54%. Those figures support the company’s cost-performance pitch, although they do not make Flash-Lite a frontier competitor.
Cyber model stays gated
Gemini 3.5 Flash Cyber is based on Gemini 3.5 Flash and tuned for cybersecurity work. Google has integrated it into CodeMender, a DeepMind code security agent that runs multiple Flash Cyber subagents in parallel and combines their findings into a report.
Google said the model scored 83.2% on the CyberGym benchmark, close to OpenAI’s GPT-5.5-Cyber at 85.6%. The company also said its Big Sleep team used Flash Cyber to search for critical flaws in Chrome and Safari, where it beat standard Flash models and Anthropic’s Claude Opus 4.6. In V8 JavaScript engine commit scanning, Google said Flash Cyber produced 55 confirmed unique findings, compared with 47 from standard Gemini 3.5 Flash and 36 from Opus 4.6.
Access is limited. Google said Gemini 3.5 Flash Cyber can be used only by governments and trusted partners through a CodeMender pilot because the model can support offensive as well as defensive security work. CodeMender itself is available more broadly in preview through the Gemini Enterprise Agent Platform, but that version runs on standard Gemini models and supports C/C++, Go, Java, Python, Ruby, Rust and TypeScript.
Pro delay leaves a gap at the top
The missing release is Gemini 3.5 Pro. Bloomberg recently reported that Google’s flagship model is months behind schedule while the company works to improve coding performance. Until that model is public, Google’s available Gemini lineup is competing more on efficiency and price than on the top end of model capability.
That position is awkward for Google. OpenAI is serving GPT-5.6 Sol, Anthropic has Fable and Mythos, and labs including Moonshot and Zhipu are pushing Kimi K3 and GLM-5.2 toward frontier performance. Meta has also released a model that beats Google’s current lineup at code writing, according to the reported comparisons. Google’s Flash releases are useful product work, but they do not fill the Pro gap.
This story draws on original reporting from The Decoder.