Google dropped three new AI models today. They are all built on the Gemini 3.5 Flash foundation. The goal here is efficiency. Speed. Reliability. The company is pushing for lower token usage across the board. But if you were waiting for the flagship heavy hitter, Gemini 3.5 Pro, that ship hasn’t sailed yet. It’s still in the oven.
Here is what actually matters from the Tuesday announcement.
Why Gemini 3.6 Flash Is the New Workhorse
Google is calling this the “workhorse.” That’s a loaded term. Usually, that means it’s boring but essential. In this case, it’s supposed to be smarter about coding and knowledge work. The multimodal performance got a boost too.
But the real selling point is the cost.
The model reduces token usage by up to 17% compared to its predecessor. Fewer tokens. Lower cost. That is a direct hit to developer budgets. Google says they built these changes based on feedback from people actually using the tools. Benchmarks back up the claim, showing gains in performance alongside the token savings.
If you are looking for how Gemini 3.6 Flash improves efficiency, the answer lies in that 17% reduction. It gets the job done without burning through credits.
Gemini 3.5 Flash-Lite: Speed for Agentic Workflows
There is a trade-off for that speed. Gemini 3.5 Flash-Lite outputs 350 tokens per second. That is fast. It is also Google’s most cost-effective option.
Why does speed matter right now? Because of agentic workflows. These models are starting to take on tasks that require rapid, sequential decision-making. The Lite version outperforms previous generations in this specific area.
It also has built-in “computer use.” It can navigate interfaces like a human would to get things done. This isn’t just about chat. It’s about action.
Where Gemini 3.5 Flash Cyber Fits In
This one is niche but dangerous in the best way. Gemini 3.5 Flash Cyber focuses on security. It finds vulnerabilities. It fixes them.
It works with an infrastructure agent named CodeMender. Together, they help teams patch issues quickly.
The results are already inside Google. The model is hunting down bugs in Android, Chrome, and YouTube. But you can’t just sign up and get access yet.
For now, it is limited to governments and trusted partners. The scope will expand later. This highlights a shift in how Google rolls out sensitive models. They start with the internal codebases. They test thoroughly. They limit external access until trust is established.
The Delay of Gemini 3.5 Pro
The flagship model. The one everyone wanted. Gemini 3.5 Pro is not ready.
Google says it is currently in testing with partners. They will make it available “as soon as it’s ready.” That phrase is dangerous. It means the date is undefined. It could be weeks. It could be months.
While we wait, Google is already looking past the current generation. Pretraining for Gemini 4 has begun. That model’s release date is also undetermined.
This creates a strange gap. The efficient models are here. The niche security tools are rolling out. But the top-tier reasoning engine you might expect to anchor the ecosystem is still absent.
How to Access These New Models
You don’t need an invitation for the general Flash models.
- Gemini 3.6 Flash is live in the Gemini API, Google AI Studio, and Android Studio. It is also available in Google Antigravity.
- Gemini 3.5 Flash-LITE is rolling out to the Gemini app and Google Search.
The distinction between the two is clear. Use 3.6 Flash for complex tasks where cost and coding matter. Use 3.5 Flash-Lite for high-volume, fast-paced agentic work where every millisecond and penny counts.
The Pro version remains on the horizon. We are left with powerful intermediates and a promise of what’s coming next. The AI landscape moves fast. Today, we get more tools. Tomorrow, perhaps the flagship arrives. Until then, the Flash series is doing the heavy lifting.






























