Gemini 3.6 Flash cuts token use as Google starts Gemini 4
Google’s new Gemini Flash lineup targets the economics of production agents: Gemini 3.6 Flash uses fewer output tokens, Gemini 3.5 Flash-Lite is built for high-throughput work, and Google says its most ambitious Gemini 4 pre-training run has begun.
Google’s July 21 Gemini release is a three-tier push toward cheaper, faster production agents. Gemini 3.6 Flash is the general workhorse, Gemini 3.5 Flash-Lite is the high-throughput option, and Gemini 3.5 Flash Cyber is a restricted cybersecurity specialist. Google’s announcement also says the Gemini 4 pre-training run has begun, although Google has not announced a release date.
The practical change is not simply a new model number. Google’s launch coverage reports that Gemini 3.6 Flash uses 17% fewer output tokens than Gemini 3.5 Flash and costs $7.50 rather than $9.00 per 1 million output tokens, while Flash-Lite targets repetitive, latency-sensitive work. For teams building an AI agent, that makes routing and completed-task cost more important than choosing one model for every step.
Definition: Google’s new Flash lineup separates general capability, high-volume throughput and controlled cybersecurity work.
Example: Google positions Flash-Lite for agentic search and document processing, while Gemini 3.6 Flash is positioned for coding and knowledge work.
Key takeaway: Compare models by the cost and time of an accepted result, not by token price alone.
Business impact: Fewer generated tokens and fewer reasoning steps can lower the operating cost of multi-step agents, if the cheaper route still meets the workflow’s quality bar.
What changes in Gemini 3.6 Flash?
Gemini 3.6 Flash is Google’s general-purpose workhorse for coding and knowledge work. Google says Gemini 3.6 Flash consumes 17% fewer output tokens than Gemini 3.5 Flash on the Artificial Analysis Index and takes fewer reasoning steps and tool calls on multi-step workflows. For an agent that repeatedly plans, calls tools and checks results, the useful test is a complete task trace rather than a standalone response.
Gemini 3.6 Flash also lowers the listed output price to $7.50 per 1 million tokens, compared with the $9.00 rate cited for Gemini 3.5 Flash; input pricing is $1.50 per 1 million tokens. The combination of lower output pricing and lower output volume is the release’s clearest operating argument, but a team should still measure retries, context growth and human review before claiming a real saving.
Google reports higher scores for Gemini 3.6 Flash than Gemini 3.5 Flash on several evaluations: DeepSWE at 49% versus 37%, MLE Bench at 63.9% versus 49.7%, OSWorld-Verified at 83.0% versus 78.4%, and GDPval-AA v2 at 1421 versus 1349. These are Google’s reported comparisons, not guarantees for every application, so operators should reproduce them on representative workloads before changing a production model.
Why is Gemini 3.5 Flash-Lite the throughput tier?
Gemini 3.5 Flash-Lite is aimed at low-latency, high-throughput workloads such as agentic search and document processing. The model is priced at $0.30 per 1 million input tokens and $2.50 per 1 million output tokens, according to the launch coverage. That combination makes Flash-Lite the natural candidate for repetitive work where volume and response time matter more than maximum capability.
Google reports that Gemini 3.5 Flash-Lite improves on Gemini 3.1 Flash-Lite in Terminal-Bench 2.1, GDM-MRCR v2 and GDPval-AA v2, and it says Flash-Lite also beats Gemini 3 Flash on SWE-Bench Pro and OSWorld-Verified. Those figures support testing the smaller tier for extraction, routing and subagent work, but they do not remove the need for structured outputs, validation and escalation when an error is expensive.
| Model | Best-supported role | Input price / 1M tokens | Output price / 1M tokens | Access described in the announcement |
|---|---|---|---|---|
| Gemini 3.6 Flash | Coding and knowledge work | $1.50 | $7.50 | Gemini app; developer access via Google Antigravity, AI Studio and Android Studio |
| Gemini 3.5 Flash-Lite | Low-latency, high-throughput agentic work | $0.30 | $2.50 | Gemini app; Search rollout; developer access via AI Studio and Android Studio |
| Gemini 3.5 Flash Cyber | Finding, validating and patching vulnerabilities | — | — | Limited-access CodeMender pilot for governments and trusted partners |
What does Gemini 3.5 Flash Cyber add?
Gemini 3.5 Flash Cyber is a specialized model for detecting, validating and patching code-security issues at scale. Google says it is built on Gemini 3.5 Flash and paired with the CodeMender security agent, using Flash’s efficiency to target a lower price per token than larger cybersecurity models. The relevant takeaway for security teams is controlled specialization, not general availability.
Gemini 3.5 Flash Cyber is initially limited to governments and trusted partners through a CodeMender pilot. Google says that restricted deployment reflects the technology’s dual-use nature: the model can help defenders find and fix critical vulnerabilities, but similar capability could also be misused. Teams should therefore treat the model as a governed security capability with access controls and review, not as an ordinary low-cost API tier.
Where does Gemini 4 fit?
Gemini 4 is not a launched product in this announcement. Google says the DeepMind team has started its most ambitious pre-training run yet for Gemini 4, while Gemini 3.5 Pro is still testing with partners. Reuters’ coverage frames the update as a lightweight-model release that leaves the flagship Pro timing unresolved.
That sequencing matters for operators. Gemini 3.6 Flash and Flash-Lite are available now, while Gemini 4 remains a forward-looking signal rather than a model teams can evaluate. The launch coverage reports availability in the Gemini app and developer access through Google Antigravity, AI Studio and Android Studio, with Flash-Lite also coming to Search; teams should verify the channels and model IDs relevant to their own accounts before migrating. More on this: Gemini app plans after AI Studio cancellation.
How should Gemini Flash models be measured?
Gemini 3.6 Flash and Gemini 3.5 Flash-Lite should be judged by the economics of a complete accepted task. Google’s launch coverage reports 17% fewer output tokens for Gemini 3.6 Flash and lower listed output prices for both Flash models; operators should therefore record cost per completed task, time to an accepted result, retry count, tool-call count and escalation rate before and after the change. For a workflow already using an AI automation stack, those measurements are more useful than token price alone.
Gemini 3.6 Flash, Gemini 3.5 Flash-Lite and Gemini 3.5 Flash Cyber make model routing a practical design choice because Google positions them for distinct workloads: general coding and knowledge work, high-throughput agentic search and document processing, and restricted vulnerability research. Operators should route by failure cost, validate the cheaper tier and keep Flash Cyber behind the limited-access controls Google describes; the launch therefore points toward a portfolio of specialized workers rather than one model applied everywhere.
Google’s Gemini 4 pre-training announcement raises the ceiling, but it does not change today’s operating decision. Builders can evaluate 3.6 Flash and Flash-Lite now, keep Gemini 3.5 Pro on the watchlist, and treat Gemini 4 as a future release until Google provides a model, access path and benchmarks. Background: Gemini for macOS adds voice control and smart dictation. Background: NotebookLM becomes Gemini Notebook, with code execution.
Frequently asked questions
What did Google launch in the Gemini Flash lineup?
Google announced Gemini 3.6 Flash, Gemini 3.5 Flash-Lite and Gemini 3.5 Flash Cyber on July 21, 2026. Gemini 3.6 Flash is the general workhorse for coding and knowledge work; Flash-Lite is the fast, lower-cost option for high-throughput work; and Flash Cyber is a specialized cybersecurity model paired with CodeMender in a limited-access pilot.
How much does Gemini 3.6 Flash cost?
Google lists Gemini 3.6 Flash at $1.50 per 1 million input tokens and $7.50 per 1 million output tokens. The output price is lower than the $9.00 rate cited for Gemini 3.5 Flash, and the launch coverage reports that Gemini 3.6 Flash uses 17% fewer output tokens than its predecessor.
What is Gemini 3.5 Flash-Lite designed for?
Gemini 3.5 Flash-Lite is designed for low-latency and high-throughput tasks such as agentic search and document processing. The model costs $0.30 per 1 million input tokens and $2.50 per 1 million output tokens. For operators, that makes Flash-Lite a candidate for repetitive extraction and routing, provided structured outputs and validation keep error costs under control.
Is Gemini 4 available now?
No. Google says its DeepMind team has started its most ambitious pre-training run yet for Gemini 4, but it did not announce a Gemini 4 release date. Google also says Gemini 3.5 Pro is testing with partners and will become broadly available when it is ready. For teams, Gemini 4 is therefore a future watch item rather than a migration target today.
Alex
Founder & Lead AI Writer
Alex is the founder of Yowox and lead AI writer since 2024, breaking down complex information into clear, actionable insights for thousands of readers every day. Alex has built AI automation systems for businesses since 2024, focusing on AI agents, workflow automation, and business process optimization.
Save hours. Save thousands.
Practical guides, real workflows, and the latest AI and automation news that matters — straight to your inbox.