News in Short
- Google has launched Gemini 3.6 Flash and Gemini 3.5 Flash-Lite.
- Gemini 3.6 Flash delivers better coding, multimodal performance and lower token usage.
- Gemini 3.5 Flash-Lite is Google’s fastest and cheapest Gemini 3.5 model.
- Google also confirmed that Gemini 4 has entered the pre-training phase.
- Both new AI models are available through Google AI Studio, Gemini API and Gemini Enterprise.
Google has expanded its Gemini AI portfolio with the launch of Gemini 3.6 Flash and Gemini 3.5 Flash-Lite. Alongside the new models, the company also revealed that development of Gemini 4 is already underway, with its most ambitious pre-training run now in progress.
The latest announcements strengthen Google’s AI offerings for developers and enterprises, focusing on faster performance, lower costs and more efficient AI agents.
Gemini 3.6 Flash focuses on smarter AI workloads
Gemini 3.6 Flash succeeds Gemini 3.5 Flash with improvements across coding, multimodal reasoning and complex knowledge tasks. Google claims the model uses up to 17 percent fewer output tokens on the Artificial Analysis Index, making it more efficient while maintaining strong performance.
The company also says the model requires fewer reasoning steps and tool calls when completing multi-step workflows, helping reduce latency and operational costs.
On industry benchmarks, Gemini 3.6 Flash outperformed its predecessor across multiple tests. It scored 49 percent on DeepSWE compared to 37 percent for Gemini 3.5 Flash. The model also achieved 63.9 percent on MLE Bench and 83 percent on OSWorld-Verified, highlighting improvements in coding, reasoning and software automation.
Google has priced Gemini 3.6 Flash at $1.50 per million input tokens and $7.50 per million output tokens.
The company has also strengthened safety protections. Gemini 3.6 Flash includes enhanced safeguards against misuse related to cyber threats as well as chemical, biological, radiological and nuclear (CBRN) risks.
Gemini 3.5 Flash-Lite prioritises speed and affordability
Alongside the flagship model, Google introduced Gemini 3.5 Flash-Lite, which it describes as the fastest and most affordable model in the Gemini 3.5 family.
Designed for high-volume applications such as document processing, agentic search and AI automation, the model can generate up to 350 output tokens per second.
Developers can also choose different reasoning levels depending on the task. Lower thinking modes reduce latency and costs, while higher reasoning settings are designed for more complex AI workflows.
Gemini 3.5 Flash-Lite is priced at just $0.30 per million input tokens and $2.50 per million output tokens, making it one of Google’s most cost-effective AI models.
Google also introduces a cybersecurity-focused Gemini model
Google also announced Gemini 3.5 Flash Cyber, a specialised version designed to detect, validate and patch software vulnerabilities.
The model powers Google’s CodeMender security agent and will not be publicly available. Instead, it will be offered through a limited pilot programme for governments and trusted partners.
Gemini 4 is already in development
Perhaps the biggest announcement was Google’s confirmation that Gemini 4 has entered its pre-training phase.
The company described it as its “most ambitious” pre-training run so far. However, Google did not reveal a launch timeline or expected release window.
Meanwhile, Gemini 3.5 Pro remains in partner testing and will receive a broader rollout once it is ready.
Availability
Gemini 3.6 Flash and Gemini 3.5 Flash-Lite are available starting today through Google AI Studio, Gemini API, Android Studio and the Gemini Enterprise platform.
The models are also accessible through the Gemini app, while Gemini 3.5 Flash-Lite is additionally rolling out to Google Search, bringing faster AI capabilities to more users.
With Gemini 4 already in development, Google’s latest announcements signal that the competition in generative AI is only accelerating.