Gemini 3.5 Flash-Lite
Google DeepMindMULTIMODAL Proprietary
- Provider
- Google DeepMind
- Modality
- MULTIMODAL
- Parameters
- —
- Context window
- —
- Weights
- Proprietary
- Released
- 21 Jul 2026
Gemini 3.5 Flash-Lite (July 2026) is Google's fastest, cheapest 3.5-class model — running at ~350 output tokens/sec with a strong price-to-performance ratio, and now with built-in computer use. It even beats the older 3 Flash on several coding and agentic benchmarks.
Best for
- High-throughput, low-latency agentic search and document processing
- Cheap, fast bulk tasks (classification, extraction, routing)
- Agentic subagent workloads (configurable thinking levels)
- When speed and cost matter more than peak reasoning
How it compares — and the India angle
- At $0.30/1M input and $2.50/1M output it's one of the best value models around — perfect for Indian high-volume apps and agents.
- 350 tokens/sec is blazing fast; use minimal thinking for cheap high-volume work, higher thinking for multi-step subagents.
- Rolling out in the Gemini API, AI Studio, the Gemini app and Google Search.
How to access
API $0.30 / 1M input, $2.50 / 1M output. ~350 output tokens/sec. Via Gemini API, AI Studio, Gemini app and Google Search.
Access Gemini 3.5 Flash-Lite