Gemini 3.5 Flash-Lite (July 2026) is Google's fastest, cheapest 3.5-class model — running at ~350 output tokens/sec with a strong price-to-performance ratio, and now with built-in computer use. It even beats the older 3 Flash on several coding and agentic benchmarks.
Best for
- High-throughput, low-latency agentic search and document processing
- Cheap, fast bulk tasks (classification, extraction, routing)
- Agentic subagent workloads (configurable thinking levels)
- When speed and cost matter more than peak reasoning
How it compares — and the India angle
- At $0.30/1M input and $2.50/1M output it's one of the best value models around — perfect for Indian high-volume apps and agents.
- 350 tokens/sec is blazing fast; use minimal thinking for cheap high-volume work, higher thinking for multi-step subagents.
- Rolling out in the Gemini API, AI Studio, the Gemini app and Google Search.