Gemini 3.5 Flash-Lite

Gemini 3.5 Flash-Lite

Google DeepMindMULTIMODAL Proprietary
Provider
Google DeepMind
Modality
MULTIMODAL
Parameters
Context window
Weights
Proprietary
Released
21 Jul 2026
Compare Gemini 3.5 Flash-Lite with others

Gemini 3.5 Flash-Lite (July 2026) is Google's fastest, cheapest 3.5-class model — running at ~350 output tokens/sec with a strong price-to-performance ratio, and now with built-in computer use. It even beats the older 3 Flash on several coding and agentic benchmarks.

Best for

  • High-throughput, low-latency agentic search and document processing
  • Cheap, fast bulk tasks (classification, extraction, routing)
  • Agentic subagent workloads (configurable thinking levels)
  • When speed and cost matter more than peak reasoning

How it compares — and the India angle

  • At $0.30/1M input and $2.50/1M output it's one of the best value models around — perfect for Indian high-volume apps and agents.
  • 350 tokens/sec is blazing fast; use minimal thinking for cheap high-volume work, higher thinking for multi-step subagents.
  • Rolling out in the Gemini API, AI Studio, the Gemini app and Google Search.

How to access

API $0.30 / 1M input, $2.50 / 1M output. ~350 output tokens/sec. Via Gemini API, AI Studio, Gemini app and Google Search.

Access Gemini 3.5 Flash-Lite
Share: