105 Days Until Gemini's Price Doubles: 3 Numbers to Check
Gemini 3.8 Flash costs $0.75 per million input tokens today and $1.50 from 1 January 2027. Google says so on its own pricing page. 3 lines of your bill double that day, and one of them is not the one people watch.
If you have been routing bulk traffic to Gemini Flash because it is a fraction of what the flagships cost, this is the number to put in your calendar. On 1 January 2027 every rate on that model doubles, and Google published the date itself when it launched Gemini 3.8 Flash on 2 September 2026.
The gap is what makes it matter. Claude Fable 5.1 and GPT-6 Astra both charge $10 per million input tokens. Gemini 3.8 Flash charges $0.75 — about a thirteenth. That is why so much production traffic sits on Flash, and why a doubling is worth planning for rather than discovering.
1. The introductory price expires on 31 December
This is not a discount somebody is offering on top of Google's price. It is Google's own price schedule, with the end date written into the launch announcement and the public pricing page.
| Rate | Through 31 Dec 2026 | From 1 Jan 2027 |
|---|---|---|
| Input / 1M tokens | $0.75 | $1.50 |
| Output / 1M tokens (thinking included) | $3.75 | $7.50 |
| Cached input / 1M | $0.075 | $0.15 |
| Cache storage / 1M / hour | $0.50 | $1.00 |
| Batch input / output | $0.375 / $1.875 | $0.75 / $3.75 |
Gemini 3.6 Flash and 3.7 Flash double on the same date. Moving down a generation to escape the rise does not work — worth knowing before planning a migration that saves nothing.
2. The line most people forget is the cache
Everyone checks input and output. Far fewer check cached reads, and on agentic work — an assistant re-reading the same system prompt, the same documents, the same growing conversation on every turn — cached reads are most of the bill.
For comparison, Claude Fable 5.1 reads cached context at $0.25 per million and Claude Sonnet 5 at $0.20. After January, Gemini Flash's cached read at $0.15 is still the cheapest of those — but the distance closes, and a routing decision made on today's numbers may not survive the change.
3. The same per-token price can still be a bigger bill
Gemini 3.8 Flash launched at exactly the rate 3.7 Flash and 3.6 Flash already charged — three generations on one price card. That reads like a free upgrade. Per task, it is not always.
This is the trap in every per-token comparison, and it cuts both ways: a model that finishes in one attempt can be cheaper than a cheaper model that needs three. Measure cost per completed task, not cost per token. It is the only number that pays your bill.
What to actually do before January
- Price your real workload at the January rate, not today's. Take last month's token usage, double the Gemini lines, and see whether the routing decision still holds.
- Check what share of your input is cached. If it is high, the cache-read and storage rows matter more to you than the headline input price.
- Measure per completed task on your own work, at the effort level you actually use. Benchmarks are run at settings you may not be running.
- Set a reminder for mid-December. Prices change without notice, and the sensible thing is to re-read Google's pricing page then rather than trust this article's numbers four months on.
Two other things that happened this fortnight
Gemini 3.8 Live and 3.8 Live Extended Thinking arrived on 15 September 2026 — live dialogue models built for voice agents, able to handle interruptions and switch languages mid-conversation, with the Extended Thinking version aimed at multi-step tasks. They are in the Gemini API and AI Studio for developers, and in Search Live and Gemini Live for everyone else. For an Indian audience the language switching is the interesting part; we will test it properly rather than repeat the launch copy.
GPT-6 Astra became the first model OpenAI has classified at the Critical level for cybersecurity under its Preparedness Framework, reported on 17 September. Microsoft made it generally available in Foundry Models the same day. That classification is why the rollout started with vetted organisations rather than everybody.
Save this summary as an image or share it.
AICreatorHub Team
The AICreatorHub editorial team is a group of hands-on AI practitioners, writers and developers based in India. We test AI tools and models ourselves, track official releases from OpenAI, Anthropic, Google, Meta and xAI, and translate them into simple, India-first guides in English and Hindi. Every article is written for real Indian use cases — pricing in rupees, free-tier tips and practical, tested steps — so you get accurate, up-to-date and genuinely useful AI information.