
Google's Gemini API pricing page lists 32 models. Thirty-one of them show one price. Three show two.
Gemini 3.6 Flash, Gemini 3.7 Flash and Gemini 3.8 Flash each carry a rate that holds "through December 31, 2026" and a second rate "starting January 1, 2027". Every one of those dated lines, 36 across the three models, is exactly double.
If you are costing a production workload on today's Flash rate, your spreadsheet is correct for four months.
What actually changes#
The standard paid tier for all three Flash models is identical, and so is the step:
No other Gemini model on the page has a second column. Gemini 3.5 Flash, 3.5 Flash-Lite, the 3.1 line, the 2.5 line, the embedding models and the robotics previews all list a single rate with no end date attached.
This is a discount ending, not a penalty#
That distinction matters, and skipping it would make the number above sound worse than it is.
Gemini 3.5 Flash, the model the current line replaced, costs $1.50 per million input tokens and $9.00 per million output tokens today, flat, with no expiry. So on 1 January the newer Flash models land at the same input price as the older one and a lower output price.
Read the two facts together and the shape is clear. Google put the whole current Flash line on half price, published the date the half price ends, and left the previous generation's rate as the floor it returns to. Nobody is being charged more than the model this one succeeded. Everyone modelling on today's number is modelling on a promotion.
That is a familiar move once you know to look for it. Anthropic did something structurally similar in the other direction, cutting the cost of a model without touching its per-token price, and the only way to see either is to read the rate card rather than the announcement.
The second multiplier is in the model card#
The Gemini 3.8 Flash card, published on 2 September 2026, describes the model as "designed for cost-effective scaling of general-purpose, production-ready agents". Further down, in Known Limitations, it adds a sentence that is not on the Gemini 3.7 Flash card:
"At times, the model might use more tokens to maximize performance, especially at higher effort levels."
That is a cost disclosure sitting in a safety section. It matters because of the middle row of the table above: output pricing on these models is billed including thinking tokens. The effort level you choose does not only change quality and latency. It changes the number the output rate multiplies.
So two things that determine your bill both have room to move upward on a known schedule. The price per token doubles on 1 January. The tokens per request are documented by the vendor as variable, in the direction of more.
What to do about it#
Three things, in order of how much they are worth.
Budget at the January number, not today's. If the workload still clears its margin at $1.50 and $7.50 per million, the discount is a bonus for four months. If it only clears at $0.75 and $3.75 per million, you do not have a viable unit cost, you have a promotion.
Pin the effort level and measure your own token spend before you commit. The card says effort level controls the mix of quality, cost and latency, and warns that higher settings may spend more tokens. A per-million headline rate tells you nothing about how many millions your actual prompts consume. Run your real traffic at a fixed effort setting for a day and read the invoice, rather than trusting the arithmetic.
If cost is the binding constraint, look at the models with no expiry date. Gemini 3.5 Flash-Lite is $0.30 per million input and $2.50 per million output, single rate, no date attached. It is a smaller model and it will not do everything the 3.8 line does, but its price is the price. This is the same calculation as picking a consumer plan, where the question is which capability you actually use rather than which tier sounds most complete.
The other thing the card says out loud#
One number on the 3.8 Flash card moved much further than the rest, and it is worth knowing about if you serve users outside English.
Google's own safety comparison against Gemini 3.7 Flash reports Multilingual Safety at +5.4pp, on a scale where lower is better. Text-to-text safety moved 0.4pp the good way, image-to-text did not move, tone moved 0.2pp, and unjustified refusals moved 1.1pp. Everything except the multilingual line sits inside roughly one point. The card states the conclusion plainly: "Safety performance across non-English languages regressed slightly relative to 3.7 Flash."
The word Google uses is "slightly". For context on what slightly means here, the equivalent line on the 3.7 Flash card, comparing it to 3.6 Flash, was -0.48pp, which was an improvement. The same measurement moved about eleven times as far this generation, in the opposite direction.
Google does not publish the base rate that percentage applies to, so the honest reading is a relative one: this is the largest single movement on the card, it is a regression, and it is confined to non-English. If your product answers in Spanish, Arabic, Hindi or anything else, that is an argument for running your own safety evaluation in those languages before upgrading, rather than treating a point release as a drop-in.
Read the footnotes on the benchmarks too#
The evaluation methodology for Gemini 3.8 Flash is a four-page PDF, and Google publishes the comparison caveats in it rather than hiding them. They are still worth knowing before you take a bar chart at face value.
"All the results for non-Gemini models are sourced from providers' self reported numbers unless otherwise mentioned below." On DeepSWE v1.1 and Terminal-Bench 2.1, Gemini's numbers are self computed while rivals' come from public leaderboards. On OSWorld 2.0, Gemini scores are "maxed over 3 runs with a single attempt per run" while the GPT-5.6 Terra figure is taken from a blog post and the Opus 5 figure from another. On LVBench, Gemini and GPT-5.6 models were given 1,024 video frames and Claude models 300, which Google attributes to API limitations. On the science benchmark HLE-Verified, Google notes that "a significant proportion of questions were blocked by content policy filters for Sonnet 5", and reports the resulting score anyway.
None of that is misconduct, and Google disclosing it is better than the alternative. It does mean the chart is not measuring all five models under the same conditions, and the sentence that tells you so is on page three of a PDF linked from a line of small text.
The general lesson holds beyond this release. The number a vendor puts on a slide is the number they chose to be able to put on a slide, which is one of the costs of building on someone else's model that rarely appears in the business case. The rate card, the model card and the methodology note are all published, all free, and all say more than the announcement does.
Sources
- Gemini Developer API pricing, Google AI for Developersai.google.dev
- Gemini 3.8 Flash Model Card, Google DeepMinddeepmind.google
- Gemini 3.7 Flash Model Card, Google DeepMinddeepmind.google
- Gemini 3.8 Flash Model evaluation: approach, methodology and results, Google DeepMindstorage.googleapis.com


Discussion