LLMCostLab

PricesCompare → Gemini 3.7 Flash vs Gemini 2.5 Flash-Lite

Gemini 3.7 Flash vs Gemini 2.5 Flash-Lite: API cost compared

Gemini 3.7 Flash lists at $0.750 in / $3.750 out per 1M tokens. Gemini 2.5 Flash-Lite lists at $0.100 in / $0.400. Sticker price only tells you so much — below is what each actually costs across eight production workloads.

advertisement
Cheaper on more workloads
Gemini 2.5 Flash-Lite
wins 8 of 8
Blended rate gap
8.6x
at 3:1 in:out

Side by side

Gemini 3.7 FlashGemini 2.5 Flash-Lite
ProviderGoogleGoogle
Input per 1M$0.750$0.100
Output per 1M$3.750$0.400
Cached input per 1Mnot publishednot published
Context window1M1M
Blended per 1M (3:1)$1.500$0.175

Monthly cost by workload

Caching applied at each workload's typical hit rate. This is where the two models actually separate.

WorkloadGemini 3.7 FlashGemini 2.5 Flash-LiteCheaperDifference
Customer support chatbot$133.12$16.00Gemini 2.5 Flash-Lite$117.12 (88%)
RAG document Q&A$165.00$20.80Gemini 2.5 Flash-Lite$144.20 (87%)
Coding agent$120.00$14.80Gemini 2.5 Flash-Lite$105.20 (88%)
Bulk summarization$1,050$136.00Gemini 2.5 Flash-Lite$914.00 (87%)
Structured data extraction$1,781$225.00Gemini 2.5 Flash-Lite$1,556 (87%)
Long-form content generation$165.00$18.00Gemini 2.5 Flash-Lite$147.00 (89%)
Email triage agent$487.50$62.00Gemini 2.5 Flash-Lite$425.50 (87%)
Translation pipeline$1,069$118.50Gemini 2.5 Flash-Lite$950.25 (89%)
advertisement

How to read this

Gemini 2.5 Flash-Lite is cheaper on the majority of these workloads, but "majority" is not the decision. The output:input price ratio is 5.0 for Gemini 3.7 Flash and 4.0 for Gemini 2.5 Flash-Lite. If your traffic is output-heavy — long-form drafting, translation — the model with the lower output price wins regardless of what the input price suggests. If you are input-heavy — RAG, bulk summarization, agents re-reading a codebase — the cached input rate matters more than either headline number.

Cost is also not the only axis. This site does not benchmark quality, latency, or rate limits, and a model that is 3x cheaper but needs two attempts per task is not cheaper. Use these numbers to size a bill, then validate on your own evals.

Run your own numbers

Cost calculator

Enter your own numbers, or start from a workload preset. Costs are monthly and update as you type.

Cheapest model
Cheapest → priciest spread
ModelIn /1MOut /1MMonthlyvs best