LLMCostLab

PricesCalculators → Long-form content generation

Long-form content generation: LLM cost calculator

Drafting long articles, reports, or descriptions. The only common workload where output tokens dominate — which flips the usual ranking, because output is priced 4-6x higher than input almost everywhere.

Default profile: 2,000 in / 4,000 out per call · 10,000 calls/mo · 30% cached
advertisement
Cheapest option
Gemini 2.5 Flash-Lite
$18.00/mo
Priciest option
GPT-5.6 Cyber
$3,182/mo
Annual gap
$37,974
same workload, different model

Every model on this workload

Input is 33% of tokens here, so output pricing dominates. Cache applied at 30%.

ModelIn /1MOut /1MNo cacheMonthlyvs best
Gemini 2.5 Flash-Lite Google$0.100$0.400$18.00$18.00cheapest
GPT-5.6 Luna OpenAI$0.100$0.600$26.00$25.461.4x
Gemini 3.1 Flash-Lite Google$0.250$1.500$65.00$65.003.6x
Gemini 3.5 Flash-Lite Google$0.300$2.500$106.00$106.005.9x
Gemini 2.5 Flash Google$0.300$2.500$106.00$106.005.9x
Gemini 3.7 Flash Google$0.750$3.750$165.00$165.009.2x
Gemini 3.6 Flash Google$0.750$3.750$165.00$165.009.2x
Claude Haiku 4.5 Anthropic$1.000$5.000$220.00$214.6011.9x
GPT-5.6 Terra OpenAI$1.000$6.000$260.00$254.6014.1x
Gemini 3.5 Flash Google$1.500$9.000$390.00$390.0021.7x
Gemini 2.5 Pro Google$1.250$10.00$425.00$425.0023.6x
Claude Sonnet 5 Anthropic$2.000$10.00$440.00$429.2023.8x
Gemini 3.1 Pro Google$2.000$12.00$520.00$520.0028.9x
GPT-5.3 Codex OpenAI$1.750$14.00$595.00$585.5532.5x
GPT-5.6 Sol OpenAI$2.500$15.00$650.00$636.5035.4x
Claude Sonnet 4.6 Anthropic$3.000$15.00$660.00$643.8035.8x
Claude Opus 5 Anthropic$5.000$25.00$1,100$1,07359.6x
Claude Opus 4.5 Anthropic$5.000$25.00$1,100$1,07359.6x
GPT chat-latest OpenAI$5.000$30.00$1,300$1,27370.7x
Claude Fable 5 Anthropic$10.00$50.00$2,200$2,146119.2x
Claude Mythos 5 Anthropic$10.00$50.00$2,200$2,146119.2x
GPT-5.6 Cyber OpenAI$12.50$75.00$3,250$3,182176.8x
advertisement

Adjust the assumptions

The defaults above are a reasonable starting profile, not your profile. Change them.

Cost calculator

Enter your own numbers, or start from a workload preset. Costs are monthly and update as you type.

Cheapest model
Cheapest → priciest spread
ModelIn /1MOut /1MMonthlyvs best

Cutting this bill in practice

Switching from GPT-5.6 Cyber to Gemini 2.5 Flash-Lite on this workload is worth $37,974 a year. Two things make a number like that collectible rather than hypothetical: being able to change model without an integration rewrite, and knowing when the volume justifies leaving per-token pricing entirely.

Switch models without rewriting your integration

The savings on this page are only real if you can actually move traffic to the cheaper model. Aggregators and gateways sit in front of multiple providers so that switch is a configuration change.

Run open-weight models on rented GPUs

Past a certain monthly volume, renting a GPU beats paying per token — but only for workloads where an open-weight model is good enough, and where utilization is high enough to amortize idle time.

Other workloads

Customer support chatbotRAG document Q&ACoding agentBulk summarizationStructured data extractionEmail triage agentTranslation pipeline