How to Reduce LLM API Costs
There is no single trick — there is a bill shape, and seven levers that act on different parts of it. An LLM invoice is tokens in × rate, plus tokens out × a much higher rate, and almost every real overspend comes from paying one of those twice. Below: the seven levers ranked by typical return, each priced against the eight models tracked here, and a calculator you can point at your own volume.
The seven levers, ranked
Ordered by what they typically return on an input-heavy production bill. Each row links to the page that measures it.
Price your own bill shape
The defaults below are the reference workload used throughout this page: 4,000 input tokens, 800 output tokens, 10,000 requests a day. Change them to your own numbers.
| Model | Input | Output | Per request | 30 days | 1 year |
|---|
Informational estimate only. Provider pricing can include caching, batch processing, tool calls, regional pricing, and other rules not represented here.
One workload, eight models
Reference workload: 4,000 input and 800 output tokens per request, 10,000 requests a day — 300,000 requests and 1.44 billion tokens a month.
| Model | Input / 1M | Output / 1M | Output ÷ input | Keep output under | 30-day cost | Output share of bill |
|---|---|---|---|---|---|---|
GPT-5.6 Sol OpenAI · promotional to Nov 21, 2026 | $4.00 | $20.00 | 5.0x | 20.0% of input tokens | $9,600.00 | 50.0% |
GPT-5.6 Terra OpenAI | $2.00 | $12.00 | 6.0x | 16.7% of input tokens | $5,280.00 | 54.5% |
GPT-5.6 Luna OpenAI | $0.20 | $1.20 | 6.0x | 16.7% of input tokens | $528.00 | 54.5% |
Claude Opus 5 Anthropic · flagship rates | $5.00 | $25.00 | 5.0x | 20.0% of input tokens | $12,000.00 | 50.0% |
Claude Sonnet 5 Anthropic | $2.00 | $10.00 | 5.0x | 20.0% of input tokens | $4,800.00 | 50.0% |
Claude Haiku 4.5 Anthropic · 200K window | $1.00 | $5.00 | 5.0x | 20.0% of input tokens | $2,400.00 | 50.0% |
Gemini 3.6 Flash Google · promotional to Dec 31, 2026 | $0.75 | $3.75 | 5.0x | 20.0% of input tokens | $1,800.00 | 50.0% |
Gemini 3.5 Flash-Lite Google | $0.30 | $2.50 | 8.3x | 12.0% of input tokens | $960.00 | 62.5% |
Seven moves, and what each returns
Each lever acts on a different line of the bill. The numbers are computed from the same eight published rate cards — no estimates of provider behaviour beyond the arithmetic shown.
1. Stop re-sending chat history
Chat APIs are stateless, so turn N re-transmits turns 1 to N−1. A 20-turn support conversation with 200-token messages and 400-token replies sends 118,000 input tokens to produce 8,000 output tokens — 96.6% of the input is re-sent history. Keeping five prior turns cuts input to 55,000; an 800-token rolling summary cuts it to 20,000.
Measure it: Conversation Cost Growth Calculator
2. Trim the prompt itself
Prompts accumulate duplicated instructions, repeated few-shot blocks, markup and politeness filler. Across three sample prompts measured on this site, the avoidable share ran from 27.9% to 53.5% — and the largest single category was different every time, so there is no universal fix to apply blind.
Measure it: Prompt Weight Analyzer
3. Convert HTML before sending it
Retrieved pages bill every div, href and aria-label at content rates. On four representative page types, Markdown conversion removed 47.4%–67.9% of input tokens and plain text removed 56.0%–74.3%. Boilerplate repeats on every fetch, so crawlers pay for it thousands of times.
Measure it: HTML → Markdown Token Savings
4. Buy the cheapest model that passes your bar
Model choice dwarfs prompt hygiene. Moving the reference workload from Claude Opus 5 to Sonnet 5 saves $7,200.00 a month; moving it to Gemini 3.5 Flash-Lite saves $11,040.00. Route easy traffic to Luna-class models and reserve flagship rates for requests that actually need them — that routing decision is a bigger lever than any prompt edit.
Compare: LLM API Pricing Comparison · OpenAI vs Claude
5. Cap the output
Every model here charges far more to generate than to read. A thousand avoidable output tokens per request costs $6,000.00 a month on GPT-5.6 Sol, $7,500.00 on Claude Opus 5 and $1,125.00 on Gemini 3.6 Flash at 10,000 requests a day. Set max_tokens, use stop sequences, and ask for structured brevity rather than prose you will discard.
Why the premium exists: What Is an AI Token?
6. Chunk past OpenAI's 272K threshold
GPT-5.6 models apply higher long-context rates above 272,000 input tokens, where input doubles and output rises 1.5x. A 1M-token prompt therefore bills $8.00 per million input on Sol instead of $4.00 — more than Claude Opus 5's $5.00. Splitting the same content into four 250,000-token requests restores the base rate.
Full arithmetic: How Many Words Is 1M Tokens?
7. Budget to the revert date, not the promo
Two of the eight rates are temporary. GPT-5.6 Sol at $4.00/$20.00 holds through at least Nov 21, 2026 and reverts to $5.00/$30.00 — the reference workload goes from $9,600.00 to $13,200.00 a month. Gemini 3.6 Flash at $0.75/$3.75 holds through Dec 31, 2026 and reverts to $1.50/$7.50, doubling it from $1,800.00 to $3,600.00.
Track it: Pricing changelog · OpenAI API Pricing
Do them in this order
Cheapest reliable wins first, changes that need evaluation last.
max_tokens, add stop sequences, and cut the verbose instructions that invite long answers. Watch Flash-Lite first: 12.0% of input tokens is a tight output budget.LLM cost reduction questions
What is the fastest way to reduce LLM API costs?
For multi-turn applications, stop re-sending the whole chat history: a 20-turn conversation with 200-token messages and 400-token replies sends 118,000 input tokens to produce 8,000 output tokens, and 96.6% of that input is re-sent history. Keeping five prior turns cuts input tokens to 55,000 (53.4% less); a rolling summary cuts them to 20,000 (83.1% less). For single-request applications the fastest win is usually model choice — the same workload costs $528 to $12,000 a month across the eight models tracked here.
Why is my output cost so high when I generate few tokens?
Because output is priced 5x to 8.3x input on every tracked model. Output overtakes input as the larger line once reply tokens exceed 20% of prompt tokens on six models, 16.7% on GPT-5.6 Terra and Luna, and 12% on Gemini 3.5 Flash-Lite. In the reference workload above, output is 16.7% of the tokens but 50% to 62.5% of the money.
How much can switching models really save?
The reference workload costs $12,000 a month on Claude Opus 5, $9,600 on GPT-5.6 Sol, $4,800 on Claude Sonnet 5, $1,800 on Gemini 3.6 Flash, $960 on Gemini 3.5 Flash-Lite and $528 on GPT-5.6 Luna — a 22.7x spread for identical traffic. Opus 5 to Sonnet 5 alone saves $7,200 a month, 60%.
Does cleaning HTML into Markdown save enough to bother?
Yes, whenever you feed retrieved pages to a model. Measured on four page types with this site's estimator, Markdown conversion removes 47.4% to 67.9% of input tokens and plain text removes 56.0% to 74.3%. Since boilerplate is identical on every page of a site, a crawler pays for the same scaffolding thousands of times over.
Do promotional prices change the plan?
They change the deadline. GPT-5.6 Sol at $4.00/$20.00 runs through at least Nov 21, 2026 and reverts to $5.00/$30.00, taking the reference workload from $9,600 to $13,200 a month. Gemini 3.6 Flash at $0.75/$3.75 runs through Dec 31, 2026 and reverts to $1.50/$7.50, doubling the same workload from $1,800 to $3,600 a month. Model your unit economics on the revert date, not the promo.
Is cutting tokens going to hurt quality?
Not for the mechanical categories. Removing duplicated sentences, boilerplate markup and script blocks does not change what the model is being asked. Removing politeness or hedging language can shift tone, and dropping conversation history can remove context the answer depended on — which is why history strategies should be measured against real transcripts before shipping. Validate prompt surgery with evals rather than assuming it is safe.
How these numbers are produced
Every figure multiplies token volumes by published per-1M rates for the eight models tracked on this site, verified against official provider pricing pages and re-checked weekly; the most recent verification across those sources is 2026-08-23. The reference workload is a stated shape — 4,000 input tokens, 800 output tokens, 10,000 requests a day — not a measurement of any specific application.
"Output ÷ input" is each model's output rate divided by its input rate. "Keep output under" is the reply-to-prompt token ratio at which output becomes the larger line, i.e. the inverse of that multiple. OpenAI's long-context tier is applied above 272,000 input tokens per request; no equivalent published tier exists for Anthropic or Google, so none is applied to those providers.
Caching, batch discounts, free tiers, regional pricing and taxes are excluded throughout — this page prices standard synchronous text calls only. Token estimates quoted for markup, prompt weight and conversation-shape tools come from this site's heuristic estimator, not from provider tokenizers, so expect a few percent of variance either way.
Official sources: OpenAI model docs, Anthropic pricing, Google Gemini API pricing. Always confirm against your invoice before making purchasing decisions.