Guide · Last verified 2026-08-23

How to Reduce LLM API Costs

There is no single trick — there is a bill shape, and seven levers that act on different parts of it. An LLM invoice is tokens in × rate, plus tokens out × a much higher rate, and almost every real overspend comes from paying one of those twice. Below: the seven levers ranked by typical return, each priced against the eight models tracked here, and a calculator you can point at your own volume.

Quick answer

The seven levers, ranked

Ordered by what they typically return on an input-heavy production bill. Each row links to the page that measures it.

1. Stop re-sending chat historyUp to 83.1% of input tokens in a 20-turn conversation. The largest single lever for anything multi-turn.
2. Trim the prompt itself27.9%–53.5% flagged as avoidable across three real sample prompts — duplicated instructions, repeated examples, filler.
3. Convert HTML before you send it47.4%–67.9% of retrieved-page tokens are markup you are billed for at content rates.
4. Buy the cheapest model that passes your barIdentical traffic ranges from $528 to $12,000 a month — a 22.7x spread.
5. Cap the outputOutput is priced 5x–8.3x input. A thousand unnecessary reply tokens costs $6,000 a month on GPT-5.6 Sol at 10,000 requests a day.
6. Chunk past OpenAI's 272K thresholdAbove it, GPT-5.6 input doubles — a 1M-token prompt costs $8.00/1M instead of $4.00, and four chunks restore the base rate.
7. Budget to the revert date, not the promoTwo tracked rates are promotional: Sol reverts after Nov 21, 2026 (+37.5% on the reference bill), Gemini 3.6 Flash after Dec 31, 2026 (+100%).
Diagnose

Price your own bill shape

The defaults below are the reference workload used throughout this page: 4,000 input tokens, 800 output tokens, 10,000 requests a day. Change them to your own numbers.

30-day requests300,000
Lowest 30-day estimate—
Standard rates: text token pricing. Excluded: caching, batch, tools and taxes. Pricing data verified —
ModelInputOutputPer request30 days1 year

Informational estimate only. Provider pricing can include caching, batch processing, tool calls, regional pricing, and other rules not represented here.

Measured

One workload, eight models

Reference workload: 4,000 input and 800 output tokens per request, 10,000 requests a day — 300,000 requests and 1.44 billion tokens a month.

ModelInput / 1MOutput / 1MOutput ÷ inputKeep output under30-day costOutput share of bill
GPT-5.6 Sol
OpenAI · promotional to Nov 21, 2026
$4.00$20.005.0x20.0% of input tokens$9,600.0050.0%
GPT-5.6 Terra
OpenAI
$2.00$12.006.0x16.7% of input tokens$5,280.0054.5%
GPT-5.6 Luna
OpenAI
$0.20$1.206.0x16.7% of input tokens$528.0054.5%
Claude Opus 5
Anthropic · flagship rates
$5.00$25.005.0x20.0% of input tokens$12,000.0050.0%
Claude Sonnet 5
Anthropic
$2.00$10.005.0x20.0% of input tokens$4,800.0050.0%
Claude Haiku 4.5
Anthropic · 200K window
$1.00$5.005.0x20.0% of input tokens$2,400.0050.0%
Gemini 3.6 Flash
Google · promotional to Dec 31, 2026
$0.75$3.755.0x20.0% of input tokens$1,800.0050.0%
Gemini 3.5 Flash-Lite
Google
$0.30$2.508.3x12.0% of input tokens$960.0062.5%
A 22.7x spread on identical traffic$528.00 a month on GPT-5.6 Luna against $12,000.00 on Claude Opus 5. Nothing else on this page moves the needle that far.
Output is 16.7% of the tokens and half the money800 output tokens against 4,000 input tokens still take 50.0%–62.5% of the bill on every model here.
Flash-Lite is the tightest trapAt 8.3x, output overtakes input once replies exceed 12.0% of prompt tokens — before almost any other model cares.
Terra costs more than Sonnet 5 hereIdentical $2.00 input rates, but $12.00 against $10.00 output: on this output-weighted shape GPT-5.6 Terra bills $5,280.00 against Claude Sonnet 5's $4,800.00.
Levers

Seven moves, and what each returns

Each lever acts on a different line of the bill. The numbers are computed from the same eight published rate cards — no estimates of provider behaviour beyond the arithmetic shown.

Up to 83.1% of input

1. Stop re-sending chat history

Chat APIs are stateless, so turn N re-transmits turns 1 to N−1. A 20-turn support conversation with 200-token messages and 400-token replies sends 118,000 input tokens to produce 8,000 output tokens — 96.6% of the input is re-sent history. Keeping five prior turns cuts input to 55,000; an 800-token rolling summary cuts it to 20,000.

Measure it: Conversation Cost Growth Calculator

27.9%–53.5% of input

2. Trim the prompt itself

Prompts accumulate duplicated instructions, repeated few-shot blocks, markup and politeness filler. Across three sample prompts measured on this site, the avoidable share ran from 27.9% to 53.5% — and the largest single category was different every time, so there is no universal fix to apply blind.

Measure it: Prompt Weight Analyzer

47.4%–67.9% of input

3. Convert HTML before sending it

Retrieved pages bill every div, href and aria-label at content rates. On four representative page types, Markdown conversion removed 47.4%–67.9% of input tokens and plain text removed 56.0%–74.3%. Boilerplate repeats on every fetch, so crawlers pay for it thousands of times.

Measure it: HTML → Markdown Token Savings

Up to 22.7x

4. Buy the cheapest model that passes your bar

Model choice dwarfs prompt hygiene. Moving the reference workload from Claude Opus 5 to Sonnet 5 saves $7,200.00 a month; moving it to Gemini 3.5 Flash-Lite saves $11,040.00. Route easy traffic to Luna-class models and reserve flagship rates for requests that actually need them — that routing decision is a bigger lever than any prompt edit.

Compare: LLM API Pricing Comparison · OpenAI vs Claude

5x–8.3x input

5. Cap the output

Every model here charges far more to generate than to read. A thousand avoidable output tokens per request costs $6,000.00 a month on GPT-5.6 Sol, $7,500.00 on Claude Opus 5 and $1,125.00 on Gemini 3.6 Flash at 10,000 requests a day. Set max_tokens, use stop sequences, and ask for structured brevity rather than prose you will discard.

Why the premium exists: What Is an AI Token?

Halves the input rate

6. Chunk past OpenAI's 272K threshold

GPT-5.6 models apply higher long-context rates above 272,000 input tokens, where input doubles and output rises 1.5x. A 1M-token prompt therefore bills $8.00 per million input on Sol instead of $4.00 — more than Claude Opus 5's $5.00. Splitting the same content into four 250,000-token requests restores the base rate.

Full arithmetic: How Many Words Is 1M Tokens?

+37.5% and +100%

7. Budget to the revert date, not the promo

Two of the eight rates are temporary. GPT-5.6 Sol at $4.00/$20.00 holds through at least Nov 21, 2026 and reverts to $5.00/$30.00 — the reference workload goes from $9,600.00 to $13,200.00 a month. Gemini 3.6 Flash at $0.75/$3.75 holds through Dec 31, 2026 and reverts to $1.50/$7.50, doubling it from $1,800.00 to $3,600.00.

Track it: Pricing changelog · OpenAI API Pricing

Order of operations

Do them in this order

Cheapest reliable wins first, changes that need evaluation last.

Week 1 — make the bill legibleCount real tokens on real traffic with the Token Counter, then split input from output per endpoint. Most teams discover one endpoint driving most of the spend.
Week 2 — remove mechanical wasteStrip markup with HTML → Markdown and duplicates with the Prompt Weight Analyzer. Neither should change what the model is asked to do.
Week 3 — re-price the routingSend each traffic class to the cheapest model that clears your quality bar, using the pricing comparison as the shortlist.
Week 4 — restrain generationCap max_tokens, add stop sequences, and cut the verbose instructions that invite long answers. Watch Flash-Lite first: 12.0% of input tokens is a tight output budget.
Then — restructure multi-turn flowsHistory strategies need eval coverage, so they come last and pay the most. Start with the Conversation Cost Growth Calculator to size the prize before touching code.
Before each renewal deadlineRe-run your numbers against the post-promo rates: Nov 21, 2026 for GPT-5.6 Sol, Dec 31, 2026 for Gemini 3.6 Flash. Both are logged in the changelog.
FAQ

LLM cost reduction questions

What is the fastest way to reduce LLM API costs?

For multi-turn applications, stop re-sending the whole chat history: a 20-turn conversation with 200-token messages and 400-token replies sends 118,000 input tokens to produce 8,000 output tokens, and 96.6% of that input is re-sent history. Keeping five prior turns cuts input tokens to 55,000 (53.4% less); a rolling summary cuts them to 20,000 (83.1% less). For single-request applications the fastest win is usually model choice — the same workload costs $528 to $12,000 a month across the eight models tracked here.

Why is my output cost so high when I generate few tokens?

Because output is priced 5x to 8.3x input on every tracked model. Output overtakes input as the larger line once reply tokens exceed 20% of prompt tokens on six models, 16.7% on GPT-5.6 Terra and Luna, and 12% on Gemini 3.5 Flash-Lite. In the reference workload above, output is 16.7% of the tokens but 50% to 62.5% of the money.

How much can switching models really save?

The reference workload costs $12,000 a month on Claude Opus 5, $9,600 on GPT-5.6 Sol, $4,800 on Claude Sonnet 5, $1,800 on Gemini 3.6 Flash, $960 on Gemini 3.5 Flash-Lite and $528 on GPT-5.6 Luna — a 22.7x spread for identical traffic. Opus 5 to Sonnet 5 alone saves $7,200 a month, 60%.

Does cleaning HTML into Markdown save enough to bother?

Yes, whenever you feed retrieved pages to a model. Measured on four page types with this site's estimator, Markdown conversion removes 47.4% to 67.9% of input tokens and plain text removes 56.0% to 74.3%. Since boilerplate is identical on every page of a site, a crawler pays for the same scaffolding thousands of times over.

Do promotional prices change the plan?

They change the deadline. GPT-5.6 Sol at $4.00/$20.00 runs through at least Nov 21, 2026 and reverts to $5.00/$30.00, taking the reference workload from $9,600 to $13,200 a month. Gemini 3.6 Flash at $0.75/$3.75 runs through Dec 31, 2026 and reverts to $1.50/$7.50, doubling the same workload from $1,800 to $3,600 a month. Model your unit economics on the revert date, not the promo.

Is cutting tokens going to hurt quality?

Not for the mechanical categories. Removing duplicated sentences, boilerplate markup and script blocks does not change what the model is being asked. Removing politeness or hedging language can shift tone, and dropping conversation history can remove context the answer depended on — which is why history strategies should be measured against real transcripts before shipping. Validate prompt surgery with evals rather than assuming it is safe.

Methodology

How these numbers are produced

Every figure multiplies token volumes by published per-1M rates for the eight models tracked on this site, verified against official provider pricing pages and re-checked weekly; the most recent verification across those sources is 2026-08-23. The reference workload is a stated shape — 4,000 input tokens, 800 output tokens, 10,000 requests a day — not a measurement of any specific application.

"Output ÷ input" is each model's output rate divided by its input rate. "Keep output under" is the reply-to-prompt token ratio at which output becomes the larger line, i.e. the inverse of that multiple. OpenAI's long-context tier is applied above 272,000 input tokens per request; no equivalent published tier exists for Anthropic or Google, so none is applied to those providers.

Caching, batch discounts, free tiers, regional pricing and taxes are excluded throughout — this page prices standard synchronous text calls only. Token estimates quoted for markup, prompt weight and conversation-shape tools come from this site's heuristic estimator, not from provider tokenizers, so expect a few percent of variance either way.

Official sources: OpenAI model docs, Anthropic pricing, Google Gemini API pricing. Always confirm against your invoice before making purchasing decisions.

Size the multi-turn prizeHistory growth curves and three strategies: Conversation Cost Growth Calculator.
Price any fixed workloadPer-request, monthly and annual across all eight models: API Cost Calculator.
Weigh your promptFind duplicated instructions, markup and filler by category: Prompt Weight Analyzer.
Clean retrieved pagesHTML → Markdown Token Savings.
Check the fit before sendingContext Window Checker · Token Counter.
Provider rate cardsAll providers · OpenAI · Anthropic · Gemini.
Why 1M in one request is a trapHow Many Words Is 1M Tokens?
How we verifySources, weekly cadence and what our estimates exclude: Methodology · Pricing changelog.