Guide · Last verified 2026-08-23

Input vs Output Tokens

Short answer: input tokens are what you send, output tokens are what the model writes back, and output costs 5.0x to 8.3x more per token on every model tracked here. Output is typically 10–20% of your tokens but 35–60% of your bill. Below: the per-model multiples, the break-even ratio that tells you which half of the invoice is winning, and what a single sentence of extra output actually costs at scale.

Definition

Two meters, one request

Every API call is billed on two separate counters. Confusing them is the most common reason a forecast misses.

Input is everything you sendSystem prompt, tool and function schemas, retrieved documents, chat history, and the current user message — all metered at the input rate.
Output is everything you get backThe visible reply, plus any tool-call arguments, structured JSON and — on providers that expose them — reasoning or thinking tokens.
They share a ceiling, not a priceInput and output draw on the same context window, but output is separately capped at 64,000 to 128,000 tokens depending on the model.
The asymmetry is deliberateReading a prompt is one parallel pass; writing a reply is one sequential pass per token. Rate cards price that difference at 5x–8.3x.
The premium

What output costs, model by model

The multiple is the output rate divided by the input rate. The break-even share is the point where output stops being the smaller half of the bill.

ModelInput / 1MOutput / 1MOutput premiumBreak-even output share
GPT-5.6 Sol
OpenAI · flagship (promo)
$4.00$20.005.00x16.7%
GPT-5.6 Terra
OpenAI · mid
$2.00$12.006.00x14.3%
GPT-5.6 Luna
OpenAI · budget
$0.20$1.206.00x14.3%
Claude Opus 5
Anthropic · flagship
$5.00$25.005.00x16.7%
Claude Sonnet 5
Anthropic · mid
$2.00$10.005.00x16.7%
Claude Haiku 4.5
Anthropic · budget
$1.00$5.005.00x16.7%
Gemini 3.6 Flash
Google · mid (promo rate)
$0.75$3.755.00x16.7%
Gemini 3.5 Flash-LiteWidest gap
Google · budget
$0.30$2.508.33x10.7%

Rates are published per-1M prices as verified 2026-08-23 (OpenAI) and 2026-08-09 to 2026-08-16 (Anthropic and Google). Break-even output share is input price ÷ (input price + output price): above that share of total tokens, output is the larger half of the invoice. Gemini 3.5 Flash-Lite is the only tracked model where the gap exceeds 6x — cheap input, comparatively expensive output, which makes it a poor fit for long generations.

Why

Reading is parallel. Writing is not.

The premium is not an arbitrary markup — it tracks the underlying compute.

Prefill is one pass over everything at once. When your prompt arrives, the model processes all of its tokens in a single parallel forward pass. Doubling the prompt roughly doubles one matrix workload; it does not double the number of sequential steps.

Decode is one pass per token. Generation cannot be parallelised in the same way: each output token requires a full forward pass through the model, and the entire key-value cache for the conversation has to be read again to produce it. A 2,000-token answer is 2,000 sequential passes; a 2,000-token prompt is one.

Long generations hold hardware longer. Accelerator memory stays occupied for the whole generation, and a slot held for 10 seconds is a slot another request could not use. Time-based opportunity cost lands on the output side of the rate card.

There is also no cheaper tier for output. On the rate cards tracked here, output has exactly one published price per model, while input has tiers: promotional rates, and on OpenAI a second, higher rate above 272,000 input tokens. See the OpenAI rate card for where that threshold bites.

1 pass vs N passesA 10,000-token prompt is one parallel workload; a 1,500-token answer is 1,500 sequential ones.
Premium: 5.00x – 8.33xNarrow on the flagship tiers, widest on Gemini 3.5 Flash-Lite, where output is 8.33x input.
Output has no discount tierInput has promo rates and long-context rules; output is a single published number per model.
Caps, not discountsOutput ceilings of 64,000–128,000 tokens limit the worst case; they do not lower the rate.
A worked mix

15% output: one shape, eight bills

A realistic production request — 10,000 tokens per call — 8,500 input and 1,500 output — at 5,000 requests a day. The blended rate is what you actually pay per million tokens.

ModelOutput premiumOutput share of spendBlended rate / 1MCost per requestMonthly (150,000 requests)
GPT-5.6 Sol
OpenAI · flagship (promo)
5.00x46.9%$6.40$0.0640$9,600.00
GPT-5.6 Terra
OpenAI · mid
6.00x51.4%$3.50$0.0350$5,250.00
GPT-5.6 LunaLowest
OpenAI · budget
6.00x51.4%$0.35$0.00350$525.00
Claude Opus 5
Anthropic · flagship
5.00x46.9%$8.00$0.0800$12,000.00
Claude Sonnet 5
Anthropic · mid
5.00x46.9%$3.20$0.0320$4,800.00
Claude Haiku 4.5
Anthropic · budget
5.00x46.9%$1.60$0.0160$2,400.00
Gemini 3.6 Flash
Google · mid (promo rate)
5.00x46.9%$1.20$0.0120$1,800.00
Gemini 3.5 Flash-Lite
Google · budget
8.33x59.5%$0.63$0.00630$945.00

Output is 15% of the tokens and 46.9% – 59.5% of the money. The monthly spread is 22.9x — $525.00 on GPT-5.6 Luna against $12,000.00 on Claude Opus 5 for byte-identical traffic. Blended rate is the input and output rates weighted by this mix; it is the only number comparable across models with different premiums. Excludes caching, batch, tool use and taxes.

Live calculator

Split your own bill

Set the shape of one request and see where the money goes on all eight models — including the blended rate you are really paying.

0output tokens per request
0input tokens per request
—of the bill is output
—monthly spread across 8 models
Cheapest at this mix—
Most expensive at this mix—
ModelOutput premiumBreak-even output shareOutput share of spendBlended rate / 1MCost per requestMonthly
Blended rate: your mix weighted across both rates — the only figure you can compare against a flat per-token quote. Long-context: above 272,000 input tokens OpenAI doubles input and raises output 1.5x, which narrows its premium to 3.75x. Pricing data verified —
The compounding part

Output grows with volume. Input grows with habit.

Two things push the output share up, and neither shows up in a unit-price comparison.

Verbosity is a per-request tax. At 5,000 requests a day — 150,000 a month — every 100 extra output tokens per request costs $300.00 a month on GPT-5.6 Sol, $375.00 on Claude Opus 5 and $18.00 on GPT-5.6 Luna. The same 100 tokens of extra input cost $60.00, $75.00 and $3.00. Trimming one sentence of boilerplate from a high-volume reply is worth more than rewriting the prompt.

Chat re-sends its input every turn. Conversation history is input, and it is billed again on every turn, so a 20-turn conversation sends 20 copies of the growing transcript. On a measured support shape, 96.6% of input tokens are re-sent history — see the Conversation Cost Growth Calculator. That pushes the token mix toward input while pushing the bill toward output, because each turn still pays a fresh output premium.

Reasoning tokens are output. On the providers that expose extended thinking or reasoning, those tokens are metered at the output rate. A model that "thinks" for 3,000 tokens before answering bills those 3,000 tokens like any other generated text.

+100 output tokens: $300/moOn GPT-5.6 Sol at 5,000 requests a day. The same 100 input tokens: $60/mo — one fifth.
+100 output tokens: $18/moOn GPT-5.6 Luna — the identical reply, 16.7x cheaper than the flagship.
96.6% of input is historyRe-sent chat turns on a 20-turn support conversation; a rolling summary cuts input 83.1%.
Thinking is billedReasoning and extended-thinking tokens are metered as output where providers expose them.
Worst case

What one maxed-out response costs

Every model caps a single reply well below its context window. These are the ceilings, priced at published output rates — the most one call can cost you on the generation side.

ModelOutput capCost of a full-length replyAt OpenAI's long-context output rate
GPT-5.6 Sol
OpenAI · flagship (promo)
128,000$2.56$3.84
GPT-5.6 Terra
OpenAI · mid
128,000$1.54$2.30
GPT-5.6 Luna
OpenAI · budget
128,000$0.1536$0.2304
Claude Opus 5
Anthropic · flagship
128,000$3.20n/a — flat rate
Claude Sonnet 5
Anthropic · mid
128,000$1.28n/a — flat rate
Claude Haiku 4.5
Anthropic · budget
64,000$0.3200n/a — flat rate
Gemini 3.6 Flash
Google · mid (promo rate)
65,536$0.2458n/a — flat rate
Gemini 3.5 Flash-Lite
Google · budget
65,536$0.1638n/a — flat rate

Output caps are provider-published maximums, not the length you should target. Only OpenAI applies a second, higher output rate, and only above 272,000 input tokens in the same request. A runaway generation at 5,000 requests a day would multiply these figures by 150,000 — which is why a deliberate max_tokens matters more than any prompt edit.

Levers

Six ways to spend less on output

Ranked by how often each one pays off in practice.

1. Set max_tokens on purpose. It caps rather than reduces, but it converts an unbounded risk into a bounded one — $3.20 worst case on Opus 5 instead of whatever a stuck loop would have produced.

2. Ask for a length, not for brevity. "Answer in 150 words" and "three bullets, one sentence each" are enforceable; "be concise" is not. A stated ceiling is the cheapest output control there is.

3. Use stop sequences. If the model pads after the useful part, a stop string ends generation there instead of paying for the tail.

4. Stop asking for what you discard. Do not request a restatement of the question, a summary you already have, or reasoning you will not display. Measure what your prompts actually contain with the Prompt Weight Analyzer.

5. Move long generations to a cheaper output tier. Output rates span $1.20 to $25.00 per 1M — a 20.8x gap. Drafting on GPT-5.6 Luna and reserving a flagship for the final pass is usually cheaper than generating everything on the flagship.

6. Do not pay twice for the same answer. Identical requests are pure waste. Cache at the application layer, and remember that streaming changes nothing about what you are billed.

For the full ranking across input and output together, see How to Reduce LLM API Costs.

20.8x output spread$1.20 / 1M on GPT-5.6 Luna to $25.00 / 1M on Claude Opus 5 — the widest single lever on the output side.
Cap, then measureA max_tokens ceiling plus a per-request token count turns the worst case into a line item.
Trim input tooCleaning scraped markup removes 47–68% of input tokens: HTML Token Savings.
Price your real mixToken volumes into monthly spend: API Cost Calculator.
FAQ

Input and output token questions

What is the difference between input and output tokens?

Input tokens are everything you send in one request: system prompt, tool schemas, retrieved documents, chat history and the current message. Output tokens are everything the model generates in reply. They are metered separately and priced separately — output costs 5.0 to 8.3 times more per token on the eight models tracked here.

Why are output tokens more expensive than input tokens?

Reading a prompt is one parallel forward pass over every token at once, and that work can be cached and reused. Writing a reply is sequential: each output token requires its own full pass through the model, with the whole conversation's key-value cache read again for every one. A long generation also holds accelerator memory for the entire time it runs. Providers pass that asymmetry on as a 5x to 8.3x output premium.

How much more do output tokens cost?

On the eight models tracked here, output costs 5.00x input on GPT-5.6 Sol, Claude Opus 5, Claude Sonnet 5, Claude Haiku 4.5 and Gemini 3.6 Flash; 6.00x on GPT-5.6 Terra and GPT-5.6 Luna; and 8.33x on Gemini 3.5 Flash-Lite, which has the widest gap of any tracked model.

What ratio of input to output tokens should I aim for?

For a balanced bill, keep output at or below the break-even share, which is input price divided by input plus output price. That is 16.7% of total tokens on the 5x models, 14.3% on the two 6x models and 10.7% on Gemini 3.5 Flash-Lite. Above that line, output is the larger half of the invoice.

Do system prompts and chat history count as input tokens?

Yes, and they are re-sent on every turn. On a 20-turn support conversation, 96.6% of input tokens are previously exchanged history being paid for again. Use a sliding window or a rolling summary instead — measured on the Conversation Cost Growth Calculator, those cut input by 53.4% and 83.1% respectively.

Does setting a lower max_tokens reduce my bill?

It caps it rather than reduces it — you are billed for tokens actually generated, not for the ceiling you set. A cap does protect you from runaway generations: a single maxed-out response costs $2.56 on GPT-5.6 Sol, $3.20 on Claude Opus 5 and $0.1638 on Gemini 3.5 Flash-Lite. Streaming changes nothing about cost.

Methodology

Where these numbers come from.

Rates are published per-1M input and output prices, re-verified weekly against official provider documentation: OpenAI 2026-08-23, Anthropic and Google 2026-08-09 to 2026-08-16. Premium multiples and break-even shares are arithmetic on those two published numbers, not estimates. OpenAI long-context rules are applied automatically above 272,000 input tokens, where input doubles and output rises 1.5x. All token volumes, mixes and request rates on this page are illustrative shapes, not measurements of any specific workload. Caching, batch, tool use and taxes are excluded throughout — always confirm against your own invoice before making purchasing decisions.

Official sources: OpenAI model docs, Anthropic pricing, Google Gemini API pricing.

Price your own mixInput, output and volume into monthly spend: API Cost Calculator.
The full cost playbookSeven levers ranked by return, from re-sent history to a 22.7x model spread: How to Reduce LLM API Costs.
Stop paying twiceChat re-sends its history every turn — 96.6% of input tokens on a 20-turn conversation: Conversation Cost Growth Calculator.
What a token isDefinition, conversion by content type, and why output costs 5–8x input: What Is an AI Token?
Count before you sendPaste text and see tokens, words, characters and per-model input cost: Token Counter.
A 200-page PDF133,333 tokens, and the one model where it nearly does not fit: Can a 200-Page PDF Fit in an LLM?
Windows and capsOutput caps from 64,000 to 128,000 and what they cost to fill: What Is a Context Window?
Provider rate cardsOpenAI · Claude · Gemini · all three.
Words to tokensConversion by content type: How Many Tokens in 1,000 Words?
How we verifySources, weekly cadence and what our estimates exclude: Methodology · Pricing changelog.