Input vs Output Tokens
Short answer: input tokens are what you send, output tokens are what the model writes back, and output costs 5.0x to 8.3x more per token on every model tracked here. Output is typically 10–20% of your tokens but 35–60% of your bill. Below: the per-model multiples, the break-even ratio that tells you which half of the invoice is winning, and what a single sentence of extra output actually costs at scale.
Two meters, one request
Every API call is billed on two separate counters. Confusing them is the most common reason a forecast misses.
What output costs, model by model
The multiple is the output rate divided by the input rate. The break-even share is the point where output stops being the smaller half of the bill.
| Model | Input / 1M | Output / 1M | Output premium | Break-even output share |
|---|---|---|---|---|
GPT-5.6 Sol OpenAI · flagship (promo) | $4.00 | $20.00 | 5.00x | 16.7% |
GPT-5.6 Terra OpenAI · mid | $2.00 | $12.00 | 6.00x | 14.3% |
GPT-5.6 Luna OpenAI · budget | $0.20 | $1.20 | 6.00x | 14.3% |
Claude Opus 5 Anthropic · flagship | $5.00 | $25.00 | 5.00x | 16.7% |
Claude Sonnet 5 Anthropic · mid | $2.00 | $10.00 | 5.00x | 16.7% |
Claude Haiku 4.5 Anthropic · budget | $1.00 | $5.00 | 5.00x | 16.7% |
Gemini 3.6 Flash Google · mid (promo rate) | $0.75 | $3.75 | 5.00x | 16.7% |
Gemini 3.5 Flash-LiteWidest gap Google · budget | $0.30 | $2.50 | 8.33x | 10.7% |
Rates are published per-1M prices as verified 2026-08-23 (OpenAI) and 2026-08-09 to 2026-08-16 (Anthropic and Google). Break-even output share is input price ÷ (input price + output price): above that share of total tokens, output is the larger half of the invoice. Gemini 3.5 Flash-Lite is the only tracked model where the gap exceeds 6x — cheap input, comparatively expensive output, which makes it a poor fit for long generations.
Reading is parallel. Writing is not.
The premium is not an arbitrary markup — it tracks the underlying compute.
Prefill is one pass over everything at once. When your prompt arrives, the model processes all of its tokens in a single parallel forward pass. Doubling the prompt roughly doubles one matrix workload; it does not double the number of sequential steps.
Decode is one pass per token. Generation cannot be parallelised in the same way: each output token requires a full forward pass through the model, and the entire key-value cache for the conversation has to be read again to produce it. A 2,000-token answer is 2,000 sequential passes; a 2,000-token prompt is one.
Long generations hold hardware longer. Accelerator memory stays occupied for the whole generation, and a slot held for 10 seconds is a slot another request could not use. Time-based opportunity cost lands on the output side of the rate card.
There is also no cheaper tier for output. On the rate cards tracked here, output has exactly one published price per model, while input has tiers: promotional rates, and on OpenAI a second, higher rate above 272,000 input tokens. See the OpenAI rate card for where that threshold bites.
15% output: one shape, eight bills
A realistic production request — 10,000 tokens per call — 8,500 input and 1,500 output — at 5,000 requests a day. The blended rate is what you actually pay per million tokens.
| Model | Output premium | Output share of spend | Blended rate / 1M | Cost per request | Monthly (150,000 requests) |
|---|---|---|---|---|---|
GPT-5.6 Sol OpenAI · flagship (promo) | 5.00x | 46.9% | $6.40 | $0.0640 | $9,600.00 |
GPT-5.6 Terra OpenAI · mid | 6.00x | 51.4% | $3.50 | $0.0350 | $5,250.00 |
GPT-5.6 LunaLowest OpenAI · budget | 6.00x | 51.4% | $0.35 | $0.00350 | $525.00 |
Claude Opus 5 Anthropic · flagship | 5.00x | 46.9% | $8.00 | $0.0800 | $12,000.00 |
Claude Sonnet 5 Anthropic · mid | 5.00x | 46.9% | $3.20 | $0.0320 | $4,800.00 |
Claude Haiku 4.5 Anthropic · budget | 5.00x | 46.9% | $1.60 | $0.0160 | $2,400.00 |
Gemini 3.6 Flash Google · mid (promo rate) | 5.00x | 46.9% | $1.20 | $0.0120 | $1,800.00 |
Gemini 3.5 Flash-Lite Google · budget | 8.33x | 59.5% | $0.63 | $0.00630 | $945.00 |
Output is 15% of the tokens and 46.9% – 59.5% of the money. The monthly spread is 22.9x — $525.00 on GPT-5.6 Luna against $12,000.00 on Claude Opus 5 for byte-identical traffic. Blended rate is the input and output rates weighted by this mix; it is the only number comparable across models with different premiums. Excludes caching, batch, tool use and taxes.
Split your own bill
Set the shape of one request and see where the money goes on all eight models — including the blended rate you are really paying.
| Model | Output premium | Break-even output share | Output share of spend | Blended rate / 1M | Cost per request | Monthly |
|---|
Output grows with volume. Input grows with habit.
Two things push the output share up, and neither shows up in a unit-price comparison.
Verbosity is a per-request tax. At 5,000 requests a day — 150,000 a month — every 100 extra output tokens per request costs $300.00 a month on GPT-5.6 Sol, $375.00 on Claude Opus 5 and $18.00 on GPT-5.6 Luna. The same 100 tokens of extra input cost $60.00, $75.00 and $3.00. Trimming one sentence of boilerplate from a high-volume reply is worth more than rewriting the prompt.
Chat re-sends its input every turn. Conversation history is input, and it is billed again on every turn, so a 20-turn conversation sends 20 copies of the growing transcript. On a measured support shape, 96.6% of input tokens are re-sent history — see the Conversation Cost Growth Calculator. That pushes the token mix toward input while pushing the bill toward output, because each turn still pays a fresh output premium.
Reasoning tokens are output. On the providers that expose extended thinking or reasoning, those tokens are metered at the output rate. A model that "thinks" for 3,000 tokens before answering bills those 3,000 tokens like any other generated text.
What one maxed-out response costs
Every model caps a single reply well below its context window. These are the ceilings, priced at published output rates — the most one call can cost you on the generation side.
| Model | Output cap | Cost of a full-length reply | At OpenAI's long-context output rate |
|---|---|---|---|
GPT-5.6 Sol OpenAI · flagship (promo) | 128,000 | $2.56 | $3.84 |
GPT-5.6 Terra OpenAI · mid | 128,000 | $1.54 | $2.30 |
GPT-5.6 Luna OpenAI · budget | 128,000 | $0.1536 | $0.2304 |
Claude Opus 5 Anthropic · flagship | 128,000 | $3.20 | n/a — flat rate |
Claude Sonnet 5 Anthropic · mid | 128,000 | $1.28 | n/a — flat rate |
Claude Haiku 4.5 Anthropic · budget | 64,000 | $0.3200 | n/a — flat rate |
Gemini 3.6 Flash Google · mid (promo rate) | 65,536 | $0.2458 | n/a — flat rate |
Gemini 3.5 Flash-Lite Google · budget | 65,536 | $0.1638 | n/a — flat rate |
Output caps are provider-published maximums, not the length you should target. Only OpenAI applies a second, higher output rate, and only above 272,000 input tokens in the same request. A runaway generation at 5,000 requests a day would multiply these figures by 150,000 — which is why a deliberate max_tokens matters more than any prompt edit.
Six ways to spend less on output
Ranked by how often each one pays off in practice.
1. Set max_tokens on purpose. It caps rather than reduces, but it converts an unbounded risk into a bounded one — $3.20 worst case on Opus 5 instead of whatever a stuck loop would have produced.
2. Ask for a length, not for brevity. "Answer in 150 words" and "three bullets, one sentence each" are enforceable; "be concise" is not. A stated ceiling is the cheapest output control there is.
3. Use stop sequences. If the model pads after the useful part, a stop string ends generation there instead of paying for the tail.
4. Stop asking for what you discard. Do not request a restatement of the question, a summary you already have, or reasoning you will not display. Measure what your prompts actually contain with the Prompt Weight Analyzer.
5. Move long generations to a cheaper output tier. Output rates span $1.20 to $25.00 per 1M — a 20.8x gap. Drafting on GPT-5.6 Luna and reserving a flagship for the final pass is usually cheaper than generating everything on the flagship.
6. Do not pay twice for the same answer. Identical requests are pure waste. Cache at the application layer, and remember that streaming changes nothing about what you are billed.
For the full ranking across input and output together, see How to Reduce LLM API Costs.
Input and output token questions
What is the difference between input and output tokens?
Input tokens are everything you send in one request: system prompt, tool schemas, retrieved documents, chat history and the current message. Output tokens are everything the model generates in reply. They are metered separately and priced separately — output costs 5.0 to 8.3 times more per token on the eight models tracked here.
Why are output tokens more expensive than input tokens?
Reading a prompt is one parallel forward pass over every token at once, and that work can be cached and reused. Writing a reply is sequential: each output token requires its own full pass through the model, with the whole conversation's key-value cache read again for every one. A long generation also holds accelerator memory for the entire time it runs. Providers pass that asymmetry on as a 5x to 8.3x output premium.
How much more do output tokens cost?
On the eight models tracked here, output costs 5.00x input on GPT-5.6 Sol, Claude Opus 5, Claude Sonnet 5, Claude Haiku 4.5 and Gemini 3.6 Flash; 6.00x on GPT-5.6 Terra and GPT-5.6 Luna; and 8.33x on Gemini 3.5 Flash-Lite, which has the widest gap of any tracked model.
What ratio of input to output tokens should I aim for?
For a balanced bill, keep output at or below the break-even share, which is input price divided by input plus output price. That is 16.7% of total tokens on the 5x models, 14.3% on the two 6x models and 10.7% on Gemini 3.5 Flash-Lite. Above that line, output is the larger half of the invoice.
Do system prompts and chat history count as input tokens?
Yes, and they are re-sent on every turn. On a 20-turn support conversation, 96.6% of input tokens are previously exchanged history being paid for again. Use a sliding window or a rolling summary instead — measured on the Conversation Cost Growth Calculator, those cut input by 53.4% and 83.1% respectively.
Does setting a lower max_tokens reduce my bill?
It caps it rather than reduces it — you are billed for tokens actually generated, not for the ceiling you set. A cap does protect you from runaway generations: a single maxed-out response costs $2.56 on GPT-5.6 Sol, $3.20 on Claude Opus 5 and $0.1638 on Gemini 3.5 Flash-Lite. Streaming changes nothing about cost.
Where these numbers come from.
Rates are published per-1M input and output prices, re-verified weekly against official provider documentation: OpenAI 2026-08-23, Anthropic and Google 2026-08-09 to 2026-08-16. Premium multiples and break-even shares are arithmetic on those two published numbers, not estimates. OpenAI long-context rules are applied automatically above 272,000 input tokens, where input doubles and output rises 1.5x. All token volumes, mixes and request rates on this page are illustrative shapes, not measurements of any specific workload. Caching, batch, tool use and taxes are excluded throughout — always confirm against your own invoice before making purchasing decisions.
Official sources: OpenAI model docs, Anthropic pricing, Google Gemini API pricing.