Conversation Cost Growth Calculator
Every message in a chat pays for every message before it. The API is stateless, so turn N transmits turns 1 to N−1 again — input tokens grow with the square of the conversation length while output grows in a straight line. On the default shape here that means 118,000 input tokens to produce 8,000 output tokens, or 96.6% of the conversation spent re-reading itself. Set your own turn sizes below and compare three ways to carry history forward.
The curve, stated once
Defaults: 200 tokens per user message, 400 tokens per reply, 20 turns, 500 conversations a day.
Model your own conversation shape
Runs entirely in your browser. Full history is priced turn by turn, so OpenAI's 272,000-token long-context rate applies only to the requests that actually exceed it.
How a single conversation snowballs
Input tokens per request under full history. Long conversations are summarised in the middle to keep the table readable.
| Request | Input tokens | Output tokens | Billed this turn |
|---|
Three history strategies, eight models
Monthly assumes conversations run every day. "Saved / year" is the cheaper of the two trimming strategies against full history.
| Model | Full history / conversation | Full / month | Sliding / month | Summarised / month | Saved / year |
|---|
Three conversation shapes, priced
Token arithmetic is exact for these shapes; cost figures apply each model's published per-1M rates, including OpenAI's long-context tier where a request crosses 272,000 input tokens.
| Shape | Input tokens / conversation | Output tokens | Re-sent share | Cheapest monthly | Most expensive monthly |
|---|---|---|---|---|---|
Support chat 200 / 400 tokens · 20 turns · 500 per day | 118,000 | 8,000 | 96.6% | $498.00GPT-5.6 Luna | $11,850.00Claude Opus 5 |
Long-document Q&A 8,000 / 2,000 tokens · 30 turns · 50 per day | 4,590,000 | 60,000 | 94.8% | $1,749.60GPT-5.6 Luna | $36,675.00Claude Opus 5 |
Agent loop 1,200 / 600 tokens · 40 turns · 100 per day | 1,452,000 | 24,000 | 96.7% | $957.60GPT-5.6 Luna | $23,580.00Claude Opus 5 |
What you are actually paying for
The API has no memory. Anything you want the model to "remember", you re-send — and re-pay for — on every turn.
Three ways to carry history forward
All three are priced live in the calculator above. They trade memory against money in very different proportions.
Full history
Send every turn, every time. Perfect recall, quadratic cost. Correct answer when the conversation is short, audit-critical, or the earlier turns really are load-bearing.
Cost: 118,000 input tokens on the default shape.
Sliding window
Keep the last N turns and drop the rest. Cheap to implement and predictable, but it forgets silently — anything older than the window simply disappears from the model's world.
Cost: 55,000 input tokens at five prior turns.
Rolling summary
Compress older turns into a fixed-size summary and send that instead. Keeps the substance while cutting the tokens, at the price of one extra summarisation call and a loss of exact wording.
Cost: 20,000 input tokens at an 800-token summary.
Conversation cost questions
Why does a 20-turn chat cost so much more than one request?
Because turn N re-sends turns 1 to N−1. With 200-token user messages and 400-token replies, the first request sends 200 input tokens and the twentieth sends 11,600 — 58 times more. Across the whole conversation that is 118,000 input tokens to produce 8,000 output tokens, 96.6% of which is re-sent history. On Claude Sonnet 5 that is $0.3160 against $0.0044 for a single turn: 72 times the price for 210 times the tokens, because output is priced five times higher than input.
How much does trimming the history actually save?
On the default shape, keeping only the last five turns cuts input tokens from 118,000 to 55,000 — 53.4% less. Replacing everything older than the current turn with an 800-token rolling summary cuts input to 20,000 tokens, 83.1% less. At 500 conversations a day on Claude Sonnet 5 that is $4,740 a month down to $2,850 or $1,800 — up to $35,280 a year.
Does a bigger context window make this cheaper?
No. A larger window removes the ceiling, not the bill. Claude Opus 5 and Sonnet 5 both accept 1,000,000 tokens, so nothing forces you to drop earlier turns — you simply keep paying the published input rate on every re-sent token. See What Is a Context Window? for how that budget is shared with the reply.
When does OpenAI's long-context surcharge apply to a conversation?
Above 272,000 input tokens in a single request on GPT-5.6 models, where input doubles and output rises 1.5x. A support-sized chat never gets close — even after 40 turns a request is under 24,000 tokens. Document-sized conversations do: in the long-document preset, turn 28 crosses the threshold, and turns 28–30 add $3.5160 to a $19.5600 GPT-5.6 Sol conversation, roughly 18% more for three turns.
Does dropping old turns break the assistant's answers?
It can, which is why the two strategies differ in risk. A sliding window silently loses anything older than the window, so a question about the start of the conversation is answered without that context. A rolling summary keeps the substance while discarding verbatim tokens, so answers degrade more gracefully — at the cost of a summarisation call on your side. Measure the failure rate on real transcripts before shipping either; this page prices the saving, not the regression.
Does this page upload my conversation parameters?
No. The six numbers you enter are processed in your browser and never leave the tab. This page has no form, no backend endpoint, no analytics script, no local storage and no network requests beyond loading its own HTML, CSS and JavaScript. See Privacy.
How these numbers are produced
Turn sizes are yours: input tokens for turn t are modelled as (t − 1) × (user + assistant) + user, which is what an application re-sends when it keeps full history in the request. Output is assistant tokens per turn, once per turn. Token estimates do not come from a provider tokenizer — if you want a measured count rather than a planning shape, paste real text into the AI Token Counter.
Costs multiply those token volumes by published per-1M rates taken from official provider pricing pages and re-verified weekly. OpenAI's long-context tier is applied per request, only to requests above 272,000 input tokens; no equivalent tier is published by Anthropic or Google, so none is applied. Caching, batch discounts, free tiers and the cost of generating the summary itself are excluded throughout.
Official sources: OpenAI model docs, Anthropic pricing, Google Gemini API pricing. Always confirm against your invoice before making purchasing decisions.