Tool · Last verified 2026-08-23

Conversation Cost Growth Calculator

Every message in a chat pays for every message before it. The API is stateless, so turn N transmits turns 1 to N−1 again — input tokens grow with the square of the conversation length while output grows in a straight line. On the default shape here that means 118,000 input tokens to produce 8,000 output tokens, or 96.6% of the conversation spent re-reading itself. Set your own turn sizes below and compare three ways to carry history forward.

Quick answer

The curve, stated once

Defaults: 200 tokens per user message, 400 tokens per reply, 20 turns, 500 conversations a day.

Input is quadratic, output is linearTurn 1 sends 200 tokens; turn 20 sends 11,600 — 58x the first request, for the same size of question.
96.6% of conversation input is history118,000 input tokens per conversation, of which only 4,000 are new user text. The rest is tokens you already paid for once.
A conversation costs 72x one turn$0.0044 for a single turn against $0.3160 for all twenty on Claude Sonnet 5 — 210x the tokens, but only 72x the money, because output is the pricier half.
History strategy is the biggest lever on this pageA rolling summary cuts input from 118,000 to 20,000 tokens — 83.1% — worth $35,280 a year on Sonnet 5 at 500 conversations a day.
Calculator

Model your own conversation shape

Runs entirely in your browser. Full history is priced turn by turn, so OpenAI's 272,000-token long-context rate applies only to the requests that actually exceed it.

0input tokens per conversation
0output tokens generated
0tokens billed in total
—of input is re-sent history
Conversation shapes:
First → last turn input—
Sliding window cuts input—
Rolling summary cuts input—
Turn by turn

How a single conversation snowballs

Input tokens per request under full history. Long conversations are summarised in the middle to keep the table readable.

RequestInput tokensOutput tokensBilled this turn
Priced

Three history strategies, eight models

Monthly assumes conversations run every day. "Saved / year" is the cheaper of the two trimming strategies against full history.

ModelFull history / conversationFull / monthSliding / monthSummarised / monthSaved / year
Cheapest per month, full history—
Most expensive per month, full history—
Largest annual saving available—
Model only: this prices the tokens, not caching, batch discounts or a summariser's own cost. Private: no upload, no network request. Pricing data verified —
Measured

Three conversation shapes, priced

Token arithmetic is exact for these shapes; cost figures apply each model's published per-1M rates, including OpenAI's long-context tier where a request crosses 272,000 input tokens.

ShapeInput tokens / conversationOutput tokensRe-sent shareCheapest monthlyMost expensive monthly
Support chat
200 / 400 tokens · 20 turns · 500 per day
118,0008,00096.6%$498.00GPT-5.6 Luna$11,850.00Claude Opus 5
Long-document Q&A
8,000 / 2,000 tokens · 30 turns · 50 per day
4,590,00060,00094.8%$1,749.60GPT-5.6 Luna$36,675.00Claude Opus 5
Agent loop
1,200 / 600 tokens · 40 turns · 100 per day
1,452,00024,00096.7%$957.60GPT-5.6 Luna$23,580.00Claude Opus 5
The support shape is a 23.8x spread$498.00 a month on GPT-5.6 Luna against $11,850.00 on Claude Opus 5 for identical traffic — the model choice dwarfs the trimming strategy on this shape.
Only the document shape trips the OpenAI surchargeTurn 28 crosses 272,000 input tokens, and those three turns add $3.5160 to a $19.5600 GPT-5.6 Sol conversation — about 18% for 10% of the turns.
Long shapes reward trimming hardestKeeping two prior turns cuts the document shape's input by 82.4%; a summary cuts it by 94.0%. The longer the history, the more there is to stop paying for.
Agent loops are the quiet budget-killerThey look like chats but run unattended: 40 turns of tool output re-sent on every step, 100 times a day, is $18,864.00 a month on GPT-5.6 Sol before any of it is trimmed.
Where it goes

What you are actually paying for

The API has no memory. Anything you want the model to "remember", you re-send — and re-pay for — on every turn.

Re-reading is not freeAttention over history costs compute, which is why providers bill input at all. Re-sending history buys coherence, not insight.
The system prompt multipliesIt is part of every request. A 2,000-token system prompt re-sent over 20 turns costs 40,000 input tokens — nearly half of a 100K token budget, before a word of history is counted.
Output stays expensiveOutput is priced 5x–8.3x input across the eight tracked models, so the 6.3% of tokens that are replies take roughly a quarter of the bill.
Window size is not cost controlBigger windows let you send more; they never charge you less. Check the fit with the Context Window Checker, then price it here.
Strategies

Three ways to carry history forward

All three are priced live in the calculator above. They trade memory against money in very different proportions.

Baseline

Full history

Send every turn, every time. Perfect recall, quadratic cost. Correct answer when the conversation is short, audit-critical, or the earlier turns really are load-bearing.

Cost: 118,000 input tokens on the default shape.

53.4% cheaper input

Sliding window

Keep the last N turns and drop the rest. Cheap to implement and predictable, but it forgets silently — anything older than the window simply disappears from the model's world.

Cost: 55,000 input tokens at five prior turns.

83.1% cheaper input

Rolling summary

Compress older turns into a fixed-size summary and send that instead. Keeps the substance while cutting the tokens, at the price of one extra summarisation call and a loss of exact wording.

Cost: 20,000 input tokens at an 800-token summary.

FAQ

Conversation cost questions

Why does a 20-turn chat cost so much more than one request?

Because turn N re-sends turns 1 to N−1. With 200-token user messages and 400-token replies, the first request sends 200 input tokens and the twentieth sends 11,600 — 58 times more. Across the whole conversation that is 118,000 input tokens to produce 8,000 output tokens, 96.6% of which is re-sent history. On Claude Sonnet 5 that is $0.3160 against $0.0044 for a single turn: 72 times the price for 210 times the tokens, because output is priced five times higher than input.

How much does trimming the history actually save?

On the default shape, keeping only the last five turns cuts input tokens from 118,000 to 55,000 — 53.4% less. Replacing everything older than the current turn with an 800-token rolling summary cuts input to 20,000 tokens, 83.1% less. At 500 conversations a day on Claude Sonnet 5 that is $4,740 a month down to $2,850 or $1,800 — up to $35,280 a year.

Does a bigger context window make this cheaper?

No. A larger window removes the ceiling, not the bill. Claude Opus 5 and Sonnet 5 both accept 1,000,000 tokens, so nothing forces you to drop earlier turns — you simply keep paying the published input rate on every re-sent token. See What Is a Context Window? for how that budget is shared with the reply.

When does OpenAI's long-context surcharge apply to a conversation?

Above 272,000 input tokens in a single request on GPT-5.6 models, where input doubles and output rises 1.5x. A support-sized chat never gets close — even after 40 turns a request is under 24,000 tokens. Document-sized conversations do: in the long-document preset, turn 28 crosses the threshold, and turns 28–30 add $3.5160 to a $19.5600 GPT-5.6 Sol conversation, roughly 18% more for three turns.

Does dropping old turns break the assistant's answers?

It can, which is why the two strategies differ in risk. A sliding window silently loses anything older than the window, so a question about the start of the conversation is answered without that context. A rolling summary keeps the substance while discarding verbatim tokens, so answers degrade more gracefully — at the cost of a summarisation call on your side. Measure the failure rate on real transcripts before shipping either; this page prices the saving, not the regression.

Does this page upload my conversation parameters?

No. The six numbers you enter are processed in your browser and never leave the tab. This page has no form, no backend endpoint, no analytics script, no local storage and no network requests beyond loading its own HTML, CSS and JavaScript. See Privacy.

Methodology

How these numbers are produced

Turn sizes are yours: input tokens for turn t are modelled as (t − 1) × (user + assistant) + user, which is what an application re-sends when it keeps full history in the request. Output is assistant tokens per turn, once per turn. Token estimates do not come from a provider tokenizer — if you want a measured count rather than a planning shape, paste real text into the AI Token Counter.

Costs multiply those token volumes by published per-1M rates taken from official provider pricing pages and re-verified weekly. OpenAI's long-context tier is applied per request, only to requests above 272,000 input tokens; no equivalent tier is published by Anthropic or Google, so none is applied. Caching, batch discounts, free tiers and the cost of generating the summary itself are excluded throughout.

Official sources: OpenAI model docs, Anthropic pricing, Google Gemini API pricing. Always confirm against your invoice before making purchasing decisions.

The full cost-reduction playbookSeven quantified levers in one place: How to Reduce LLM API Costs.
Price any request shapeFixed input and output volumes across all eight models: API Cost Calculator.
Check the window firstConfirm long conversations still fit alongside the reply: Context Window Checker.
Trim before you re-sendHistory is not the only thing billed twice — repeated instructions and markup are too: Prompt Weight Analyzer.
Stop paying for markupRetrieved HTML peanuts the token budget: HTML → Markdown Token Savings.
Convert tokens to meaningSanity-check whether these volumes are even large: Tokens ↔ Words Converter.
Compare headline ratesAll eight tracked models side by side: LLM API Pricing Comparison.
Why output costs moreInput versus output pricing, and what a token actually is: What Is an AI Token?
How we verifySources, weekly cadence and what our estimates exclude: Methodology · Pricing changelog.