Can a 200-Page PDF Fit in an LLM?
Short answer: yes, on all eight tracked models — a text-based 200-page document is about 133,333 tokens and the smallest window here is 200,000. But it is closer than it looks: on Claude Haiku 4.5 a 200-page PDF plus a full-length 64,000-token answer leaves just 2,167 tokens, about 1% of the window. Cost per document runs from $0.0292 on GPT-5.6 Luna to $0.7192 on Claude Opus 5 — a 24.7x spread for the same file.
Pages to words to tokens
No provider bills by the page, so every PDF question is really a token question in disguise.
How many tokens is a document of N pages?
Ordinary English prose at 500 words per page and 0.75 words per token. Planning estimates, not tokenizer output.
| Pages | Words | Tokens | Fits the smallest window? | Fits a 1M window? |
|---|---|---|---|---|
| 50 | 25,000 | 33,333 | Yes | Yes |
| 100 | 50,000 | 66,667 | Yes | Yes |
| 200 | 100,000 | 133,333 | Yes | Yes |
| 300 | 150,000 | 200,000 | Exactly at the limit | Yes |
| 500 | 250,000 | 333,333 | No | Yes |
| 800 | 400,000 | 533,333 | No | Yes |
| 1,000 | 500,000 | 666,667 | No | Yes |
| 1,500 | 750,000 | 1,000,000 | No | Exactly at the limit |
Smallest window is Claude Haiku 4.5 at 200,000 tokens. "Exactly at the limit" leaves no room for a system prompt or a reply, so treat it as a fail in practice. Page and word figures are planning estimates at 0.75 English words per token; the same document meters several percent differently across providers, and code, tables and non-English text run higher.
Maximum pages each model can hold
The first figure is the whole window filled with document text. The second reserves room for the longest reply the model can produce.
| Model | Context window | Output cap | Pages (input only) | Pages (with a full-length answer) |
|---|---|---|---|---|
GPT-5.6 Sol OpenAI · flagship (promo) | 1,050,000 | 128,000 | 1,575 | 1,383 |
GPT-5.6 Terra OpenAI · mid | 1,050,000 | 128,000 | 1,575 | 1,383 |
GPT-5.6 Luna OpenAI · budget | 1,050,000 | 128,000 | 1,575 | 1,383 |
Claude Opus 5 Anthropic · flagship | 1,000,000 | 128,000 | 1,500 | 1,308 |
Claude Sonnet 5 Anthropic · mid | 1,000,000 | 128,000 | 1,500 | 1,308 |
Claude Haiku 4.5Tightest Anthropic · budget · smallest window | 200,000 | 64,000 | 300 | 204 |
Gemini 3.6 Flash Google · mid (promo rate) | 1,048,576 | 65,536 | 1,573 | 1,475 |
Gemini 3.5 Flash-Lite Google · budget | 1,048,576 | 65,536 | 1,573 | 1,475 |
Computed as window ÷ 667 tokens per page, and (window − output cap) ÷ 667. Claude Haiku 4.5 is the only tracked model where 200 pages is anywhere near the ceiling: its practical limit with a full-length answer is 204 pages, so a 200-page document uses 98% of it. Everything else has six times the headroom it needs.
On Claude Haiku 4.5, 200 pages is 98% of the ceiling.
Every other model treats a 200-page document as routine. Haiku 4.5 does not, and the reason is arithmetic rather than capability.
The document is 133,333 tokens. Add a 500-token system prompt and question and you are sending 133,833 input tokens into a 200,000-token window — 66.9% of it. That leaves 66,167 tokens, and the output cap takes 64,000 of them if you want a full-length answer. What remains is 2,167 tokens, about 1.1% of the window, for tool output, retries and any answer longer than you planned for. With the 2,000-token answer priced in the cost table below, headroom is 64,167 tokens.
Density is what breaks it. At 600 words per page — a journal article or a legal filing — the same 200 pages is 160,000 tokens. That still fits the window on its own, but with a 64,000-token answer reserved it needs 224,500 tokens and Haiku 4.5 rejects the request. On the 1M-token models the same document needs 288,500 and fits with room to spare.
The practical rule: below 200 pages every model works; between 200 and 300 pages, check before you batch. Paste the extracted text into the Context Window Checker rather than trusting the page count.
Will your document fit, and what will it cost?
Set the page count, the density and the answer length you need. OpenAI's 272K long-context surcharge switches on automatically where it applies.
| Model | Context window | Input needed | Fits | Headroom | Cost per document | Monthly |
|---|
What one 200-page document costs
133,833 input tokens (the document plus a 500-token prompt) and a 2,000-token answer, at published per-1M rates. Monthly assumes 100 documents a day.
| Model | Input cost | Output cost | Per document | Monthly (3,000 documents) |
|---|---|---|---|---|
GPT-5.6 Sol OpenAI · flagship (promo) | $0.5353 | $0.0400 | $0.5753 | $1,726.00 |
GPT-5.6 Terra OpenAI · mid | $0.2677 | $0.0240 | $0.2917 | $875.00 |
GPT-5.6 LunaLowest OpenAI · budget | $0.0268 | $0.00240 | $0.0292 | $87.50 |
Claude Opus 5 Anthropic · flagship | $0.6692 | $0.0500 | $0.7192 | $2,157.50 |
Claude Sonnet 5 Anthropic · mid | $0.2677 | $0.0200 | $0.2877 | $863.00 |
Claude Haiku 4.5 Anthropic · budget · tightest fit | $0.1338 | $0.0100 | $0.1438 | $431.50 |
Gemini 3.6 Flash Google · mid (promo rate) | $0.1004 | $0.00750 | $0.1079 | $323.62 |
Gemini 3.5 Flash-Lite Google · budget | $0.0401 | $0.00500 | $0.0451 | $135.45 |
At 133,833 input tokens every model stays under OpenAI's 272,000-token threshold, so no long-context surcharge applies to a 200-page document. The 24.7x spread — $87.50 to $2,157.50 a month for identical throughput — is pure per-token rate, not window size. Output is only 1.5% of the tokens here and 6.9% to 11.1% of the cost, because output rates run 5–8.3x input: see Input vs Output Tokens.
Fitting is rarely the actual problem.
Four things fail before the window does, and none of them appear on a rate card.
Scanned PDFs have no text layer. A page that is really an image extracts to nothing at all. You need OCR first, and OCR output is noisy — broken words tokenise into more tokens than clean text, so a 200-page scan can cost more than a 300-page original.
Tables and markup explode. A table that reads as 200 words can extract into several hundred tokens of spacing, pipes and repeated headers. Converting PDF content to clean text or Markdown before sending removes 47–68% of input tokens on the samples measured in HTML Token Savings.
Retrieval quality decays with length. Material in the middle of a very long context is recalled less reliably than material at the beginning or end. Feeding the 20 relevant pages usually beats feeding all 200.
The answer is part of the budget. Input and output share one window, and the output cap is a second, lower limit. A request that is close to the limit has no room for tool output or a longer reply than expected — see What Is a Context Window?.
An 800-page manual is a different problem.
Scale the same conversion up and two things change at once.
A window stops fitting. 800 pages is 533,333 tokens — over twice Claude Haiku 4.5's 200,000-token window. It is not close, and no amount of prompt trimming changes that.
OpenAI's surcharge switches on. Above 272,000 input tokens, GPT-5.6 input doubles and output rises 1.5x. An 800-page document costs $4.33 in one request on GPT-5.6 Sol — more than Claude Opus 5's $2.72, despite Sol being the cheaper model per token at this size. Split into sub-272K chunks, the same content costs roughly half on Sol.
On Anthropic and Google the input rate is flat across the whole window, so chunking there costs marginally more because you repeat the system prompt. The reason to chunk is reliability and retrieval quality, not price — read the ranking of all seven levers in How to Reduce LLM API Costs.
PDF and context-window questions
How many tokens is a 200-page PDF?
About 133,333 tokens for a text-based document at 500 words per page and 0.75 English words per token. A dense academic PDF at 600 words per page is closer to 160,000 tokens, and a PDF with heavy tables or markup can run higher still. Extract the text and count it with a token counter rather than trusting the page number.
Which models can read a 200-page PDF in one request?
All eight models tracked here. The smallest window is Claude Haiku 4.5 at 200,000 tokens, which holds 300 pages of ordinary prose — or 204 pages if you reserve room for its full 64,000-token answer. Everything else has a 1,000,000-token window or larger and holds 1,308 to 1,475 pages even with a maximum-length reply.
How many pages fit in a 1M context window?
About 1,500 pages of ordinary prose at 500 words per page. GPT-5.6 models hold 1,575 pages, Claude Opus 5 and Sonnet 5 hold 1,500, and both Gemini models hold 1,573. Reserve room for the response and the figures drop to 1,383, 1,308 and 1,475 respectively.
Why did my PDF fail even though it should fit?
Four usual causes. The output cap: input and output share one window, so reserving a long answer leaves less room for the document. Overhead: system prompt, tool schemas and instructions are billed as input alongside the document. A scanned PDF with no text layer extracts to nothing or to OCR noise. And verbose extraction markup, which inflates tokens well above the clean text count.
Is it cheaper to split a long PDF into chunks?
On OpenAI, yes above 272,000 input tokens, where input doubles — an 800-page manual costs $4.33 in one request on GPT-5.6 Sol but roughly half that in sub-272K chunks. On Anthropic and Google the input rate is flat, so chunking costs slightly more because you repeat the system prompt; the reason to chunk there is reliability and retrieval quality, not price.
Do images and tables in a PDF count as tokens?
Images are converted to tokens under each provider's own rules and billed at input rates; the tables on this page cover extracted text only. Tables are the bigger surprise — a table that reads as 200 words can extract into several hundred tokens of whitespace and delimiters, which is why converting PDF content to clean text or Markdown before sending usually costs less.
Where these numbers come from.
Context windows and output caps are provider-published values, re-verified weekly against official model documentation (OpenAI 2026-08-23; Anthropic and Google 2026-08-09 to 2026-08-16). Page and word conversions are planning estimates at 0.75 English words per token and 500 words per page for ordinary prose — they are not tokenizer output, and the same document meters several percent differently across providers. Cost figures multiply the token estimate by published per-1M rates; OpenAI long-context rates are applied automatically above 272,000 input tokens. Image, audio, tool use, caching, batch and taxes are excluded throughout. Always confirm against your own invoice before making purchasing decisions.
Official sources: OpenAI model docs, Anthropic pricing, Google Gemini API pricing.