Claude vs Gemini API Pricing: Same Workload, Real Numbers
Short answer: Gemini is cheaper at every tier we track — 3.6 Flash's promotional rate beats Sonnet 5 by ~62% on a typical document workload, and Flash-Lite is the cheapest model on either roster. Claude's advantages are structural: 128K max output (double Gemini's), a 1M context mid-tier, and a budget model that becomes the cheaper answer when the Gemini promo ends on January 1, 2027.
The verdict in four lines
Head-to-head pricing table
Standard first-party text rates per 1,000,000 tokens. Claude rates verified 2026-08-09 to 2026-08-16; Gemini rates verified 2026-08-16.
| Tier | Claude | Gemini | Input verdict | Output verdict |
|---|---|---|---|---|
Flagship Hardest reasoning | Claude Opus 5 — $5.00 / $25.00 | — no flagship tracked | — | — |
Mid Balanced default | Claude Sonnet 5 — $2.00 / $10.00 | Gemini 3.6 Flash — $0.75 / $3.75Promo | Gemini −62.5% | Gemini −62.5% |
Budget High-volume tasks | Claude Haiku 4.5 — $1.00 / $5.00 | Gemini 3.5 Flash-Lite — $0.30 / $2.50Lowest | Gemini −70% | Gemini −50% |
Rates in USD per 1M input/output tokens. Gemini 3.6 Flash's $0.75/$3.75 is promotional through Dec 31, 2026 and becomes $1.50/$7.50 from Jan 1, 2027. Sonnet 5's planned Sep 1, 2026 increase was cancelled. Caching, batch, tools and taxes excluded.
Identical document workload, priced monthly
Fixed profile: 60K input + 1K output tokens per request, 1,000 requests per day — 1.8B input and 30M output tokens per 30-day month. No list-price comparisons; this is the same work billed two ways.
| Model | Input cost / mo | Output cost / mo | 30-day total |
|---|---|---|---|
Claude Opus 5 Anthropic · flagship | $9,000.00 | $750.00 | $9,750.00 |
Claude Sonnet 5 Anthropic · mid | $3,600.00 | $300.00 | $3,900.00 |
Claude Haiku 4.5 Anthropic · budget · 200K context | $1,800.00 | $150.00 | $1,950.00 |
Gemini 3.6 Flash Google · mid (promo rate) | $1,350.00 | $112.50 | $1,462.50 |
Gemini 3.5 Flash-LiteLowest Google · budget | $540.00 | $75.00 | $615.00 |
Computed from published per-1M rates on the fixed profile above using Gemini 3.6 Flash's promotional rate; from Jan 1, 2027 its row becomes $2,925.00 ($2,700.00 input + $225.00 output). Your tokenization and caching behavior will differ — run your own volumes below.
Your workload on Claude vs Gemini
Pick a preset or enter per-request token volumes and daily requests. All five models from both providers are priced side by side at current (promotional) rates.
| Model | Per request | Per day | 30-day estimate |
|---|
The break-even is a date, not a volume.
Gemini 3.6 Flash is cheaper than Sonnet 5 on both axes, so there is no usage level where Sonnet 5 wins on arithmetic — the gap only scales: $1.25 per 1M input plus $6.25 per 1M output during the promo, worth $2,437.50/month on the workload above. The decision between them is a quality and capability call, not a volume break-even.
The real break-even is January 1, 2027, when 3.6 Flash doubles to $1.50/$7.50. Two things change that day: the Sonnet 5 gap shrinks from $2,437.50 to $975/month on the same workload, and Claude Haiku 4.5 ($1.00/$5.00 → $1,950/month here) becomes cheaper than Google's mid-tier Flash ($2,925/month). If you are signing an annual budget in 2026, price the second half at post-promo rates.
The one structural Claude win that no promo erases: max output. Opus 5 and Sonnet 5 generate up to 128,000 tokens per request; both Gemini models cap at 65,536. Long reports and full-file code generation fit Claude's single-request budget.
Both take 1M tokens — with one 200K exception.
Neither provider surcharges long prompts in current published rates: Sonnet 5 and Opus 5 accept 1,000,000 input tokens at flat prices, and both Gemini models accept 1,048,576. The exception is Haiku 4.5 at 200K tokens — long documents cannot route to the cheapest Claude, while even the cheapest Gemini takes the full 1M.
Tokenizers differ between Anthropic and Google, so identical text can meter several percent apart — measure real documents with the Token Counter, and confirm window fit with the Context Window Checker. Caching and batch discounts exist on both platforms but are not modeled on this page.
Claude vs Gemini pricing questions
Is Claude or Gemini cheaper?
Gemini is cheaper at every tier tracked here. Gemini 3.6 Flash costs $0.75/$3.75 per 1M input/output tokens on promotion (through Dec 31, 2026) versus Claude Sonnet 5 at $2.00/$10.00 — about 62% cheaper on a mixed document workload. Gemini 3.5 Flash-Lite ($0.30/$2.50) is the cheapest model on this page, undercutting Claude Haiku 4.5 ($1.00/$5.00) by roughly 68% on the same workload.
What happens to Gemini pricing on January 1, 2027?
Gemini 3.6 Flash's promotional rate ends December 31, 2026 and doubles to $1.50/$7.50 per 1M input/output tokens. It stays cheaper than Claude Sonnet 5 ($2.00/$10.00), but Claude Haiku 4.5 ($1.00/$5.00) then undercuts 3.6 Flash — so the cheapest-answer between the two providers flips on that date for budget workloads.
Claude vs Gemini: which is better for long outputs?
Claude. Claude Opus 5 and Sonnet 5 support up to 128,000 output tokens per request, while Gemini 3.6 Flash and 3.5 Flash-Lite cap at 65,536. Long reports, full-file code generation and book-chapter drafting either fit Claude's single-request budget or require chunked generation on Gemini.
Is there a long-context surcharge on Claude or Gemini?
Neither provider applies a long-context surcharge in their current published rates. Claude Sonnet 5 and Opus 5 accept 1,000,000 input tokens at flat rates; Gemini models accept 1,048,576. The exception is Claude Haiku 4.5, whose context window is 200,000 tokens — long documents must route to Sonnet 5, Opus 5, or a Gemini model.
Which is cheapest for document processing: Claude or Gemini?
On cost alone, Gemini 3.6 Flash. On a workload of 1.8B input and 30M output tokens per month it costs about $1,462.50 on promotion versus $3,900 on Claude Sonnet 5 and $1,950 on Claude Haiku 4.5. If your outputs exceed 65,536 tokens or you need Claude-specific behavior, Sonnet 5 is the practical Claude default — benchmark quality on your own documents before routing.
Where these numbers come from.
All rates are taken from the providers' official pricing documentation and re-verified weekly. Workload costs multiply fixed token volumes by published per-million-token rates; they do not estimate tokenization and do not apply caching or batch discounts. "Best fit" on this page is a cost result, not a quality ranking — benchmark both providers on your own prompts.
Official sources: Anthropic pricing and Gemini API pricing. Always confirm the invoice before making purchasing decisions.