Provider comparison · Last verified 2026-08-23

OpenAI vs Gemini API Pricing: Same Workload, Real Numbers

Short answer: Google wins the mid and flagship tiers and keeps winning after both promos expire — Gemini 3.6 Flash's $0.75/$3.75 beats GPT-5.6 Terra's $2.00/$12.00 by 62.5% on input. OpenAI wins the budget tier — GPT-5.6 Luna ($0.20/$1.20) undercuts Flash-Lite ($0.30/$2.50) — but that verdict flips once a prompt crosses 272K input tokens, where OpenAI's long-context surcharge kicks in and Gemini's flat rates do not.

TL;DR

The verdict in four lines

Mid tier: Gemini wins by 62.5% / 68.75%Gemini 3.6 Flash $0.75/$3.75 vs GPT-5.6 Terra $2.00/$12.00 per 1M tokens, through Dec 31, 2026. After the promo ends it is $1.50/$7.50 — still cheaper.
Budget tier: OpenAI wins below 272KGPT-5.6 Luna $0.20/$1.20 vs Gemini 3.5 Flash-Lite $0.30/$2.50 — 33% cheaper on input, 52% on output. Neither rate is promotional.
Flagship: no Gemini counterpart trackedGPT-5.6 Sol ($4.00/$20.00 promo) has no comparable Google flagship in the lineups tracked here. Against 3.6 Flash, Sol costs 5.3x more on input.
Long context: the verdict flips at 272KAbove 272K input tokens OpenAI doubles input rates and Gemini does not — a 400K-input request costs ~$1.62 on Terra vs $0.3038 on 3.6 Flash, and Flash-Lite overtakes Luna.
Rate card

Head-to-head pricing table

Standard first-party text rates per 1,000,000 tokens. OpenAI rates verified 2026-08-23; Gemini rates verified 2026-08-16. The tier structures do not line up neatly — OpenAI fields three tiers, Google two, with no tracked flagship.

TierOpenAIGeminiInput verdictOutput verdict
Flagship
Hardest reasoning
GPT-5.6 Sol — $4.00 / $20.00Promo— no flagship tracked
Mid
Balanced default
GPT-5.6 Terra — $2.00 / $12.00Gemini 3.6 Flash — $0.75 / $3.75PromoGemini −62.5%Gemini −68.75%
Budget
High-volume tasks
GPT-5.6 Luna — $0.20 / $1.20LowestGemini 3.5 Flash-Lite — $0.30 / $2.50OpenAI −33.3%OpenAI −52%

Rates in USD per 1M input/output tokens. GPT-5.6 Sol's $4.00/$20.00 is promotional through at least Nov 21, 2026 (previously $5.00/$30.00). Gemini 3.6 Flash's $0.75/$3.75 is promotional through Dec 31, 2026 and becomes $1.50/$7.50 from Jan 1, 2027. Prompts over 272K input tokens on GPT-5.6 models use higher long-context rates, not shown here. Caching, batch, tools and taxes excluded.

Same workload

Identical document workload, priced monthly

Fixed profile: 60K input + 1K output tokens per request, 1,000 requests per day — 1.8B input and 30M output tokens per 30-day month. This is input-heavy work, which is where OpenAI's long-context exposure is highest.

ModelInput cost / moOutput cost / mo30-day total
GPT-5.6 Sol
OpenAI · flagship (promo)
$7,200.00$600.00$7,800.00
GPT-5.6 Terra
OpenAI · mid
$3,600.00$360.00$3,960.00
GPT-5.6 LunaLowest
OpenAI · budget
$360.00$36.00$396.00
Gemini 3.6 Flash
Google · mid (promo rate)
$1,350.00$112.50$1,462.50
Gemini 3.5 Flash-Lite
Google · budget
$540.00$75.00$615.00

Computed from published per-1M rates on the fixed profile above. From Jan 1, 2027 the 3.6 Flash row becomes $2,925.00 ($2,700.00 input + $225.00 output) — still below Terra's $3,960.00. Your tokenization and caching behavior will differ — run your own volumes below.

Output-heavy workload

Flip the ratio: chatbot profile, priced monthly

Second fixed profile: 2K input + 500 output tokens per request, 5,000 requests per day — 300M input and 75M output tokens per 30-day month. Output is billed 5–8x higher than input, so this profile tells a different story than the document one.

ModelInput cost / moOutput cost / mo30-day totalShare of bill from output
GPT-5.6 Sol
OpenAI · flagship (promo)
$1,200.00$1,500.00$2,700.0055.6%
GPT-5.6 Terra
OpenAI · mid
$600.00$900.00$1,500.0060.0%
GPT-5.6 LunaLowest
OpenAI · budget
$60.00$90.00$150.0060.0%
Gemini 3.6 Flash
Google · mid (promo rate)
$225.00$281.25$506.2555.6%
Gemini 3.5 Flash-Lite
Google · budget
$90.00$187.50$277.5067.6%

Same computation method as the table above, on the second fixed profile. Note that GPT-5.6 Luna wins both profiles on cost — but only while prompts stay under 272K input tokens.

Live calculator

Your workload on OpenAI vs Gemini

Pick a preset or enter per-request token volumes and daily requests. All five models from both providers are priced side by side; OpenAI long-context rates apply automatically above 272K input tokens, so you can watch the crossover happen.

Presets:
Cheapest 30-day total
Most expensive 30-day total
ModelPer requestPer day30-day estimate
Current rates: first-party OpenAI and Google API pricing, both at promotional rates where applicable. Excluded: caching, batch, tools and taxes. Pricing data verified —
Break-even

Two promos expire — the direction does not change.

The calendar. GPT-5.6 Sol's promotional $4.00/$20.00 runs through at least November 21, 2026; Gemini 3.6 Flash's promotional $0.75/$3.75 runs through December 31, 2026 and then doubles to $1.50/$7.50. Neither date flips the mid-tier verdict: at post-promo rates 3.6 Flash still costs less than GPT-5.6 Terra on both axes ($1.50 vs $2.00 input, $7.50 vs $12.00 output). What changes is the size of the saving — on the reference workload the Flash-vs-Terra gap shrinks from $2,497.50/month to $1,035/month.

The token count. The sharp break-even on this page is 272K input tokens. Below it, OpenAI's budget model is the cheapest option on either roster. Above it, OpenAI's long-context rates apply — input ×2, output ×1.5 — while Gemini stays flat to 1,048,576 input tokens. A 400K-input, 1K-output request costs roughly $1.62 on GPT-5.6 Terra versus $0.3038 on Gemini 3.6 Flash.

The ratio. Above the 272K threshold, Luna's long-context rates ($0.40 in / $1.80 out) are cheaper than Flash-Lite on output but dearer on input ($0.30 / $2.50). That makes the crossover a ratio rather than a threshold: Luna stays cheaper only when output tokens exceed input tokens ÷ 7. At 400K input that means more than 57,143 output tokens; at the 1K output used above, Flash-Lite wins at $0.1225 versus Luna's $0.1618.

Mid-tier gap (promo)3.6 Flash saves $1.25 per 1M input + $8.25 per 1M output vs Terra — $2,497.50/mo on the reference workload.
Mid-tier gap (post-promo)Shrinks to $1,035/mo from Jan 1, 2027. Still cheaper — the gap changes, the winner does not.
272K flipAbove 272K input tokens, check both providers with the Context Window Checker before assuming Luna is cheapest.
Output ceilingGPT-5.6 models generate up to 128,000 output tokens per request; both Gemini models cap at 65,536 — half.
Long context

One request, 400K input tokens, priced per call

Per-request cost for a single 400,000-token input prompt with a 1,000-token response — above OpenAI's 272K long-context threshold, inside Gemini's 1,048,576-token limit.

ModelEffective input rate / 1MInput costOutput costPer request
GPT-5.6 Sol
OpenAI · long-context rate
$8.00$3.2000$0.0300≈ $3.23
GPT-5.6 Terra
OpenAI · long-context rate
$4.00$1.6000$0.0180≈ $1.62
GPT-5.6 Luna
OpenAI · long-context rate
$0.40$0.1600$0.0018$0.1618
Gemini 3.6 Flash
Google · flat rate (promo)
$0.75$0.3000$0.0038$0.3038
Gemini 3.5 Flash-LiteLowest
Google · flat rate
$0.30$0.1200$0.0025$0.1225

Long-context rates on GPT-5.6 models: input price doubles and output price rises 1.5x once a prompt exceeds 272,000 input tokens. Gemini applies no long-context surcharge at these rates. Computed on 400,000 input and 1,000 output tokens; post-promo 3.6 Flash would cost $0.6075 on the same request.

Context fit & caveats

Both take 1M tokens — the ceiling is on the way out.

On input the two providers are effectively tied: 1,050,000 tokens on GPT-5.6 versus 1,048,576 on Gemini — a difference of 1,424 tokens, or 0.14%. The real structural difference is on the output side, where GPT-5.6 models generate up to 128,000 tokens per request and both Gemini models cap at 65,536. Long reports, full-file code generation and book-length drafting either fit a single OpenAI request or need chunking on Gemini.

Tokenizers differ between OpenAI and Google, so identical text can meter several percent apart — measure real documents with the Token Counter before committing, and confirm fit with the Context Window Checker. Prompt caching exists on both platforms (cached input on GPT-5.6 Sol is $0.40 per 1M) and is not modeled on this page. Neither is batch pricing, tool use, or tax.

Deep dive: OpenAIGPT-5.6 rates, the 272K surcharge and the Sol promo window: OpenAI API Pricing.
Deep dive: Gemini3.6 Flash promo window and Flash-Lite budget rates: Gemini API Pricing.
Budget the winnerModel monthly and annual spend with the API Cost Calculator.
FAQ

OpenAI vs Gemini pricing questions

Is OpenAI or Gemini cheaper?

Google wins the mid and flagship tiers, OpenAI wins the budget tier below 272K input tokens. Gemini 3.6 Flash costs $0.75/$3.75 per 1M input/output versus GPT-5.6 Terra at $2.00/$12.00 — 62.5% cheaper on input, 68.75% on output. But GPT-5.6 Luna ($0.20/$1.20) undercuts Gemini 3.5 Flash-Lite ($0.30/$2.50) on both axes, so OpenAI is cheaper for short, high-volume requests.

What happens when both promotional prices expire?

The gap shrinks but the direction does not change. GPT-5.6 Sol's promotional $4.00/$20.00 runs through at least Nov 21, 2026; Gemini 3.6 Flash's promotional $0.75/$3.75 runs through Dec 31, 2026 and becomes $1.50/$7.50. Even at post-promo rates, 3.6 Flash stays cheaper than both GPT-5.6 Terra ($2.00/$12.00) and post-promo Sol ($5.00/$30.00) — on the 1.8B-in/30M-out reference workload it costs $2,925/month versus $3,960 on Terra.

Which is cheaper for long documents: OpenAI or Gemini?

Gemini, and the gap widens with length. OpenAI applies long-context rates above 272K input tokens (input doubles, output rises 1.5x) while Gemini holds flat rates to 1,048,576 input tokens. A 400K-input / 1K-output request costs about $1.62 on GPT-5.6 Terra versus $0.3038 on Gemini 3.6 Flash — roughly 5.3x more on OpenAI. Gemini also becomes cheaper than GPT-5.6 Luna at that length, because Luna's long-context input rate is $0.40 per 1M versus Flash-Lite's flat $0.30.

Why does GPT-5.6 Luna stop being the cheapest model above 272K tokens?

Because OpenAI's long-context surcharge doubles Luna's input rate from $0.20 to $0.40 per 1M tokens and raises its output rate from $1.20 to $1.80. That makes Luna more expensive than Gemini 3.5 Flash-Lite ($0.30/$2.50) on input but still cheaper on output, so the break-even becomes a ratio: Luna only stays cheaper when output tokens exceed input tokens divided by 7. On a 400K-input / 1K-output request Luna costs about $0.1618 versus Flash-Lite's $0.1225 — Gemini wins.

Do OpenAI and Gemini count tokens the same way?

No. OpenAI and Google use different tokenizers, so the same text can meter several percent differently, and both will differ from any third-party estimate. Before migrating a workload, measure your real prompts with a token counter and validate against each provider's own tokenizer or usage API.

Methodology

Where these numbers come from.

All rates are taken from the providers' official pricing documentation and re-verified weekly. Workload costs multiply fixed token volumes by published per-million-token rates; they do not estimate tokenization and do not apply caching or batch discounts. "Best fit" on this page is a cost result, not a quality ranking — benchmark both providers on your own prompts.

Official sources: OpenAI model docs and Gemini API pricing. Always confirm the invoice before making purchasing decisions.

All three providersAdd Anthropic to the picture on the LLM API Pricing Comparison page.
OpenAI vs ClaudeHow OpenAI compares to Anthropic's lineup: OpenAI vs Claude Pricing.
Claude vs GeminiThe other leg of the triangle, including the Jan 1, 2027 flip: Claude vs Gemini Pricing.
What 100K tokens is~75,000 words, 150 pages, and the chunk size that stays under OpenAI's 272K surcharge: How Many Words Is 100K Tokens?.
Context windows explainedInput and output share one budget — window sizes, output caps and what it costs to fill one: What Is a Context Window?.
What a token isDefinition, conversion by content type, and why output costs 5–8x input: What Is an AI Token?
HTML token savingsCleaning scraped markup removes 47–68% of input tokens: HTML Token Savings.
1M tokens in words750,000 words, 1,500 pages — and why only 5 of 8 models accept it in one request: How Many Words Is 1M Tokens?
Trim the promptFind repeated instructions, markup and filler: Prompt Weight Analyzer
How we verifySources, weekly cadence and what our estimates exclude: Methodology · Pricing changelog.