OpenAI vs Gemini API Pricing: Same Workload, Real Numbers
Short answer: Google wins the mid and flagship tiers and keeps winning after both promos expire — Gemini 3.6 Flash's $0.75/$3.75 beats GPT-5.6 Terra's $2.00/$12.00 by 62.5% on input. OpenAI wins the budget tier — GPT-5.6 Luna ($0.20/$1.20) undercuts Flash-Lite ($0.30/$2.50) — but that verdict flips once a prompt crosses 272K input tokens, where OpenAI's long-context surcharge kicks in and Gemini's flat rates do not.
The verdict in four lines
Head-to-head pricing table
Standard first-party text rates per 1,000,000 tokens. OpenAI rates verified 2026-08-23; Gemini rates verified 2026-08-16. The tier structures do not line up neatly — OpenAI fields three tiers, Google two, with no tracked flagship.
| Tier | OpenAI | Gemini | Input verdict | Output verdict |
|---|---|---|---|---|
Flagship Hardest reasoning | GPT-5.6 Sol — $4.00 / $20.00Promo | — no flagship tracked | — | — |
Mid Balanced default | GPT-5.6 Terra — $2.00 / $12.00 | Gemini 3.6 Flash — $0.75 / $3.75Promo | Gemini −62.5% | Gemini −68.75% |
Budget High-volume tasks | GPT-5.6 Luna — $0.20 / $1.20Lowest | Gemini 3.5 Flash-Lite — $0.30 / $2.50 | OpenAI −33.3% | OpenAI −52% |
Rates in USD per 1M input/output tokens. GPT-5.6 Sol's $4.00/$20.00 is promotional through at least Nov 21, 2026 (previously $5.00/$30.00). Gemini 3.6 Flash's $0.75/$3.75 is promotional through Dec 31, 2026 and becomes $1.50/$7.50 from Jan 1, 2027. Prompts over 272K input tokens on GPT-5.6 models use higher long-context rates, not shown here. Caching, batch, tools and taxes excluded.
Identical document workload, priced monthly
Fixed profile: 60K input + 1K output tokens per request, 1,000 requests per day — 1.8B input and 30M output tokens per 30-day month. This is input-heavy work, which is where OpenAI's long-context exposure is highest.
| Model | Input cost / mo | Output cost / mo | 30-day total |
|---|---|---|---|
GPT-5.6 Sol OpenAI · flagship (promo) | $7,200.00 | $600.00 | $7,800.00 |
GPT-5.6 Terra OpenAI · mid | $3,600.00 | $360.00 | $3,960.00 |
GPT-5.6 LunaLowest OpenAI · budget | $360.00 | $36.00 | $396.00 |
Gemini 3.6 Flash Google · mid (promo rate) | $1,350.00 | $112.50 | $1,462.50 |
Gemini 3.5 Flash-Lite Google · budget | $540.00 | $75.00 | $615.00 |
Computed from published per-1M rates on the fixed profile above. From Jan 1, 2027 the 3.6 Flash row becomes $2,925.00 ($2,700.00 input + $225.00 output) — still below Terra's $3,960.00. Your tokenization and caching behavior will differ — run your own volumes below.
Flip the ratio: chatbot profile, priced monthly
Second fixed profile: 2K input + 500 output tokens per request, 5,000 requests per day — 300M input and 75M output tokens per 30-day month. Output is billed 5–8x higher than input, so this profile tells a different story than the document one.
| Model | Input cost / mo | Output cost / mo | 30-day total | Share of bill from output |
|---|---|---|---|---|
GPT-5.6 Sol OpenAI · flagship (promo) | $1,200.00 | $1,500.00 | $2,700.00 | 55.6% |
GPT-5.6 Terra OpenAI · mid | $600.00 | $900.00 | $1,500.00 | 60.0% |
GPT-5.6 LunaLowest OpenAI · budget | $60.00 | $90.00 | $150.00 | 60.0% |
Gemini 3.6 Flash Google · mid (promo rate) | $225.00 | $281.25 | $506.25 | 55.6% |
Gemini 3.5 Flash-Lite Google · budget | $90.00 | $187.50 | $277.50 | 67.6% |
Same computation method as the table above, on the second fixed profile. Note that GPT-5.6 Luna wins both profiles on cost — but only while prompts stay under 272K input tokens.
Your workload on OpenAI vs Gemini
Pick a preset or enter per-request token volumes and daily requests. All five models from both providers are priced side by side; OpenAI long-context rates apply automatically above 272K input tokens, so you can watch the crossover happen.
| Model | Per request | Per day | 30-day estimate |
|---|
Two promos expire — the direction does not change.
The calendar. GPT-5.6 Sol's promotional $4.00/$20.00 runs through at least November 21, 2026; Gemini 3.6 Flash's promotional $0.75/$3.75 runs through December 31, 2026 and then doubles to $1.50/$7.50. Neither date flips the mid-tier verdict: at post-promo rates 3.6 Flash still costs less than GPT-5.6 Terra on both axes ($1.50 vs $2.00 input, $7.50 vs $12.00 output). What changes is the size of the saving — on the reference workload the Flash-vs-Terra gap shrinks from $2,497.50/month to $1,035/month.
The token count. The sharp break-even on this page is 272K input tokens. Below it, OpenAI's budget model is the cheapest option on either roster. Above it, OpenAI's long-context rates apply — input ×2, output ×1.5 — while Gemini stays flat to 1,048,576 input tokens. A 400K-input, 1K-output request costs roughly $1.62 on GPT-5.6 Terra versus $0.3038 on Gemini 3.6 Flash.
The ratio. Above the 272K threshold, Luna's long-context rates ($0.40 in / $1.80 out) are cheaper than Flash-Lite on output but dearer on input ($0.30 / $2.50). That makes the crossover a ratio rather than a threshold: Luna stays cheaper only when output tokens exceed input tokens ÷ 7. At 400K input that means more than 57,143 output tokens; at the 1K output used above, Flash-Lite wins at $0.1225 versus Luna's $0.1618.
One request, 400K input tokens, priced per call
Per-request cost for a single 400,000-token input prompt with a 1,000-token response — above OpenAI's 272K long-context threshold, inside Gemini's 1,048,576-token limit.
| Model | Effective input rate / 1M | Input cost | Output cost | Per request |
|---|---|---|---|---|
GPT-5.6 Sol OpenAI · long-context rate | $8.00 | $3.2000 | $0.0300 | ≈ $3.23 |
GPT-5.6 Terra OpenAI · long-context rate | $4.00 | $1.6000 | $0.0180 | ≈ $1.62 |
GPT-5.6 Luna OpenAI · long-context rate | $0.40 | $0.1600 | $0.0018 | $0.1618 |
Gemini 3.6 Flash Google · flat rate (promo) | $0.75 | $0.3000 | $0.0038 | $0.3038 |
Gemini 3.5 Flash-LiteLowest Google · flat rate | $0.30 | $0.1200 | $0.0025 | $0.1225 |
Long-context rates on GPT-5.6 models: input price doubles and output price rises 1.5x once a prompt exceeds 272,000 input tokens. Gemini applies no long-context surcharge at these rates. Computed on 400,000 input and 1,000 output tokens; post-promo 3.6 Flash would cost $0.6075 on the same request.
Both take 1M tokens — the ceiling is on the way out.
On input the two providers are effectively tied: 1,050,000 tokens on GPT-5.6 versus 1,048,576 on Gemini — a difference of 1,424 tokens, or 0.14%. The real structural difference is on the output side, where GPT-5.6 models generate up to 128,000 tokens per request and both Gemini models cap at 65,536. Long reports, full-file code generation and book-length drafting either fit a single OpenAI request or need chunking on Gemini.
Tokenizers differ between OpenAI and Google, so identical text can meter several percent apart — measure real documents with the Token Counter before committing, and confirm fit with the Context Window Checker. Prompt caching exists on both platforms (cached input on GPT-5.6 Sol is $0.40 per 1M) and is not modeled on this page. Neither is batch pricing, tool use, or tax.
OpenAI vs Gemini pricing questions
Is OpenAI or Gemini cheaper?
Google wins the mid and flagship tiers, OpenAI wins the budget tier below 272K input tokens. Gemini 3.6 Flash costs $0.75/$3.75 per 1M input/output versus GPT-5.6 Terra at $2.00/$12.00 — 62.5% cheaper on input, 68.75% on output. But GPT-5.6 Luna ($0.20/$1.20) undercuts Gemini 3.5 Flash-Lite ($0.30/$2.50) on both axes, so OpenAI is cheaper for short, high-volume requests.
What happens when both promotional prices expire?
The gap shrinks but the direction does not change. GPT-5.6 Sol's promotional $4.00/$20.00 runs through at least Nov 21, 2026; Gemini 3.6 Flash's promotional $0.75/$3.75 runs through Dec 31, 2026 and becomes $1.50/$7.50. Even at post-promo rates, 3.6 Flash stays cheaper than both GPT-5.6 Terra ($2.00/$12.00) and post-promo Sol ($5.00/$30.00) — on the 1.8B-in/30M-out reference workload it costs $2,925/month versus $3,960 on Terra.
Which is cheaper for long documents: OpenAI or Gemini?
Gemini, and the gap widens with length. OpenAI applies long-context rates above 272K input tokens (input doubles, output rises 1.5x) while Gemini holds flat rates to 1,048,576 input tokens. A 400K-input / 1K-output request costs about $1.62 on GPT-5.6 Terra versus $0.3038 on Gemini 3.6 Flash — roughly 5.3x more on OpenAI. Gemini also becomes cheaper than GPT-5.6 Luna at that length, because Luna's long-context input rate is $0.40 per 1M versus Flash-Lite's flat $0.30.
Why does GPT-5.6 Luna stop being the cheapest model above 272K tokens?
Because OpenAI's long-context surcharge doubles Luna's input rate from $0.20 to $0.40 per 1M tokens and raises its output rate from $1.20 to $1.80. That makes Luna more expensive than Gemini 3.5 Flash-Lite ($0.30/$2.50) on input but still cheaper on output, so the break-even becomes a ratio: Luna only stays cheaper when output tokens exceed input tokens divided by 7. On a 400K-input / 1K-output request Luna costs about $0.1618 versus Flash-Lite's $0.1225 — Gemini wins.
Do OpenAI and Gemini count tokens the same way?
No. OpenAI and Google use different tokenizers, so the same text can meter several percent differently, and both will differ from any third-party estimate. Before migrating a workload, measure your real prompts with a token counter and validate against each provider's own tokenizer or usage API.
Where these numbers come from.
All rates are taken from the providers' official pricing documentation and re-verified weekly. Workload costs multiply fixed token volumes by published per-million-token rates; they do not estimate tokenization and do not apply caching or batch discounts. "Best fit" on this page is a cost result, not a quality ranking — benchmark both providers on your own prompts.
Official sources: OpenAI model docs and Gemini API pricing. Always confirm the invoice before making purchasing decisions.