Gemini API Pricing 2026: 3.8 Flash, 3.1 Pro and Flash-Lite
Google's Gemini 3.8 Flash price doubles on 1 January 2027. Here's what every Gemini 3 model costs direct from Google, through Requesty and OpenRouter, and on SayGm, where 3.8 Flash costs half Google's standard rate.

On this page
Gemini API pricing for the Gemini 3 text models runs from $0.25 per million input tokens for 3.1 Flash-Lite to $2.00 for 3.1 Pro, and SayGm serves every one of them below Google's standard rate, through the Gemini SDK you already use. The timing matters more than usual this quarter. Google's prices for Gemini 3.8, 3.7 and 3.6 Flash are introductory and double on 1 January 2027, so a budget built on today's rates will be wrong in three months.
Table of Contents
- How much does the Gemini API cost?
- Gemini API pricing for every Gemini 3 text model
- Gemini 3.8 Flash pricing and the January change
- Gemini 3.1 Pro pricing above 200K tokens
- Gemini Flash pricing at the low end: Flash-Lite
- Where to buy Gemini: Google, OpenRouter, Requesty and SayGm
- How private is Gemini on SayGm?
- Questions about Gemini API costs
- What to budget for Gemini
- About SayGm
How much does the Gemini API cost?
SayGm is the cheapest way to buy Gemini at standard speed. Gemini 3.8 Flash costs $0.37 per million input tokens and $1.86 output on SayGm, half Google's direct rate of $0.75 and $3.75, and below Requesty and OpenRouter, which both add a fee on top of Google's price. Across every Gemini 3 text model, SayGm is 25% to 51% below buying direct.
On an 80/20 input-to-output mix, 100 million 3.8 Flash tokens costs $67 on SayGm, $135 direct from Google and $142 through Requesty or OpenRouter. Google's 3.8 Flash rate doubles to $1.50 and $7.50 on 1 January 2027.
Gemini API pricing for every Gemini 3 text model

Here is what each Gemini 3 text model costs direct from Google, on SayGm, through Requesty and through OpenRouter. Google's rates are its standard paid tier, from Google's pricing page. Requesty adds its 5% markup and OpenRouter its 5.5% card fee to Google's rate. All checked 5 October 2026.
Each cell shows input / output per million tokens, then the cost of 100 million tokens at an 80/20 mix.
| Model | Google direct | SayGm | Requesty | OpenRouter |
|---|---|---|---|---|
| Gemini 3.8 Flash | $0.75 / $3.75 · $135 | $0.37 / $1.86 · $67 | $0.79 / $3.94 · $142 | $0.79 / $3.96 · $142 |
| Gemini 3.7 Flash | $0.75 / $3.75 · $135 | $0.56 / $2.81 · $101 | $0.79 / $3.94 · $142 | $0.79 / $3.96 · $142 |
| Gemini 3.5 Flash | $1.50 / $9.00 · $300 | $1.13 / $6.75 · $225 | $1.58 / $9.45 · $315 | $1.58 / $9.49 · $316 |
| Gemini 3.5 Flash-Lite | $0.30 / $2.50 · $74 | $0.23 / $1.88 · $56 | $0.32 / $2.62 · $78 | $0.32 / $2.64 · $78 |
| Gemini 3.1 Flash-Lite | $0.25 / $1.50 · $50 | $0.19 / $1.13 · $38 | $0.26 / $1.58 · $52 | $0.26 / $1.58 · $53 |
| Gemini 3.1 Pro (preview) | $2.00 / $12.00 · $400 | $1.50 / $9.00 · $300 | $2.10 / $12.60 · $420 | $2.11 / $12.66 · $422 |
Gemini 3.6 Flash is priced the same as 3.7 on Google and SayGm. Google's 3.8, 3.7 and 3.6 Flash rates double on 1 January 2027; the others don't change. 3.1 Pro's rates are for prompts up to 200K tokens.
The comparison is at standard speed. Google also sells Flex and Batch at half its standard rate for work that can wait, and on 3.5 Flash, Flash-Lite and 3.1 Pro those tiers come in below SayGm.
Gemini 3.8 Flash pricing and the January change
Gemini 3.8 Flash is Google's newest model, now generally available, and Google calls it its most intelligent Flash model, built for long-running coding work and agents. It has a 1M-token context window and returns up to 64K tokens per response.
Three pricing details matter more than the headline rate.
The rate is introductory. $0.75 and $3.75 applies through 31 December 2026. From 1 January 2027 Google's standard rate becomes $1.50 and $7.50, and the same change applies to 3.7 and 3.6 Flash. Cached input doubles too, from $0.075 to $0.15.
Thinking tokens bill as output. 3.8 Flash reasons before it answers, at low, medium or high effort, and Google charges those reasoning tokens at the output rate. A request that thinks hard can bill several times the visible answer's length. Medium is the default; if a task doesn't need it, low effort is the quickest way to cut a 3.8 Flash bill.
Google sells four speeds. Standard is the rate everyone quotes. Flex and Batch are half that, $0.375 and $1.875, for work that can wait. Priority is $1.35 and $6.75 for traffic that can't. All four double on 1 January.
SayGm's 3.8 Flash rate, $0.37 and $1.86, sits at about Google's Flex price. The Gemini 3.8 Flash model page shows the live rate and measured throughput.
Gemini 3.1 Pro pricing above 200K tokens
Gemini 3.1 Pro is still in preview, and it's the only model in this table without a free tier. Google charges $2.00 input and $12.00 output for prompts up to 200K tokens. Past that, every token in the request costs more: $4.00 input and $18.00 output.
That threshold is easy to cross by accident. A long codebase, a stack of PDFs or an agent carrying a growing conversation history can push a single request past 200K, and the whole request then bills at the higher rate.
SayGm serves 3.1 Pro at $1.50 and $9.00, 25% below Google's standard rate, and the discount holds past the threshold: $3.00 input and $13.50 output above 200K tokens, against Google's $4.00 and $18.00. For anyone feeding 3.1 Pro whole codebases or document sets, that long-context rate is where most of the saving sits.
Gemini Flash pricing at the low end: Flash-Lite
Flash-Lite is the model for high-volume work that doesn't need reasoning depth: classification, extraction, routing, short summaries. 3.1 Flash-Lite is the cheapest Gemini 3 model at $0.25 and $1.50 on Google, and $0.19 and $1.13 on SayGm. 3.5 Flash-Lite is newer and costs a little more, $0.30 and $2.50.
Neither is affected by the January change.
Where to buy Gemini: Google, OpenRouter, Requesty and SayGm
Gemini is a closed model, so every route ends at Google. The options differ in how you're billed, what you pay on top, and who else is in the path.
Google AI Studio and Vertex AI

Buying direct gets you every tier: Standard, Flex, Batch and Priority, plus a free tier on the Flash models. The free tier is rate-limited, and Google's pricing page notes that free-tier usage is used to improve its products, which the paid tier is not. Vertex AI is the same models under Google Cloud billing, for teams that already buy cloud from Google.
OpenRouter

OpenRouter lists six Gemini 3.8 Flash endpoints, all Google's own (AI Studio and Vertex at the Flex, Standard and Priority tiers), and passes those rates through. The cost is its credit fee: 5.5% by card, 5% by crypto, which we break down in our OpenRouter alternatives comparison. On Standard that makes 100 million 3.8 Flash tokens about $142.
Requesty

Requesty passes Google's rate through with a 5% markup, so 3.8 Flash comes to about $0.79 and $3.94 per million, or $142 per 100 million tokens.
SayGm

$0.37 and $1.86 for 3.8 Flash, and between 25% and 51% below Google's standard rate across the Gemini 3 range, with no subscription and no credit fee. It also carries Claude, GPT and open-weight models behind the same key.
Switching is a base URL change in the Google GenAI SDK; the model name stays the same:
import os
from google import genai
from google.genai import types
client = genai.Client(
api_key=os.environ["GM_API_KEY"],
http_options=types.HttpOptions(base_url="https://api.saygm.com"),
)
response = client.models.generate_content(
model="gemini-3.8-flash",
contents="Summarise this contract in five bullet points.",
)
print(response.text)The Gemini compatibility guide covers streaming and REST.
How private is Gemini on SayGm?
Gemini is a frontier model, so Google still receives the prompt; that's true on every route. What SayGm changes is who sees it on the way. Requests use anonymous routing: Google gets the request but never sees who sent it, and the gateway in between runs inside hardware-attested confidential compute.
SayGm can also strip sensitive data out of a prompt before Google ever receives it. Guardrails are optional and set per API key in the dashboard. Once switched on, they redact personal data, secrets and credentials, match your own regular expressions and filter blocked words. They run inside the gateway's enclave, so the redaction happens before the request leaves for Google. They are pattern-based, which means they will miss some sensitive values and catch some harmless ones, so treat them as one layer of protection rather than a complete data-loss-prevention setup.
If you need a prompt nobody can read, including the model's operator, that guarantee comes from SayGm's open-weight models, which run inside the enclave itself.
Questions about Gemini API costs
Is the Gemini API free?
Partly. Google offers a free tier on the Flash and Flash-Lite models, with rate limits, and notes that free-tier usage is used to improve its products. Gemini 3.1 Pro has no free tier. For production traffic, the paid tier is the realistic option.
What does Gemini cost per 1,000 tokens?
To get Gemini's token cost per 1,000 tokens, divide the per-million rate by 1,000. Gemini 3.8 Flash on Google's standard tier is $0.00075 per 1,000 input tokens and $0.00375 per 1,000 output tokens until January. On SayGm it's about $0.00037 and $0.00186.
Will SayGm's Gemini prices change in January?
SayGm's rates track what providers charge, so expect them to move when Google's do. SayGm's model pages always show the current figure.
Which Gemini model should I use?
3.8 Flash for most work: it's the newest and, until January, among the cheapest. 3.1 Pro when you need the strongest reasoning and can stay under 200K tokens per request. Flash-Lite for high-volume tasks where speed and price matter more than depth.
What to budget for Gemini
If you're budgeting Gemini into 2027, use January's rates. Gemini 3.8 Flash costs $0.75 and $3.75 on Google until 31 December, then doubles. On SayGm it's $0.37 and $1.86 today, with Gemini 3.1 Pro at 25% below Google's rate and Flash-Lite from $0.19.
Set your Gemini client's base URL to api.saygm.com, keep the model name, and say gm to AI.
About SayGm
SayGm is a drop-in inference gateway for teams who don't want to just take a company's word that their prompts are private. Route to a confidential model and your prompt is sealed inside hardware-verified execution, where not even SayGm can see it. That's not a policy, it's provable. Swap in your existing OpenAI, Anthropic, or Gemini code and you're covered in minutes, at transparent, published rates with no hidden markup.
Say gm to AI at saygm.com.
- AI Inference
- Gemini
- Pricing


