Google Gemini API Pricing (2026)
Current Gemini API token prices, with Gemini 3.8 Flash's temporary launch pricing separated from the higher rates Google has scheduled for January 1, 2027.
Verified against Google pricing documentation: September 3, 2026Current Gemini API prices
Prices below are for the Gemini Developer API paid tier and are quoted per one million tokens. They are not Vertex AI contract quotes. Audio input can carry a different price, and tool usage such as grounding may add separate charges.
| Model | Standard input | Standard output | Cached input | Batch input / output |
|---|---|---|---|---|
| Gemini 3.8 Flash GA |
$0.75 | $3.75 | $0.075 | $0.375 / $1.875 |
| Gemini 3.7 Flash Available |
$0.75 | $3.75 | $0.075 | $0.375 / $1.875 |
| Gemini 3.6 Flash Available |
$0.75 | $3.75 | $0.075 | $0.375 / $1.875 |
| Gemini 3.5 Flash Available |
$1.50 | $9.00 | $0.15 | $0.75 / $4.50 |
| Gemini 3.5 Flash-Lite Available |
$0.30 | $2.50 | $0.03 | $0.15 / $1.25 |
| Gemini 3.1 Flash-Lite Lowest standard text rate |
$0.25* | $1.50 | $0.025* | $0.125 / $0.75* |
*Gemini 3.1 Flash-Lite prices shown are for text, image and video input. Audio input is $0.50/M standard and $0.25/M for Batch/Flex. Listed output prices include thinking tokens where Google specifies this.
Gemini 3.8 Flash: launch pricing and 2027 rates
Google publishes separate Standard, Batch, Flex and Priority schedules for gemini-3.8-flash. Batch and Flex currently use the same token rates, but they are different service modes: choose based on latency, processing and capacity requirements rather than price alone.
| Mode | Through Dec 31, 2026 input / output |
From Jan 1, 2027 input / output |
Cached input now โ 2027 |
|---|---|---|---|
| Standard | $0.75 / $3.75 | $1.50 / $7.50 | $0.075 โ $0.15 |
| Batch | $0.375 / $1.875 | $0.75 / $3.75 | $0.0375 โ $0.075 |
| Flex | $0.375 / $1.875 | $0.75 / $3.75 | $0.0375 โ $0.075 |
| Priority | $1.35 / $6.75 | $2.70 / $13.50 | $0.135 โ $0.27 |
Gemini 3.8 cache storage costs $0.50 per million tokens per hour through December 31, 2026 and $1.00 from January 1, 2027. Token prices alone do not include separately billed tools or grounding.
Which Gemini model is cheapest?
Lowest standard token rate
Gemini 3.1 Flash-Lite is the cheapest general-purpose option listed here for text, image and video workloads: $0.25/M input and $1.50/M output.
Newest Flash model
Gemini 3.8 Flash is the current GA Flash release. Its launch rates are higher than 3.1 Flash-Lite but materially lower than its scheduled 2027 rates.
Asynchronous bulk work
Batch halves Gemini 3.8's current standard token rates. It suits work that does not need an immediate interactive response.
Predictable capacity
Priority costs more than Standard. The premium is for the service mode, not a different model output-quality tier.
Free tier versus paid tier
| Area | Free tier | Paid tier |
|---|---|---|
| Model access | Limited access to certain models | Higher limits and access to advanced models |
| Input/output tokens | Free for supported models within limits | Usage billed at published rates |
| Context caching | Not generally available | Available, with cached-token and storage charges |
| Batch API | Not available | Available at lower token rates |
| Data use | Content may be used to improve Google products | Content is not used to improve Google products |
Google can change model availability and rate limits by tier and project. Check the API console for the limits that apply to your account rather than relying on a fixed requests-per-day estimate.
Long context, caching and usage-dependent costs
Gemini 3.8 Flash supports a 1,048,576-token input limit and up to 65,536 output tokens. Its published Gemini Developer API table does not add a separate long-context token tier. Large prompts still cost more because they contain more billable tokens.
- Cached context: discounted input tokens plus an hourly storage charge.
- Thinking: included in the published output-token price for the models listed above.
- Grounding and tools: may be billed separately and can make the final request cost usage-dependent.
- Vertex AI: pricing, commitments and enterprise terms can differ from the Gemini Developer API.
What happened to Gemini 2.5 and older models?
Gemini 2.5 Flash and Gemini 2.5 Flash-Lite base model codes still have no shutdown date listed in Google's deprecation schedule, but they are no longer the current Flash generation. Gemini 2.0 models were shut down on June 1, 2026, and multiple dated preview aliases have also been retired. This page therefore uses current 3.x models for new pricing decisions rather than presenting Gemini 1.5 or 2.0 prices as current options.
Frequently asked questions
What does Gemini 3.8 Flash cost?
Through December 31, 2026, Standard costs $0.75/M input and $3.75/M output. Google schedules $1.50/M input and $7.50/M output from January 1, 2027.
Which Gemini API model is cheapest?
Gemini 3.1 Flash-Lite has the lowest general-purpose standard text rate in this comparison at $0.25/M input and $1.50/M output. Audio input is priced differently.
Does the Gemini API have a free tier?
Yes. It provides limited access to certain models with free input and output. Availability and rate limits vary; Batch and paid context caching are not part of the general free tier.
How do Batch and Flex differ from standard pricing?
For Gemini 3.8 Flash through December 2026, both are priced at $0.375/M input and $1.875/M output, half the Standard token rates. Their processing and service characteristics differ, so they are not interchangeable for every workload.
When does Gemini 3.8 pricing increase?
The scheduled increase takes effect January 1, 2027. The lower launch schedule remains current through December 31, 2026.
Sources and methodology
This guide uses Google's official Gemini Developer API pricing, Gemini 3.8 Flash model documentation and model deprecation schedule. Rates are normalized to US dollars per one million tokens. Scheduled future rates are labeled separately and are not treated as active prices.