Home/Google/Gemini 3 Flash
Google pricing

Gemini 3 Flash API Pricing & Cost Calculator

Gemini 3 Flash balances speed and capability with 1M context. Native audio + vision. $0.05/1M cached input makes it dominant for high-volume multimodal pipelines.

Input
$0.50
per 1M tokens
Output
$3.00
per 1M tokens
Context window
1049K
tokens
Released
2025-12
Cutoff 2025-10
≈ Estimated tokenizer·$0.50 in·$3.00 out (per 1M)
Quick start with a use case
Total cost per call$0.002000
Input$0.000500
Output$0.001500
Cost comparison
Standard
$0.002000
With Caching
$0.001775
Save 11% ↓
With Batch
$0.001000
Save 50% ↓
Cutting costs? Open models start at $0.28/1M

DeepSeek, Kimi & 200+ open models on Novita AI — often 5–20x cheaper than proprietary APIs for comparable tasks.

Try Novita AI

Affiliate link — we may earn a commission at no extra cost to you.

Detailed pricing

Gemini 3 Flash pricing breakdown

All pricing dimensions including caching and batch discounts.

TypePrice (per 1M tokens)
Input$0.5000
Output$3.0000
Cached input$0.0500
Batch input$0.2500
Batch output$1.5000
Cutting costs? Open models start at $0.28/1M

DeepSeek, Kimi & 200+ open models on Novita AI — often 5–20x cheaper than proprietary APIs for comparable tasks.

Try Novita AI

Affiliate link — we may earn a commission at no extra cost to you.

Last verified 2026-08-01 · Google official pricing · ⚠️ Spotted a wrong price? Report in 30s →

How it compares

Gemini 3 Flash vs alternatives

Single-call cost (1000 input + 500 output tokens) ranked from cheapest.

ModelPer call
Gemini 3 Flash
Google · this page
$0.002000
GPT-6 Luna
OpenAI
$0.000350
DeepSeek V3.2
DeepSeek
$0.000480
GPT-5.6 Luna
OpenAI
$0.000800
DeepSeek V4-Flash
DeepSeek
$0.000900
Gemini 3.1 Flash Lite
Google
$0.001000
Recommended use

When to choose Gemini 3 Flash

Gemini 3 Flash shines for general-purpose tasks, image and document understanding, high-throughput, low-latency tasks, and audio understanding and generation. Token counts are estimated within ~10-20% margin.

✓Context window of 1049K tokens handles entire codebases or book-length documents.
✓Prompt caching available — significant savings for repeated system prompts.
✓Batch API support for non-realtime workloads at ~50% discount.
✓Tool / function calling supported.
FAQ

Frequently asked questions

Gemini 3 Flash costs $0.50 per 1M input tokens and $3.00 per 1M output tokens. A typical chat call (1000 input + 500 output tokens) costs approximately $0.0020. Use the calculator above to estimate your specific use case.
Gemini 3 Flash supports a 1049K token context window with a max output of 65,535 tokens. Knowledge cutoff: 2025-10.
Yes. Cached input is priced at $0.05 per 1M tokens — 90% cheaper than uncached input. This is especially valuable for repeated system prompts, long-context retrieval, and chat threads with shared history.
Yes. Gemini 3 Flash supports the Batch API at $0.25 input / $1.50 output per 1M tokens — typically 50% off standard pricing. Batch jobs complete within ~24 hours, ideal for non-realtime workloads like overnight data processing or content generation pipelines.
Yes, Gemini 3 Flash accepts image input. Vision token pricing is generally calculated based on image dimensions and folded into the input token count. Specific image-specific pricing varies — refer to Google's official documentation.
Token counts for Gemini 3 Flash are estimated from character ratios (~10-20% margin). Cost calculations use prices verified on 2026-08-01. For final billing accuracy, always verify with Google's usage dashboard.
Yes, Google offers fine-tuning for Gemini 3 Flash. Fine-tuned model pricing is separate from base model pricing — check Google's pricing page for current fine-tuning rates.