Home/Google/Gemini 3.5 Flash Lite
Google pricing

Gemini 3.5 Flash Lite API Pricing & Cost Calculator

Gemini 3.5 Flash Lite is Google's low-cost multimodal tier: 1M context, native audio + vision, at $0.30/$2.50 per 1M tokens — 5x cheaper than Gemini 3.5 Flash. Cached input at $0.03/1M.

Google plans to retire Gemini 3.5 Flash Lite on 2027-07-21

Pricing below is still current. Worth factoring the retirement date into anything you are building on it long-term.

Still available from Google: Gemini 3.1 Pro, Gemini 3.5 Flash, Gemini 3.6 Flash

Input
$0.30
per 1M tokens
Output
$2.50
per 1M tokens
Context window
1049K
tokens
Released
2026-07
Cutoff
≈ Estimated tokenizer·$0.30 in·$2.50 out (per 1M)
Quick start with a use case
Total cost per call$0.001550
Input$0.000300
Output$0.001250
Cost comparison
Standard
$0.001550
With Caching
$0.001415
Save 9% ↓
With Batch
$0.000775
Save 50% ↓
Cutting costs? Open models start at $0.28/1M

DeepSeek, Kimi & 200+ open models on Novita AI — often 5–20x cheaper than proprietary APIs for comparable tasks.

Try Novita AI

Affiliate link — we may earn a commission at no extra cost to you.

Detailed pricing

Gemini 3.5 Flash Lite pricing breakdown

All pricing dimensions including caching and batch discounts.

TypePrice (per 1M tokens)
Input$0.3000
Output$2.5000
Cached input$0.0300
Batch input$0.1500
Batch output$1.2500
Cutting costs? Open models start at $0.28/1M

DeepSeek, Kimi & 200+ open models on Novita AI — often 5–20x cheaper than proprietary APIs for comparable tasks.

Try Novita AI

Affiliate link — we may earn a commission at no extra cost to you.

Last verified 2026-08-01 · Google official pricing · ⚠️ Spotted a wrong price? Report in 30s →

How it compares

Gemini 3.5 Flash Lite vs alternatives

Single-call cost (1000 input + 500 output tokens) ranked from cheapest.

ModelPer call
Gemini 3.5 Flash Lite
Google · this page
$0.001550
DeepSeek V3.2
DeepSeek
$0.000480
GPT-5.6 Luna
OpenAI
$0.000800
DeepSeek V4-Flash
DeepSeek
$0.000900
Gemini 3.1 Flash Lite
Google
$0.001000
GPT-5 mini
OpenAI
$0.001250
Recommended use

When to choose Gemini 3.5 Flash Lite

Gemini 3.5 Flash Lite shines for general-purpose tasks, image and document understanding, high-throughput, low-latency tasks, and audio understanding and generation. Token counts are estimated within ~10-20% margin.

Context window of 1049K tokens handles entire codebases or book-length documents.
Prompt caching available — significant savings for repeated system prompts.
Batch API support for non-realtime workloads at ~50% discount.
Tool / function calling supported.
FAQ

Frequently asked questions

Gemini 3.5 Flash Lite costs $0.30 per 1M input tokens and $2.50 per 1M output tokens. A typical chat call (1000 input + 500 output tokens) costs approximately $0.0015. Use the calculator above to estimate your specific use case.
Gemini 3.5 Flash Lite supports a 1049K token context window with a max output of 65,536 tokens.
Yes. Cached input is priced at $0.03 per 1M tokens — 90% cheaper than uncached input. This is especially valuable for repeated system prompts, long-context retrieval, and chat threads with shared history.
Yes. Gemini 3.5 Flash Lite supports the Batch API at $0.15 input / $1.25 output per 1M tokens — typically 50% off standard pricing. Batch jobs complete within ~24 hours, ideal for non-realtime workloads like overnight data processing or content generation pipelines.
Yes, Gemini 3.5 Flash Lite accepts image input. Vision token pricing is generally calculated based on image dimensions and folded into the input token count. Specific image-specific pricing varies — refer to Google's official documentation.
Token counts for Gemini 3.5 Flash Lite are estimated from character ratios (~10-20% margin). Cost calculations use prices verified on 2026-08-01. For final billing accuracy, always verify with Google's usage dashboard.
Yes, Google offers fine-tuning for Gemini 3.5 Flash Lite. Fine-tuned model pricing is separate from base model pricing — check Google's pricing page for current fine-tuning rates.