Pricing (GA May 2026, current)
Verified against Google's official AI pricing page (ai.google.dev). All prices in USD per 1M tokens. Flash-Lite has flat pricing regardless of context length.
| Tier | Input (text/image/video) | Input (audio) | Cached Input | Output | Notes |
|---|---|---|---|---|---|
| Standard | $0.25 | $0.50 | $0.025 / $0.05(audio) | $1.50 | Default rate. Flat across all context lengths |
| Batch | $0.125 | $0.25 | $0.0125 / $0.025 | $0.75 | 50% off Standard. Up to 24h completion |
| Flex | $0.125 | $0.25 | $0.0125 / $0.025 | $0.75 | Lower priority. 50% off Standard |
| Priority | $0.45 | $0.90 | $0.045 / $0.09 | $2.70 | High priority. 1.8× Standard |
Primary sources: ai.google.dev/gemini-api/docs/pricing · Google Cloud Docs: Gemini 3.1 Flash-Lite (verified 2026-08-10)
Model Specifications
Sources: Google Cloud: Gemini 3.1 Flash-Lite · Google DeepMind Model Card
Performance — Matches 2.5 Flash Key Capabilities at Lower Cost
Built on the Gemini 3 Pro architecture and optimized for high throughput. Significant quality uplift over 2.5 Flash-Lite, matching 2.5 Flash across key capability areas.
- Response quality: Matches Gemini 2.5 Flash performance across key capability areas
- Instruction following: Improved for complex chatbot and instruction-heavy workflows
- Audio input: Improved ASR quality for speech-to-text pipelines
- Multimodal: Full text, image, audio, and video input support
- Thinking levels: 4 levels (minimal–high) for fine-grained reasoning control. Minimal for bulk, medium/high for complex cases
- vs Gemini 2.5 Flash: 17% cheaper input ($0.25 vs $0.30), 40% cheaper output ($1.50 vs $2.50)
- vs Gemini 3 Flash: Half the price (input $0.25 vs $0.50, output $1.50 vs $3.00)
Google positions Flash-Lite as "our most cost-efficient Gemini model, optimized for low latency use cases for high-volume, cost-sensitive LLM traffic." Ideal for translation, data extraction, and code completion where per-token economics determine viability.
Sources: Google Cloud Docs · Google DeepMind Model Card · Google Blog
Competitor Price Comparison
Flash-Lite is among the cheapest models, though GPT-5.6 Luna's 80% price cut has intensified competition. All prices in USD per 1M tokens (Standard tier, text input).
| Model | Provider | Input | Output | vs Flash-Lite | Notes |
|---|---|---|---|---|---|
| Gemini 3.1 Flash-Lite | $0.25 | $1.50 | — | Baseline. Multimodal input | |
| GPT-5.6 Luna | OpenAI | $0.20 | $1.20 | Luna cheaper | 20% cheaper input & output. Text + image only |
| Gemini 3.5 Flash-Lite | $0.30 | $2.50 | 3.1 cheaper | 3.1: 17% cheaper input, 40% cheaper output | |
| Gemini 2.5 Flash-Lite | $0.10 | $0.40 | 2.5 cheaper | Oldest gen. Large capability gap | |
| Claude Haiku 4.5 | Anthropic | $1.00 | $5.00 | Flash-Lite cheaper | Flash-Lite: 75% cheaper input, 70% cheaper output |
| DeepSeek V4 Flash | DeepSeek | $0.14 | $0.28 | DeepSeek cheaper | Current rate. Future price increase announced |
* GPT-5.6 Luna comparison is text-input only. Gemini 3.1 Flash-Lite accepts audio and video input as well, making it potentially more cost-effective for multimodal workloads. DeepSeek V4 Flash currently lists at $0.14 per 1M cache-miss input and $0.28 per 1M output. DeepSeek has announced a future overall price increase, but has not published new rates or an effective date.
Sources: official pricing pages from each provider (verified 2026-08-10). See full pricing comparison. GPT-5.6 Luna details →
Ideal Use Cases
High-throughput, cost-sensitive, multimodal workloads.
Translation & Multi-Language
Bulk text translation pipelines. Use minimal thinking for maximum token efficiency.
Data Extraction & Classification
Structured data extraction, content moderation, sentiment analysis at scale.
Speech Recognition (ASR)
Audio input with improved ASR quality over 2.5 Flash-Lite. Ideal for transcription pipelines.
Code Completion & Review
Inline suggestions, lint-style review, refactor proposals at IDE/CI/CD scale.
Batch Processing & Nightly Jobs
Batch/Flex for 50% additional savings. Ideal for async large-scale data processing.
Multimodal Bulk Processing
Classification and metadata extraction from mixed media including images, video, and audio.
Gemini 3.x Family Selection Guide
| Model | Input | Output | When to Use |
|---|---|---|---|
| Gemini 3.1 Pro Preview | $2.00 | $12.00 | Hardest reasoning, agents, complex multi-step problems |
| Gemini 3.6 Flash | $1.50 | $7.50 | Speed-focused production workloads. Grounding included |
| Gemini 3.5 Flash | $1.50 | $9.00 | Standard Flash tier. Slightly higher output cost than 3.6 |
| Gemini 3.1 Flash-Lite ⬅ | $0.25 | $1.50 | High-throughput, cost-sensitive, multimodal bulk processing |
💡 Strategy: Handle routine high-frequency tasks with Flash-Lite, escalate to Flash when quality matters, and reserve Pro for the hardest problems.
Estimate Your Actual Costs
Compare token costs for Gemini 3.1 Flash-Lite and all major models. Just enter your token counts.
🧮 Open Token Calculator →📊 Full Pricing Comparison · 🌙 Compare with GPT-5.6 Luna · 🗓️ Release Timeline · 🔗 Google Official Pricing
⚠️ Disclaimer
- Pricing verified against Google's official AI page (ai.google.dev) on 2026-08-10. Subject to change without notice. Always check official pages before contracting.
- This site is informational only. We are not affiliated with or endorsed by any provider.
- Performance claims are based on Google's published data and third-party evaluations. Your mileage may vary.