⚡ Cheapest in Gemini 3 Series — $0.25/$1.50 (GA May 2026)

Gemini 3.1 Flash-Lite

The cheapest and fastest Gemini 3 series model. Built on the Gemini 3 Pro architecture, matching 2.5 Flash across key capability areas at 17% lower input and 40% lower output pricing.

$0.25
Input / 1M tokens
$1.50
Output / 1M tokens
1,048,576
Context window
65,536
Max output tokens
Last updated: | Sources: Google AI Pricing · Google Cloud Docs | All Pricing · Calculator · 日本語版

Pricing (GA May 2026, current)

Verified against Google's official AI pricing page (ai.google.dev). All prices in USD per 1M tokens. Flash-Lite has flat pricing regardless of context length.

TierInput (text/image/video)Input (audio)Cached InputOutputNotes
Standard$0.25$0.50$0.025 / $0.05(audio)$1.50Default rate. Flat across all context lengths
Batch$0.125$0.25$0.0125 / $0.025$0.7550% off Standard. Up to 24h completion
Flex$0.125$0.25$0.0125 / $0.025$0.75Lower priority. 50% off Standard
Priority$0.45$0.90$0.045 / $0.09$2.70High priority. 1.8× Standard
💡 Thinking level cost control: Gemini 3.1 Flash-Lite supports 4 thinking levels: minimal, low, medium, high. Thinking tokens are billed as output tokens, so level choice directly impacts cost. Use minimal for bulk processing, medium/high for edge cases requiring deeper reasoning.
⚠️ Audio input is 2×: Audio input costs double the text/image/video rate. Factor this into cost estimates for audio-heavy workloads.

Primary sources: ai.google.dev/gemini-api/docs/pricing · Google Cloud Docs: Gemini 3.1 Flash-Lite (verified 2026-08-10)

Model Specifications

Preview Release
March 3, 2026
GA Release
May 7, 2026
Knowledge Cutoff
January 2025 (3rd party)
Context Window
1,048,576 tokens
Max Output
65,536 tokens
Input Modalities
Text + Image + Audio + Video
Output Modalities
Text
Tool Use
Function Calling ✓
Base Architecture
Gemini 3 Pro
Model ID
gemini-3.1-flash-lite

Sources: Google Cloud: Gemini 3.1 Flash-Lite · Google DeepMind Model Card

Performance — Matches 2.5 Flash Key Capabilities at Lower Cost

Built on the Gemini 3 Pro architecture and optimized for high throughput. Significant quality uplift over 2.5 Flash-Lite, matching 2.5 Flash across key capability areas.

  • Response quality: Matches Gemini 2.5 Flash performance across key capability areas
  • Instruction following: Improved for complex chatbot and instruction-heavy workflows
  • Audio input: Improved ASR quality for speech-to-text pipelines
  • Multimodal: Full text, image, audio, and video input support
  • Thinking levels: 4 levels (minimal–high) for fine-grained reasoning control. Minimal for bulk, medium/high for complex cases
  • vs Gemini 2.5 Flash: 17% cheaper input ($0.25 vs $0.30), 40% cheaper output ($1.50 vs $2.50)
  • vs Gemini 3 Flash: Half the price (input $0.25 vs $0.50, output $1.50 vs $3.00)

Google positions Flash-Lite as "our most cost-efficient Gemini model, optimized for low latency use cases for high-volume, cost-sensitive LLM traffic." Ideal for translation, data extraction, and code completion where per-token economics determine viability.

Sources: Google Cloud Docs · Google DeepMind Model Card · Google Blog

Competitor Price Comparison

Flash-Lite is among the cheapest models, though GPT-5.6 Luna's 80% price cut has intensified competition. All prices in USD per 1M tokens (Standard tier, text input).

ModelProviderInputOutputvs Flash-LiteNotes
Gemini 3.1 Flash-LiteGoogle$0.25$1.50Baseline. Multimodal input
GPT-5.6 LunaOpenAI$0.20$1.20Luna cheaper20% cheaper input & output. Text + image only
Gemini 3.5 Flash-LiteGoogle$0.30$2.503.1 cheaper3.1: 17% cheaper input, 40% cheaper output
Gemini 2.5 Flash-LiteGoogle$0.10$0.402.5 cheaperOldest gen. Large capability gap
Claude Haiku 4.5Anthropic$1.00$5.00Flash-Lite cheaperFlash-Lite: 75% cheaper input, 70% cheaper output
DeepSeek V4 FlashDeepSeek$0.14$0.28DeepSeek cheaperCurrent rate. Future price increase announced

* GPT-5.6 Luna comparison is text-input only. Gemini 3.1 Flash-Lite accepts audio and video input as well, making it potentially more cost-effective for multimodal workloads. DeepSeek V4 Flash currently lists at $0.14 per 1M cache-miss input and $0.28 per 1M output. DeepSeek has announced a future overall price increase, but has not published new rates or an effective date.

Sources: official pricing pages from each provider (verified 2026-08-10). See full pricing comparison. GPT-5.6 Luna details →

Ideal Use Cases

High-throughput, cost-sensitive, multimodal workloads.

🌐

Translation & Multi-Language

Bulk text translation pipelines. Use minimal thinking for maximum token efficiency.

📊

Data Extraction & Classification

Structured data extraction, content moderation, sentiment analysis at scale.

🎤

Speech Recognition (ASR)

Audio input with improved ASR quality over 2.5 Flash-Lite. Ideal for transcription pipelines.

💻

Code Completion & Review

Inline suggestions, lint-style review, refactor proposals at IDE/CI/CD scale.

🔄

Batch Processing & Nightly Jobs

Batch/Flex for 50% additional savings. Ideal for async large-scale data processing.

📹

Multimodal Bulk Processing

Classification and metadata extraction from mixed media including images, video, and audio.

Gemini 3.x Family Selection Guide

ModelInputOutputWhen to Use
Gemini 3.1 Pro Preview$2.00$12.00Hardest reasoning, agents, complex multi-step problems
Gemini 3.6 Flash$1.50$7.50Speed-focused production workloads. Grounding included
Gemini 3.5 Flash$1.50$9.00Standard Flash tier. Slightly higher output cost than 3.6
Gemini 3.1 Flash-Lite ⬅$0.25$1.50High-throughput, cost-sensitive, multimodal bulk processing

💡 Strategy: Handle routine high-frequency tasks with Flash-Lite, escalate to Flash when quality matters, and reserve Pro for the hardest problems.

Estimate Your Actual Costs

Compare token costs for Gemini 3.1 Flash-Lite and all major models. Just enter your token counts.

🧮 Open Token Calculator →

📊 Full Pricing Comparison · 🌙 Compare with GPT-5.6 Luna · 🗓️ Release Timeline · 🔗 Google Official Pricing

⚠️ Disclaimer

  • Pricing verified against Google's official AI page (ai.google.dev) on 2026-08-10. Subject to change without notice. Always check official pages before contracting.
  • This site is informational only. We are not affiliated with or endorsed by any provider.
  • Performance claims are based on Google's published data and third-party evaluations. Your mileage may vary.