💰 July 30, 2026: 80% price cut — $1.00→$0.20 / $6.00→$1.20

GPT-5.6 Luna

OpenAI's fastest and cheapest model gets a dramatic price cut. Now cheaper than Gemini 3.1 Flash-Lite, with GPT-5.5-beating performance at 1/25th the cost.

$0.20
Input / 1M tokens
$1.20
Output / 1M tokens
1.05M
Context window
128K
Max output tokens
Last updated: | Sources: OpenAI API Docs · OpenAI Pricing | All Pricing · Calculator · Gemini 3.1 Flash-Lite · 日本語版

Pricing (effective July 30, 2026)

Verified against OpenAI's official pricing page (developers.openai.com). All prices in USD per 1M tokens.

TierInputCached InputOutputNotes
Standard (Short context)$0.20$0.02$1.20≤272K input tokens. Default rate
Long context$0.40$0.04$1.80>272K input tokens. 2× input, 1.5× output
Batch$0.10$0.01$0.6050% off Standard. Up to 24h completion
Flex$0.10$0.01$0.60Lower priority. 50% off Standard
Price history:
$1.00 / $6.00 (at GA launch, July 9, 2026)
→ $0.20 / $1.20 (July 30, 2026 — 80% cut, just 21 days after launch)
* Cache writes are 1.25× standard input ($0.25). Cache reads are 10% of standard input. Minimum cache life: 30 minutes.
⚠️ Short/Long context boundary: 272K input tokens (confirmed in OpenAI API docs). Prompts exceeding 272K tokens have long-context pricing applied to the entire request, including cached input.

Primary sources: developers.openai.com/api/docs/pricing · GPT-5.6 Luna Model Docs (verified 2026-08-07)

Model Specifications

GA Release
July 9, 2026
Knowledge Cutoff
Feb 16, 2026
Context Window
1,050,000 tokens
Max Output
128,000 tokens
Input Modalities
Text + Image
Output Modalities
Text
Tool Use
Function Calling ✓
Tier
Equivalent to nano tier
Regional Uplift
+10% (eligible regions)
Model ID
gpt-5.6-luna

Sources: OpenAI GPT-5.6 Luna Docs · GPT-5.6 Announcement

Performance — Beyond "Nano"

Positioned as the nano-tier equivalent, yet outperforms the previous flagship on key benchmarks.

  • Artificial Analysis Quality Index: 51 — high composite score for a lightweight model, in the upper tier
  • Agents' Last Exam: Beats GPT-5.5. Only 2.4 points behind Sol on professional agentic tasks
  • HealthBench Professional / DeepSWE: Both exceed GPT-5.5
  • Coding: BenchLM rank #6 out of 215 models
  • Vision Evals (Roboflow): 74.1% overall (#6 of 16)
  • vs Claude Fable 5: Beats Fable 5 on Agents' Last Exam at ~1/16th the estimated cost per task

OpenAI claims Luna delivers "frontier-class performance from one year ago, at ~6% of the cost per task and ~9× the speed."

Sources: Artificial Analysis · BenchLM.ai · OpenAI

Competitor Price Comparison

After the 80% cut, Luna is now among the cheapest lightweight models. All prices in USD per 1M tokens (standard, short context).

ModelProviderInputOutputvs LunaNotes
GPT-5.6 LunaOpenAI$0.20$1.20Baseline
GPT-5.4 nanoOpenAI$0.20$1.25Same inputOlder gen. Luna is more capable
Gemini 3.5 Flash-LiteGoogle$0.30$2.50Luna cheaperLuna: 33% cheaper input, 52% cheaper output
Gemini 3.1 Flash-LiteGoogle$0.25$1.50Luna cheaperDetail page. Luna: 20% cheaper input & output
Gemini 2.5 Flash-LiteGoogle$0.10$0.40Gemini cheaperOlder gen, large capability gap
Claude Haiku 4.5Anthropic$1.00$5.00Luna cheaperLuna: 80% cheaper input, 76% cheaper output
DeepSeek V4 FlashDeepSeek$0.14$0.28DeepSeek cheaperCurrent official rate. Future price increase announced

* DeepSeek V4 Flash is currently listed at $0.14 per 1M cache-miss input tokens and $0.28 per 1M output tokens. DeepSeek has announced a future overall price increase, but has not published the new rates or an effective date.

Sources: official pricing pages from each provider (verified 2026-08-07). See full pricing comparison. Gemini 3.1 Flash-Lite details →

Ideal Use Cases

High-throughput, cost-sensitive workloads. Luna excels in these scenarios:

💬

Chat & Customer Support

Real-time high-volume conversations. Low latency meets low cost.

🏷️

Classification & Extraction

Document classification, sentiment analysis, entity extraction — routine NLP.

🤖

Lightweight Agent Workflows

Automation with tool use / function calling. Agent performance beats GPT-5.5.

✍️

Summarization & Drafting

Bulk document summarization, email drafting, proofreading at scale.

🔄

Batch Processing & Pipelines

Batch/Flex for 50% additional savings. Ideal for overnight data processing.

🔍

RAG & Search-Augmented Gen

1.05M context window for large-document search and Q&A.

GPT-5.6 Family Selection Guide

ModelInputOutputWhen to Use
GPT-5.6 Sol$5.00$30.00Hardest tasks, long-horizon agents, high-failure-cost work
GPT-5.6 Terra$2.00$12.00Standard coding, analysis, production agents
GPT-5.6 Luna ⬅$0.20$1.20High-throughput, cost-sensitive, routine processing, first-pass review

💡 Strategy: Start with Luna, escalate to Terra→Sol when quality is insufficient. For text-centric workloads, Luna is the cost leader. For multimodal workloads including audio/video, also consider Gemini 3.1 Flash-Lite.

Estimate Your Actual Costs

Compare token costs for GPT-5.6 Luna and all major models. Just enter your token counts.

🧮 Open Token Calculator →

📊 Full Pricing Comparison · ⚡ Compare with Gemini 3.1 Flash-Lite · 🗓️ Release Timeline · 🔗 OpenAI Official Pricing

⚠️ Disclaimer

  • Pricing verified against OpenAI's official page (developers.openai.com) on 2026-08-07. Subject to change without notice. Always check official pages before contracting.
  • This site is informational only. We are not affiliated with or endorsed by any provider.
  • Benchmark scores are as published by evaluation organizations and OpenAI. Your mileage may vary.