The short version
- Input $2 and output $10 match GPT-6 Sol, but cached input drops to $0.10 (5% of input) from $0.20 (90% off) on GPT-6 Sol.
- Reasoning effort: low / medium (default) / high / xhigh / max.
noneandminimalare not supported — GPT-6 Sol and Luna do supportnone. - Tool calling requires the Responses API. Chat Completions works only without tools.
- Speed tiers change the price: Fast is 2x Standard, Ultrafast 6x, Batch and Flex 50% off. Regional processing adds 10%.
- Prompts above 272,000 input tokens are repriced across the whole request (2x input and cache rates, 1.5x output) — a real constraint on one-shot long-document work.
- OpenAI's model page publishes no release date. Third-party reports agree on September 29, 2026; we flag that as third-party information rather than stating it as official.
Pricing (USD per 1M tokens, from the official model page)
Standard prices are as published. Fast, Batch/Flex and Ultrafast figures are our calculations from the official multipliers (2x, 50% off, 6x).
| Item | Standard | Fast (2x) | Batch / Flex (50% off) | Ultrafast (6x) |
|---|---|---|---|---|
| Input | $2.00 | $4.00 | $1.00 | $12.00 |
| Cached input | $0.10 | $0.20 | $0.05 | $0.60 |
| Cache writes (1.25x) | $2.50 | $5.00 | $1.25 | $15.00 |
| Output | $10.00 | $20.00 | $5.00 | $60.00 |
※ Official notes: cached input is “priced at 5% of the uncached input token rate”; cache writes are “billed at 1.25x the uncached input token rate”; Fast mode is 2x Standard, Batch and Flex 50% lower, Ultrafast 6x, regional processing +10%.
Source: OpenAI — GPT-6.1 Sol (Pricing), checked 2026-10-09.
The 272K wall: one-shot long prompts are repriced
The official note reads: prompts with more than 272,000 input tokens “are priced at 2x input and cache rates and 1.5x output for the full request.” Not just the overflow.
| Case | Calculation | Cost |
|---|---|---|
| 300,000 input + 5,000 output tokens in a single request | $4/M × 0.3 input + $15/M × 0.005 output 2x input and 1.5x output because the request exceeds 272K | $1.275 |
| The same work split into three 100,000-token requests | $2/M × 0.3 input + output per call | $0.60 + output |
- “1.05M context” is not the same as “send 1.05M in one go.” Past 272K the whole request is repriced, which makes long-document work a design decision.
- For long conversations, $0.10 cached input matters — until the request crosses 272K, where cached rates double too ($0.20/M).
- All figures are our calculations from the official multipliers; confirm against your own usage reports.
How it differs from GPT-6 Sol
| Item | GPT-6.1 Sol | GPT-6 Sol (previous gen) |
|---|---|---|
| Input | $2.00 | $2.00 |
| Cached input | $0.10 (5% of input) | $0.20 (90% off) |
| Cache writes | $2.50 (1.25x input) | Not stated on the model page |
| Output | $10.00 | $10.00 |
| Reasoning effort | low / medium (default) / high / xhigh / max — no none or minimal | Includes none (which allows function calling on Chat Completions) |
| Context / max output | 1,050,000 / 128,000 (max input 922,000) | 1,050,000 / 128,000 |
| Knowledge cutoff | April 30, 2026 | April 2026 |
※ GPT-6 Sol figures are as listed on our pricing comparison (verified against OpenAI's page on 2026-09-25); GPT-6.1 Sol figures are from its model page, checked 2026-10-09.
Migration notes worth acting on
none, remove temperature, top_p and top_logprobs (plus logprobs on Chat Completions).prompt_cache_retention (GPT-5.5 and earlier) with prompt_cache_options.ttl set to "30m".Specifications
| Item | Value (official model page) |
|---|---|
| Model ID / snapshot | gpt-6.1-sol |
| Context window | 1,050,000 tokens |
| Max input tokens | 922,000 |
| Max output | 128,000 tokens |
| Modalities | text, images → text (audio and video unsupported) |
| Knowledge cutoff | April 30, 2026 (reasoning token support) |
| Endpoints | Responses, Chat Completions (without tools), Batch. Live / Realtime / Assistants / Embeddings / image generation / audio / fine-tuning are not supported. |
| Tools | web search, file search, image generation, code interpreter, computer use, MCP, apply_patch and more |
| Rate limits (Standard) | Build 5,000 RPM / 1,000,000 TPM; Launch 10,000 / 4,000,000; Grow 15,000 / 40,000,000 |
On benchmarks (our policy)
OpenAI's model page publishes no numeric benchmark table (checked 2026-10-09). What it publishes is the positioning statement: “Near-Astra performance for complex work at a lower cost.”
- Third-party trackers publish scores, but we do not print numbers the vendor has not published — filling gaps with estimates would break our rule of publishing only what primary sources support.
- The most reliable evaluation is the one the vendor itself recommends: “Compare it with Astra on your tasks to assess the tradeoff between quality and cost.”
Related pages
- GPT-6 Sol / GPT-6 Luna — the sibling models (Sol at $2/$10, Luna at $0.10/$0.50)
- LLM API pricing comparison
- Token cost calculator
- Claude Sonnet 5.5 — the same price band ($2/$10, $0.10 cache reads)
- Prompt caching explained
- Release timeline
⚠️ Disclaimer
- Pricing and specifications were checked on 2026-10-09 against OpenAI's documentation (GPT-6.1 Sol model page and pricing page).
- Fast, Batch/Flex, Ultrafast and above-272K figures are our own calculations from the official multipliers; confirm against your invoice.
- The model page states no release date, so we label the date reported by third parties as third-party information.
- This site is informational and does not endorse or represent any provider.