⚡ The balanced model of the GPT-6 family — officially “Near-Astra performance for complex work at a lower cost”

GPT-6.1 Sol

OpenAI's balance point between intelligence and cost in the GPT-6 family. The model ID is gpt-6.1-sol, and the official model page describes it as “Near-Astra performance for complex work at a lower cost.” Pricing is $2 input / $10 output, with cached input at 5% of input — $0.10.

$2.00
Input / 1M
$0.10
Cached input / 1M
$10.00
Output / 1M
1.05M
Context

The short version

  • Input $2 and output $10 match GPT-6 Sol, but cached input drops to $0.10 (5% of input) from $0.20 (90% off) on GPT-6 Sol.
  • Reasoning effort: low / medium (default) / high / xhigh / max. none and minimal are not supported — GPT-6 Sol and Luna do support none.
  • Tool calling requires the Responses API. Chat Completions works only without tools.
  • Speed tiers change the price: Fast is 2x Standard, Ultrafast 6x, Batch and Flex 50% off. Regional processing adds 10%.
  • Prompts above 272,000 input tokens are repriced across the whole request (2x input and cache rates, 1.5x output) — a real constraint on one-shot long-document work.
  • OpenAI's model page publishes no release date. Third-party reports agree on September 29, 2026; we flag that as third-party information rather than stating it as official.

Pricing (USD per 1M tokens, from the official model page)

Standard prices are as published. Fast, Batch/Flex and Ultrafast figures are our calculations from the official multipliers (2x, 50% off, 6x).

ItemStandardFast (2x)Batch / Flex (50% off)Ultrafast (6x)
Input$2.00$4.00$1.00$12.00
Cached input$0.10$0.20$0.05$0.60
Cache writes (1.25x)$2.50$5.00$1.25$15.00
Output$10.00$20.00$5.00$60.00

※ Official notes: cached input is “priced at 5% of the uncached input token rate”; cache writes are “billed at 1.25x the uncached input token rate”; Fast mode is 2x Standard, Batch and Flex 50% lower, Ultrafast 6x, regional processing +10%.

Source: OpenAI — GPT-6.1 Sol (Pricing), checked 2026-10-09.

The 272K wall: one-shot long prompts are repriced

The official note reads: prompts with more than 272,000 input tokens “are priced at 2x input and cache rates and 1.5x output for the full request.” Not just the overflow.

CaseCalculationCost
300,000 input + 5,000 output tokens in a single request$4/M × 0.3 input + $15/M × 0.005 output
2x input and 1.5x output because the request exceeds 272K
$1.275
The same work split into three 100,000-token requests$2/M × 0.3 input + output per call$0.60 + output
  • “1.05M context” is not the same as “send 1.05M in one go.” Past 272K the whole request is repriced, which makes long-document work a design decision.
  • For long conversations, $0.10 cached input matters — until the request crosses 272K, where cached rates double too ($0.20/M).
  • All figures are our calculations from the official multipliers; confirm against your own usage reports.

How it differs from GPT-6 Sol

ItemGPT-6.1 SolGPT-6 Sol (previous gen)
Input$2.00$2.00
Cached input$0.10 (5% of input)$0.20 (90% off)
Cache writes$2.50 (1.25x input)Not stated on the model page
Output$10.00$10.00
Reasoning effortlow / medium (default) / high / xhigh / max — no none or minimalIncludes none (which allows function calling on Chat Completions)
Context / max output1,050,000 / 128,000 (max input 922,000)1,050,000 / 128,000
Knowledge cutoffApril 30, 2026April 2026

※ GPT-6 Sol figures are as listed on our pricing comparison (verified against OpenAI's page on 2026-09-25); GPT-6.1 Sol figures are from its model page, checked 2026-10-09.

Migration notes worth acting on

Tool calling → Responses API
“Use the Responses API for tool calling. Chat Completions is supported without tool calling.” That is verbatim from the model page.
Drop temperature when effort ≠ none
When reasoning effort is not none, remove temperature, top_p and top_logprobs (plus logprobs on Chat Completions).
Cache TTL API changed
Replace prompt_cache_retention (GPT-5.5 and earlier) with prompt_cache_options.ttl set to "30m".
Data residency
US and EU, including Fast and Ultrafast modes. Regional processing adds a 10% premium.

Specifications

ItemValue (official model page)
Model ID / snapshotgpt-6.1-sol
Context window1,050,000 tokens
Max input tokens922,000
Max output128,000 tokens
Modalitiestext, images → text (audio and video unsupported)
Knowledge cutoffApril 30, 2026 (reasoning token support)
EndpointsResponses, Chat Completions (without tools), Batch. Live / Realtime / Assistants / Embeddings / image generation / audio / fine-tuning are not supported.
Toolsweb search, file search, image generation, code interpreter, computer use, MCP, apply_patch and more
Rate limits (Standard)Build 5,000 RPM / 1,000,000 TPM; Launch 10,000 / 4,000,000; Grow 15,000 / 40,000,000

On benchmarks (our policy)

OpenAI's model page publishes no numeric benchmark table (checked 2026-10-09). What it publishes is the positioning statement: “Near-Astra performance for complex work at a lower cost.”

  • Third-party trackers publish scores, but we do not print numbers the vendor has not published — filling gaps with estimates would break our rule of publishing only what primary sources support.
  • The most reliable evaluation is the one the vendor itself recommends: “Compare it with Astra on your tasks to assess the tradeoff between quality and cost.”

Related pages

⚠️ Disclaimer

  • Pricing and specifications were checked on 2026-10-09 against OpenAI's documentation (GPT-6.1 Sol model page and pricing page).
  • Fast, Batch/Flex, Ultrafast and above-272K figures are our own calculations from the official multipliers; confirm against your invoice.
  • The model page states no release date, so we label the date reported by third parties as third-party information.
  • This site is informational and does not endorse or represent any provider.