What is a 1M-token context window? The limit, the price step, and the accuracy ceiling

Glossary entry 4. This one covers the 1M-token context window. “Up to one million tokens” quietly mixes up two different claims: how much actually fits (capacity), and what it costs once you put it there (billing). Billing turns out to work very differently at OpenAI, Google and Anthropic: OpenAI and Google raise the rate across the whole request once you cross the line, while Anthropic stays on standard pricing for its 1M models. And a bigger window does not mean the model reads all of it equally well. We keep the numbers and the sources in, so you can work out what it means for your own workload.

“Supports 1M” does not mean “uses 1M well, at the same price”

A context window sets the physical ceiling. Whether the tokens inside are treated evenly, and whether they are billed at a flat rate, are two separate questions. This entry exists to keep those three things apart.

A picture first. A 1M context window is like being lent a very large desk. But the more you pile on it, the harder the right page is to find — and the size of the desk is tied to the shipping fee. Past a certain weight it is not just the excess that costs more: the whole shipment does. That is how OpenAI and Google structure their tiers; Anthropic is closer to "bigger desk, same shipping fee."
● Last updated:  |  Primary sources: OpenAI (GPT-6 Astra model page) · Gemini API pricing · Anthropic (Context windows) · DeepSeek (Models & Pricing) · Google Cloud (pricing footnotes)  |  Checked 2026-09-21

Short version: the ceiling is capacity, the bill is a separate rule

A context window is the maximum amount of information a model can look at in a single request. A 1M (one million) token window means the system prompt, the conversation history, the documents you paste in and the tool outputs can all add up to one million tokens in one call.
But “up to a million tokens fit” is not the same as “up to a million tokens cost the same.” The long-input rules differ by vendor. Short version first:
  • OpenAI: above 272K tokens, the whole request is billed at 2x input and 1.5x output — not just the excess.
  • Google: above 200K tokens, all tokens are charged at the long-context rate.
  • Anthropic: standard pricing all the way to 1M — the docs state a 900k-token request bills at the same per-token rate as a 9k-token one.
So the same million tokens sent to three vendors produce three different bills.

Three different numbers get mixed into the phrase “1M context”: (1) the input limit, (2) the output limit, and (3) the long-context pricing threshold. GPT-6 Astra, for example, carries 1,050,000 / 128,000 / 272K — three separate figures. This entry keeps them apart.

What a million tokens actually holds

“One million” is hard to picture, so here is what it looks like in familiar units. A token is not a character, so treat these as rough guidance only.

English text
~750,000 words
Using the simple 1 token ≈ 0.75 words convention: about 1,500 pages at 500 words a page, or seven to eight 100,000-word technical books
Japanese text
Fewer characters
Japanese tends to use more tokens per equivalent sentence than English, so fewer characters fit in the same 1M. Exact counts depend on the tokenizer — measure with the calculator
Audio (Gemini basis)
~11 hours
Gemini states 25 tokens per second of audio. 1,000,000 ÷ 25 = 40,000 seconds = about 11.1 hours (our arithmetic)
Images / PDFs (Claude)
up to 600
Anthropic states a single request can include up to 600 images or PDF pages (100 for 200k-token models)
Anthropic, verbatim: "A single request can include up to 600 images or PDF pages (100 for models with a 200k-token context window)."

Sources: Claude Platform Docs, “Context windows” and Gemini API pricing (audio token conversion), both checked 2026-09-21. The word-count conversion is ours. Converting characters to tokens varies by model and tokenizer, so always estimate from measured token counts — our calculator takes real token counts as input for exactly that reason.

The ceilings, as the vendors themselves state them

A 1M window is no longer one vendor's specialty. Here are the current ceilings, written the way each vendor writes them.

Provider / modelInput limitOutput limitLong-input billingSource (checked 2026-09-21)
OpenAI
GPT-6 Astra
1,050,000128,000Past 272K, the whole request goes to 2x input and 1.5x outputOpenAI model page
Anthropic
Claude Fable 5 / Sonnet 5 and other 1M models
1,000,000128,000Standard pricing — no long-context premiumContext windows / Pricing
Google
Gemini 3.1 Pro Preview
1,048,57665,536Past 200K, every token moves to the higher rateGemini model page / pricing
DeepSeek
deepseek-flash / deepseek-v4-pro
1M384K (max)No tiered long-context pricing stated on the official tableModels & Pricing
Anthropic's wording is the interesting one here. Their docs state that across 1M-token models, 1M is the default (no beta header needed) and long-context requests are billed at standard pricing. The pricing page goes further.
"Claude 4.6 and later models and Claude Mythos Preview include the full 1M token context window at standard pricing. (A 900k-token request is billed at the same per-token rate as a 9k-token request.) Prompt caching and batch processing discounts apply at standard rates across the full context window." In other words, a long request does not raise the per-token rate. Prompt caching and batch discounts also apply at standard rates across the whole window.

Sources: Claude Platform Docs, “Pricing” and “Context windows” (checked 2026-09-21). Anthropic's own list of 1M models reads: Claude Fable 5.1, Mythos 5.1, Fable 5, Mythos 5, Opus 5, Opus 4.8, Opus 4.7, Opus 4.6, Sonnet 5, Sonnet 4.6 and Mythos Preview — and other models such as Sonnet 4.5 have a 200k window, so check the official table before you build on it.

What crossing the threshold costs, in dollars

“The whole request at 2x” reads like a minor footnote. In money it looks like this. Every figure below is a plain tokens x rate calculation and excludes caching, batch discounts, cache writes and minimum charges. GPT-6 Astra bills standard input at $10 per 1M tokens, and $20 per 1M past 272K.

Input tokensRate appliedInput chargeChange from the row above
272,000Standard $10 / 1M$2.72—
272,100Long $20 / 1M$5.44+$2.72 (roughly 2x)
500,000Long $20 / 1M$10.00—
1,000,000Long $20 / 1M$20.00—
Overshooting the threshold by 100 tokens roughly doubles the input charge. A hundred tokens is a few hundred characters. One extra line in the prompt can double the bill — that is what a pricing step means. Where the same request also produces output, output carries 1.5x too ($50 to $75 per 1M), so the gap in the total bill widens further: input alone roughly doubles, but the total multiple depends on the input/output mix.

Google's table has the same shape: Gemini 3.1 Pro bills input at $2.00 per 1M up to 200K and $4.00 per 1M above it. 199,000 tokens costs $0.398; 201,000 tokens costs $0.804 (our arithmetic). Again, the whole prompt moves to the higher rate, not just the excess.
Google's footnote, verbatim: "If a query input context is longer than 200K tokens, all tokens (input and output) are charged at long context rates." Read alongside the tier table ($2.00 for prompts <= 200k tokens, $4.00 for prompts > 200k tokens) that confirms the design: it is not just the excess that moves, the whole request does. Output follows the same pattern ($12 rising to $18).

Rates from the OpenAI model page (verbatim: "Prompts with more than 272K input tokens are priced at 2x input and cache rates and 1.5x output for the full request.") and Gemini API pricing (checked 2026-09-21). The Google footnote is quoted verbatim from Google Cloud's Generative AI pricing page (checked 2026-09-21). Google's own pages vary between "longer than 200K" and "longer than or equal to 200K" depending on the surface, so confirm the rate table for the API and region you use. Dollar figures are our plain multiplication (tokens × rate) and exclude caching, cache writes and batch discounts. The 272K threshold is specific to GPT-6 Astra — it is not an OpenAI-wide number, so check each model's own page.

The second trap: input piles up as a conversation grows

You do not have to reach the ceiling for costs to rise. Chat APIs are usually stateless — the server does not remember the conversation — so every turn resends everything from the start.

SettingValue
Fixed prefix (system prompt + documents)8,000 tokens
Added per turn3,000 tokens
Turns10
Rate (GPT-6 Astra, no caching)$10 / 1M
Ten turns send 245,000 input tokens in total — $2.45. That is 8,000×10 + 3,000×(1+2+…+10) = 80,000 + 165,000 = 245,000. The genuinely new information was 8,000 + 3,000×10 = 38,000 tokens ($0.38). More than six times that amount was spent re-sending the same prefix.

And it scales with the square of the turn count. At 100 turns the total input sent is about 15.95 million tokens ($159.50) — without ever touching the 1M ceiling.

Two ways out: send less, and cache what you resend. Pass only the parts that matter, and put the unchanging prefix in a prompt cache. Where a fixed prefix is sent repeatedly, the cache is the single biggest lever on the bill (see entry 2 on prompt caching).

Dollar figures are our plain arithmetic (total input tokens × rate). Real bills also involve prompt caching, batch discounts and cache-write charges, so read these as upper-bound estimates. Separately, Anthropic documents context compaction (beta), which “automatically summarizes older context as conversations approach limits, increasing effective context length” (Anthropic announcement, checked 2026-09-21). Summarise-and-fold is a design direction several vendors are moving toward.

The third trap: “it fits” is not “it is used evenly”

Fill the window to the brim and will the model use all of it? Chroma's July 2025 technical report “Context Rot” evaluated 18 models (GPT-4.1, Claude 4, Gemini 2.5, Qwen3 and others) and found that performance gets less reliable as input grows.

"models do not use their context uniformly; instead, their performance grows increasingly unreliable as input length grows." A 1M context window is not a promise of 1M tokens of dependable performance.
ExperimentConditionResult (same report)
LongMemEval
(answering from a long chat history)
Full input of 113,000 tokens versus a focused prompt of about 300 tokens carrying only the relevant parts“Across all models, we see significantly higher performance on focused prompts” — adding irrelevant context adds a retrieval step and degrades reliability
DistractorsZero to four plausible-but-wrong statements placed near the needle“Even a single distractor reduces performance relative to the baseline, and adding four compounds this degradation further” — and each distractor differs in impact
Haystack structureLogically ordered text versus the same sentences randomly shuffledCounterintuitively, shuffled haystacks scored better: models are sensitive to the logical flow of context
Repeated wordsReproduce 25 to 10,000 repeated words with one unique word insertedPerformance degrades consistently across all models as length grows; position accuracy is highest when the unique word comes early
The report's conclusion is about presentation, not just presence. "Whether relevant information is present in a model's context is not all that matters; what matters more is how that information is presented." Which means “it fits, so include it” is not automatically a win for accuracy either.

Source: Chroma, “Context Rot: How Increasing Input Tokens Impacts LLM Performance” (checked 2026-09-21). That evaluation ran on models current in July 2025 (GPT-4.1, Claude 4, Gemini 2.5, Qwen3 generations). We have not verified whether the current September 2026 models show the same effect at the same strength — read it as a result from that period.

The picture: the ceiling, the price step and the accuracy ceiling are three different lines

Three numbers behind a 1M context window: the input ceiling (GPT-6 Astra 1,050,000 tokens), the long-context price step (past 272K the entire request is billed at 2x input), and the range where accuracy stays stable, which is narrower still. Crossing the price step makes the bill jump “1M” hides three different lines Example: GPT-6 Astra (input 1,050,000 / long-context step 272K / output 128,000) Standard rate Long-context rate (entire request: 2x input, 1.5x output) 272K step one token past this doubles the whole prompt 1,050,000 hard input ceiling Range where accuracy stays steadier performance grows less reliable as input grows (Chroma, July 2025, 18 models) At the same 113,000 tokens, focused prompts scored significantly higher than full ones. A single distractor already lowers performance; four compound the damage. Sources: OpenAI model page / ai.google.dev pricing / Anthropic Context windows / Chroma Context Rot (checked 2026-09-21) Figures are vendor-published values as we read them. The accuracy band is conceptual, not a measured threshold.
The top band is billing; the bottom band is a conceptual sketch of the accuracy trend, not a safe ceiling. Fitting inside the window does not mean getting even quality. Thresholds differ per model, so check each vendor's page.

Everyday version: a huge desk and a shipping charge

A 1M context window is like being handed an enormous desk. But the more you pile on it, the longer it takes to find the page you need. And the awkward part is that the size of the desk is tied to the shipping fee. Past a certain weight, it is not the excess that gets charged more — the whole shipment goes up. That is how OpenAI and Google structure long-context pricing. Anthropic's approach is closer to “a bigger desk, same shipping fee.” Same word, different design.
In everyday termsIn context-window terms
The size of the deskThe input limit (1M tokens and so on)
Total weight you can pile onSystem prompt + history + documents + tool output combined
The weight where shipping jumpsThe long-context threshold (272K, 200K and so on)
A huge desk where you still cannot find the pagePerformance degrading with longer input (context rot)
Bringing every box to every meetingStateless APIs re-billing the entire input on every turn

One takeaway: a bigger desk is not automatically a bargain. How much you pile on, how you arrange it, and where the price step sits are the three things worth checking together.

How to use this at work

Three checks before you send a long prompt

  • Where is the step, and does it apply to the whole request or just the excess? — GPT-6 Astra: past 272K, the whole request doubles on input and goes 1.5x on output. Gemini: past 200K, all tokens move to the higher rate. Anthropic: standard pricing to 1M. Getting this wrong can distort an estimate by more than 2x.
  • Do you need to send all of it? — In the LongMemEval comparison, a 113,000-token full prompt was clearly beaten by a roughly 300-token focused one. Retrieving first and sending less tends to be both cheaper and more accurate.
  • Is the repeated prefix cached? — Of $2.45 spent over ten turns, most was re-sending the same prefix. Put the unchanging part in a prompt cache.
  • Do not choose a provider on the ceiling number alone. The same “1M” comes with different long-context rules, different output limits and different cache treatment. Read this alongside the pricing comparison.
  • Test it on your own task. Needle in a Haystack — hiding one fact in a long document — is a retrieval test, not a test of interpreting ambiguous instructions. A high score there does not transfer automatically to your workload.
  • Plan to fold, not to fill. Anthropic documents context compaction, which summarises older context as a conversation approaches its limit. Designing for summarise-and-drop is often more stable than designing for a permanently full window.
  • Always read a context figure next to a rate card. Every dollar figure here is our own arithmetic; caching, batch discounts and cache writes change the real number.

In one sentence

A 1M context window is a capacity claim — price and accuracy are separate stories

  • 1M tokens is roughly 750,000 English words, or about 11 hours of audio. Japanese typically needs more tokens for the same sentence.
  • Long-input billing differs sharply. OpenAI doubles the whole request past 272K; Google moves every token to the higher rate past 200K; Anthropic stays on standard pricing at 1M.
  • Crossing the step makes the bill jump: 272,000 tokens $2.72 versus 272,100 tokens $5.44 (GPT-6 Astra, our arithmetic).
  • Costs grow even before you reach the ceiling. Ten turns send 245,000 input tokens while adding only 38,000 tokens of new information.
  • And fitting is not the same as being used evenly — published research found performance degrading as input length grows.

The easiest way to misread this term is to let the headline number travel on its own. What actually changes your outcome is the price step just below the ceiling, and the accuracy ceiling well below that. Next time you see “supports a 1M context window,” look for the rate step and the cache rules in the same paragraph.

Next

The rate difference when you resend a prefix is in entry 2 (prompt caching), and how to read benchmark numbers is in entry 1. Current rates are in the pricing comparison.

🧪 Back to the glossary →

Glossary · Prompt caching · Reading benchmark numbers · Pricing · Token calculator

⚠️ Disclaimer

  • Rates, limits and quotations were checked against primary sources (OpenAI, Google, Anthropic and DeepSeek official pages) on 2026-09-21. Pricing and specifications change without notice — always confirm on the vendor's page before you commit.
  • Every dollar figure is our own plain arithmetic (tokens × rate) and excludes caching, batch discounts, cache writes and minimum charges. Real invoices will differ.
  • The finding that performance degrades with longer input comes from Chroma's July 2025 report (18 models; the GPT-4.1, Claude 4, Gemini 2.5 and Qwen3 generations). We have not verified whether current models show the same effect at the same strength.
  • Thresholds such as 272K are model-specific. Do not apply the figures here to other models.
  • Vendor wording is quoted verbatim; the surrounding explanation is ours. Where the two differ, the vendor's wording governs.
  • We are not affiliated with any provider, and this is not a recommendation of any product.