What is a 1M-token context window? The limit, the price step, and the accuracy ceiling
Glossary entry 4. This one covers the 1M-token context window. “Up to one million tokens” quietly mixes up two different claims: how much actually fits (capacity), and what it costs once you put it there (billing). Billing turns out to work very differently at OpenAI, Google and Anthropic: OpenAI and Google raise the rate across the whole request once you cross the line, while Anthropic stays on standard pricing for its 1M models. And a bigger window does not mean the model reads all of it equally well. We keep the numbers and the sources in, so you can work out what it means for your own workload.
“Supports 1M” does not mean “uses 1M well, at the same price”
Short version: the ceiling is capacity, the bill is a separate rule
- OpenAI: above 272K tokens, the whole request is billed at 2x input and 1.5x output — not just the excess.
- Google: above 200K tokens, all tokens are charged at the long-context rate.
- Anthropic: standard pricing all the way to 1M — the docs state a 900k-token request bills at the same per-token rate as a 9k-token one.
Three different numbers get mixed into the phrase “1M context”: (1) the input limit, (2) the output limit, and (3) the long-context pricing threshold. GPT-6 Astra, for example, carries 1,050,000 / 128,000 / 272K — three separate figures. This entry keeps them apart.
What a million tokens actually holds
“One million” is hard to picture, so here is what it looks like in familiar units. A token is not a character, so treat these as rough guidance only.
Sources: Claude Platform Docs, “Context windows” and Gemini API pricing (audio token conversion), both checked 2026-09-21. The word-count conversion is ours. Converting characters to tokens varies by model and tokenizer, so always estimate from measured token counts — our calculator takes real token counts as input for exactly that reason.
The ceilings, as the vendors themselves state them
A 1M window is no longer one vendor's specialty. Here are the current ceilings, written the way each vendor writes them.
| Provider / model | Input limit | Output limit | Long-input billing | Source (checked 2026-09-21) |
|---|---|---|---|---|
| OpenAI GPT-6 Astra | 1,050,000 | 128,000 | Past 272K, the whole request goes to 2x input and 1.5x output | OpenAI model page |
| Anthropic Claude Fable 5 / Sonnet 5 and other 1M models | 1,000,000 | 128,000 | Standard pricing — no long-context premium | Context windows / Pricing |
| Google Gemini 3.1 Pro Preview | 1,048,576 | 65,536 | Past 200K, every token moves to the higher rate | Gemini model page / pricing |
| DeepSeek deepseek-flash / deepseek-v4-pro | 1M | 384K (max) | No tiered long-context pricing stated on the official table | Models & Pricing |
Sources: Claude Platform Docs, “Pricing” and “Context windows” (checked 2026-09-21). Anthropic's own list of 1M models reads: Claude Fable 5.1, Mythos 5.1, Fable 5, Mythos 5, Opus 5, Opus 4.8, Opus 4.7, Opus 4.6, Sonnet 5, Sonnet 4.6 and Mythos Preview — and other models such as Sonnet 4.5 have a 200k window, so check the official table before you build on it.
What crossing the threshold costs, in dollars
“The whole request at 2x” reads like a minor footnote. In money it looks like this. Every figure below is a plain tokens x rate calculation and excludes caching, batch discounts, cache writes and minimum charges. GPT-6 Astra bills standard input at $10 per 1M tokens, and $20 per 1M past 272K.
| Input tokens | Rate applied | Input charge | Change from the row above |
|---|---|---|---|
| 272,000 | Standard $10 / 1M | $2.72 | — |
| 272,100 | Long $20 / 1M | $5.44 | +$2.72 (roughly 2x) |
| 500,000 | Long $20 / 1M | $10.00 | — |
| 1,000,000 | Long $20 / 1M | $20.00 | — |
Google's table has the same shape: Gemini 3.1 Pro bills input at $2.00 per 1M up to 200K and $4.00 per 1M above it. 199,000 tokens costs $0.398; 201,000 tokens costs $0.804 (our arithmetic). Again, the whole prompt moves to the higher rate, not just the excess.
Rates from the OpenAI model page (verbatim: "Prompts with more than 272K input tokens are priced at 2x input and cache rates and 1.5x output for the full request.") and Gemini API pricing (checked 2026-09-21). The Google footnote is quoted verbatim from Google Cloud's Generative AI pricing page (checked 2026-09-21). Google's own pages vary between "longer than 200K" and "longer than or equal to 200K" depending on the surface, so confirm the rate table for the API and region you use. Dollar figures are our plain multiplication (tokens × rate) and exclude caching, cache writes and batch discounts. The 272K threshold is specific to GPT-6 Astra — it is not an OpenAI-wide number, so check each model's own page.
The second trap: input piles up as a conversation grows
You do not have to reach the ceiling for costs to rise. Chat APIs are usually stateless — the server does not remember the conversation — so every turn resends everything from the start.
| Setting | Value |
|---|---|
| Fixed prefix (system prompt + documents) | 8,000 tokens |
| Added per turn | 3,000 tokens |
| Turns | 10 |
| Rate (GPT-6 Astra, no caching) | $10 / 1M |
And it scales with the square of the turn count. At 100 turns the total input sent is about 15.95 million tokens ($159.50) — without ever touching the 1M ceiling.
Two ways out: send less, and cache what you resend. Pass only the parts that matter, and put the unchanging prefix in a prompt cache. Where a fixed prefix is sent repeatedly, the cache is the single biggest lever on the bill (see entry 2 on prompt caching).
Dollar figures are our plain arithmetic (total input tokens × rate). Real bills also involve prompt caching, batch discounts and cache-write charges, so read these as upper-bound estimates. Separately, Anthropic documents context compaction (beta), which “automatically summarizes older context as conversations approach limits, increasing effective context length” (Anthropic announcement, checked 2026-09-21). Summarise-and-fold is a design direction several vendors are moving toward.
The third trap: “it fits” is not “it is used evenly”
Fill the window to the brim and will the model use all of it? Chroma's July 2025 technical report “Context Rot” evaluated 18 models (GPT-4.1, Claude 4, Gemini 2.5, Qwen3 and others) and found that performance gets less reliable as input grows.
| Experiment | Condition | Result (same report) |
|---|---|---|
| LongMemEval (answering from a long chat history) | Full input of 113,000 tokens versus a focused prompt of about 300 tokens carrying only the relevant parts | “Across all models, we see significantly higher performance on focused prompts” — adding irrelevant context adds a retrieval step and degrades reliability |
| Distractors | Zero to four plausible-but-wrong statements placed near the needle | “Even a single distractor reduces performance relative to the baseline, and adding four compounds this degradation further” — and each distractor differs in impact |
| Haystack structure | Logically ordered text versus the same sentences randomly shuffled | Counterintuitively, shuffled haystacks scored better: models are sensitive to the logical flow of context |
| Repeated words | Reproduce 25 to 10,000 repeated words with one unique word inserted | Performance degrades consistently across all models as length grows; position accuracy is highest when the unique word comes early |
Source: Chroma, “Context Rot: How Increasing Input Tokens Impacts LLM Performance” (checked 2026-09-21). That evaluation ran on models current in July 2025 (GPT-4.1, Claude 4, Gemini 2.5, Qwen3 generations). We have not verified whether the current September 2026 models show the same effect at the same strength — read it as a result from that period.
The picture: the ceiling, the price step and the accuracy ceiling are three different lines
Everyday version: a huge desk and a shipping charge
| In everyday terms | In context-window terms |
|---|---|
| The size of the desk | The input limit (1M tokens and so on) |
| Total weight you can pile on | System prompt + history + documents + tool output combined |
| The weight where shipping jumps | The long-context threshold (272K, 200K and so on) |
| A huge desk where you still cannot find the page | Performance degrading with longer input (context rot) |
| Bringing every box to every meeting | Stateless APIs re-billing the entire input on every turn |
One takeaway: a bigger desk is not automatically a bargain. How much you pile on, how you arrange it, and where the price step sits are the three things worth checking together.
How to use this at work
Three checks before you send a long prompt
- Where is the step, and does it apply to the whole request or just the excess? — GPT-6 Astra: past 272K, the whole request doubles on input and goes 1.5x on output. Gemini: past 200K, all tokens move to the higher rate. Anthropic: standard pricing to 1M. Getting this wrong can distort an estimate by more than 2x.
- Do you need to send all of it? — In the LongMemEval comparison, a 113,000-token full prompt was clearly beaten by a roughly 300-token focused one. Retrieving first and sending less tends to be both cheaper and more accurate.
- Is the repeated prefix cached? — Of $2.45 spent over ten turns, most was re-sending the same prefix. Put the unchanging part in a prompt cache.
- Do not choose a provider on the ceiling number alone. The same “1M” comes with different long-context rules, different output limits and different cache treatment. Read this alongside the pricing comparison.
- Test it on your own task. Needle in a Haystack — hiding one fact in a long document — is a retrieval test, not a test of interpreting ambiguous instructions. A high score there does not transfer automatically to your workload.
- Plan to fold, not to fill. Anthropic documents context compaction, which summarises older context as a conversation approaches its limit. Designing for summarise-and-drop is often more stable than designing for a permanently full window.
- Always read a context figure next to a rate card. Every dollar figure here is our own arithmetic; caching, batch discounts and cache writes change the real number.
In one sentence
A 1M context window is a capacity claim — price and accuracy are separate stories
- 1M tokens is roughly 750,000 English words, or about 11 hours of audio. Japanese typically needs more tokens for the same sentence.
- Long-input billing differs sharply. OpenAI doubles the whole request past 272K; Google moves every token to the higher rate past 200K; Anthropic stays on standard pricing at 1M.
- Crossing the step makes the bill jump: 272,000 tokens $2.72 versus 272,100 tokens $5.44 (GPT-6 Astra, our arithmetic).
- Costs grow even before you reach the ceiling. Ten turns send 245,000 input tokens while adding only 38,000 tokens of new information.
- And fitting is not the same as being used evenly — published research found performance degrading as input length grows.
The easiest way to misread this term is to let the headline number travel on its own. What actually changes your outcome is the price step just below the ceiling, and the accuracy ceiling well below that. Next time you see “supports a 1M context window,” look for the rate step and the cache rules in the same paragraph.
Next
The rate difference when you resend a prefix is in entry 2 (prompt caching), and how to read benchmark numbers is in entry 1. Current rates are in the pricing comparison.
🧪 Back to the glossary →Glossary · Prompt caching · Reading benchmark numbers · Pricing · Token calculator
⚠️ Disclaimer
- Rates, limits and quotations were checked against primary sources (OpenAI, Google, Anthropic and DeepSeek official pages) on 2026-09-21. Pricing and specifications change without notice — always confirm on the vendor's page before you commit.
- Every dollar figure is our own plain arithmetic (tokens × rate) and excludes caching, batch discounts, cache writes and minimum charges. Real invoices will differ.
- The finding that performance degrades with longer input comes from Chroma's July 2025 report (18 models; the GPT-4.1, Claude 4, Gemini 2.5 and Qwen3 generations). We have not verified whether current models show the same effect at the same strength.
- Thresholds such as 272K are model-specific. Do not apply the figures here to other models.
- Vendor wording is quoted verbatim; the surrounding explanation is ours. Where the two differ, the vendor's wording governs.
- We are not affiliated with any provider, and this is not a recommendation of any product.