What is prompt caching? Reading “cache hit” in dollars
Glossary, entry 2. We take common terms in pricing tables and vendor announcements and walk them back to primary sources. This one is prompt caching, starting from a single line on DeepSeek's pricing page. The same input, billed at 1/50 the price — could you explain why?
“Input (cache hit) $0.003 per 1M tokens”
Three terms, and you are done
Every pricing table reduces to these three. Once they are separated, the rest is arithmetic.
How it works: only the matching prefix is reused
- Prefix matching is the shared idea across all four providers. OpenAI describes avoiding recomputation of a prompt prefix; Anthropic states that caching covers the full prompt prefix (tools, system, messages); Gemini's docs advise putting large, common content at the beginning of the prompt and sending similar prefixes close together in time.
- Everything after the prefix stays flexible. Fix the prefix (system instructions, reference documents, tool definitions) and put the variable part last. That single change moves your hit rate.
- You can verify hits from the response. OpenAI returns
usage.input_tokens_details.cached_tokensandcache_write_tokens; Anthropic returnscache_read_input_tokensandcache_creation_input_tokens; Gemini returnsusage.total_cached_tokens. Measure first, optimise second.
Rules per provider (primary sources)
Write charges, lifetimes and the minimum cacheable length all differ. These are the figures exactly as the official documentation states them.
| Provider | Write charge | Read (cache hit) | Cache lifetime (TTL) | Minimum cacheable |
|---|---|---|---|---|
| OpenAI (GPT-5.6 and later) | 1.25x the standard input rate | 0.1x — up to 90% off | 30 minutes (prompt_cache_options.ttl supports only "30m", which is also the default) | — (model-dependent, automatic) |
| Anthropic (Claude) | 5-minute TTL: 1.25x 1-hour TTL: 2x | 0.1x Except Fable 5.1 / Mythos 5.1: 0.025x | 5 minutes by default (refreshed free on each use); 1 hour costs extra | 512 tokens (Fable 5.1, Mythos 5.1, Opus 5) / 2,048 (others) / 4,096 (Haiku 4.5) |
| Google (Gemini) | Explicit caching adds a separate storage charge | 10% of input | Implicit caching is automatic; explicit caches take a TTL | 4,096 (3.x Flash family, 3.1 Pro Preview) / 2,048 (2.5 family) |
| DeepSeek | No write charge documented (the first send bills as normal input) | 1/50 of a cache miss | Automatic (on by default, no code change) | — (prefix units, automatic) |
Sources: OpenAI prompt caching ・ Anthropic prompt caching (multipliers and minimum lengths) ・ Gemini context caching (implicit caching on by default for Gemini 2.5 and newer; minimum token counts) ・ DeepSeek context caching (all checked 2026-09-17). Anthropic's headline multipliers are 5-minute writes at 1.25x, 1-hour writes at 2x, and reads (hits and refreshes) at 0.1x; Fable 5.1 and Mythos 5.1 are documented at 0.025x for reads.
Real rates (USD per 1M tokens)
| Model | Standard input | cache hit | cache write |
|---|---|---|---|
| GPT-6 Astra | $10.00 | $1.00 | $12.50 |
| Claude Fable 5.1 | $10.00 | $0.25 | $12.50 (5m) / $20 (1h) |
| Gemini 3.8 Flash (through 12/31) | $0.75 | $0.075 | storage $0.50 / 1M / hour |
| DeepSeek V4.1 Flash (off-peak) | $0.15 | $0.003 | none (first send bills as input) |
| Grok 4.6 | $2.00 | $0.50 | — |
| GLM-5.3-Flash | $0.15 | $0.03 | storage free for a limited time |
| MAI-Thinking-1 (Azure Global) | $2.00 | $0.20 | — |
Values come from each provider's official pricing page (verified on this site; full source URLs are on the pricing comparison). Gemini rates shown are the introductory rates available through December 31, 2026; from January 1, 2027 they rise to $1.50 / $0.15.
What 10 requests actually cost
Assumptions: an 8,000-token shared prefix, sent in the same order across 10 requests. Output ignored. The write happens once, on request 1; requests 2–10 all hit the cache.
| Model | No cache | With cache | Difference |
|---|---|---|---|
| GPT-6 Astra ($10 / $1.00 / $12.50) | $0.80 | $0.172 | ~78% less |
| Claude Fable 5.1 ($10 / $0.25 / $12.50) | $0.80 | $0.118 | ~85% less |
| DeepSeek V4.1 Flash ($0.15 / $0.003, off-peak) | $0.012 | $0.0014 | ~88% less |
- The arithmetic (GPT-6 Astra): without caching, 8,000 x 10 = 80,000 tokens at $10 per 1M = $0.80. With caching, the first write costs 8,000 x $12.50 per 1M = $0.10, and the nine reads cost 72,000 x $1.00 per 1M = $0.072 — $0.172 in total.
- The 1.25x write premium is earned back on the second request. With two or more reuses, the write premium (0.25x) is more than offset by the read discount (0.9x). Conversely, if you send it once, the write is pure loss at $12.50 per 1M.
- A bigger percentage cut does not mean a bigger saving. DeepSeek cuts 88% and still totals $0.0014; GPT-6 Astra cuts 78% and totals $0.172 — over 100x more. Compare totals, not discount rates.
- Output is never discounted. Only input can be cached, so on output-heavy workloads the real-world saving is smaller than the table suggests. For totals, use the token calculator (it computes standard rates without caching).
Rates are the published figures from each provider; the arithmetic is ours. Your actual bill depends on hit rates, TTL expiry and output tokens.
Four common misconceptions
Summary: three things to check in a price table
When you see “cache hit”
- The multiplier vs standard input — 0.1x, 0.025x or 1/50. That is the value of caching.
- The write multiplier and the TTL — 1.25x or 2x; 5 minutes, 30 minutes or 1 hour. That decides whether it pays off.
- The minimum prompt length — 512 to 4,096 tokens. Below it, the saving is zero.
Prompt caching rewards architectures that resend the same prefix. Put differently, how you assemble the prompt visibly moves the bill. Agentic loops that resend the same system prompt, tool definitions and reference documents every turn benefit the most.
Next
See every rate in the pricing comparison, run your own numbers in the token calculator, or check worked examples on DeepSeek V4.1 Flash and GPT-6 Astra.
📊 Open the pricing tables →Glossary ・ How to read benchmark numbers ・ Token calculator ・ Release timeline ・ MS AI (MAI)
⚠️ Disclaimer
- Prices and specifications were checked against each provider's official pages and documentation on 2026-09-17. They change without notice — always confirm on the vendor's site before you commit.
- The cost examples are simple arithmetic on published rates and do not guarantee your actual bill. Real cache hit rates depend on your workload.
- Caching behaviour, defaults and multipliers can change; check each provider's current documentation before implementing.