What is prompt caching? Reading “cache hit” in dollars

Glossary, entry 2. We take common terms in pricing tables and vendor announcements and walk them back to primary sources. This one is prompt caching, starting from a single line on DeepSeek's pricing page. The same input, billed at 1/50 the price — could you explain why?

“Input (cache hit) $0.003 per 1M tokens”

Source: DeepSeek's official pricing page (V4.1 Flash, model id deepseek-flash, checked 2026-09-17) — the same table lists input (cache miss) at $0.15 off-peak per 1M. So a cache hit costs 1/50 of a cache miss. See the full rate card.

● Last updated:  |  Primary sources: OpenAI prompt caching ・ Anthropic prompt caching ・ Gemini context caching ・ DeepSeek context caching  |  Checked: 2026-09-17

Three terms, and you are done

Every pricing table reduces to these three. Once they are separated, the rest is arithmetic.

cache miss
standard input rate
The input was not in the cache. A first request is always a miss
cache hit (cached input)
discounted rate
The same prefix was still cached. This is the cheap side
cache write
1.25x to 2x
What you pay to store the prefix. This is the expensive side
Bottom line: prompt caching is not a discount, it is a prepayment. You pay a premium to write the cache once and earn it back on later reads. That is why caching a prompt you send only once costs you money. This asymmetry is where most cost models go wrong.

How it works: only the matching prefix is reused

Bottom line: caching only applies to the part of the prompt that matches from the very beginning. The provider walks your input tokens from the start and reuses the range that is identical to the previous request instead of recomputing it. Put a dynamic value first — a timestamp, a user ID, today's question — and the match breaks there, so every request is a miss.
How prompt caching works: the shared prefix is written once at 1.25x, read back at 0.1x, turning $0.80 into $0.172 across 10 requests How prompt caching works Same prefix every time, cheaper from the second call. Example: GPT-6 Astra ($10 input / $1.00 cache hit / $12.50 cache write per 1M) Prompt structure Shared prefix (system prompt, docs) 8,000 tokens This request's question Request 1 (nothing cached) cache write 1.25x $12.50 / 1M Requests 2+ (prefix matches) cache hit 0.1x $1.00 / 1M Input cost over 10 requests (8,000-token prefix sent every time) No cache $0.80 With cache $0.172 (~78% less)
Only the shared prefix is cached. It is written at 1.25x on the first request and read back at 0.1x afterwards (rates from OpenAI's GPT-6 Astra pricing).
  • Prefix matching is the shared idea across all four providers. OpenAI describes avoiding recomputation of a prompt prefix; Anthropic states that caching covers the full prompt prefix (tools, system, messages); Gemini's docs advise putting large, common content at the beginning of the prompt and sending similar prefixes close together in time.
  • Everything after the prefix stays flexible. Fix the prefix (system instructions, reference documents, tool definitions) and put the variable part last. That single change moves your hit rate.
  • You can verify hits from the response. OpenAI returns usage.input_tokens_details.cached_tokens and cache_write_tokens; Anthropic returns cache_read_input_tokens and cache_creation_input_tokens; Gemini returns usage.total_cached_tokens. Measure first, optimise second.

Rules per provider (primary sources)

Write charges, lifetimes and the minimum cacheable length all differ. These are the figures exactly as the official documentation states them.

ProviderWrite chargeRead (cache hit)Cache lifetime (TTL)Minimum cacheable
OpenAI
(GPT-5.6 and later)
1.25x the standard input rate0.1x — up to 90% off30 minutes (prompt_cache_options.ttl supports only "30m", which is also the default)— (model-dependent, automatic)
Anthropic
(Claude)
5-minute TTL: 1.25x
1-hour TTL: 2x
0.1x
Except Fable 5.1 / Mythos 5.1: 0.025x
5 minutes by default (refreshed free on each use); 1 hour costs extra512 tokens (Fable 5.1, Mythos 5.1, Opus 5) / 2,048 (others) / 4,096 (Haiku 4.5)
Google
(Gemini)
Explicit caching adds a separate storage charge10% of inputImplicit caching is automatic; explicit caches take a TTL4,096 (3.x Flash family, 3.1 Pro Preview) / 2,048 (2.5 family)
DeepSeekNo write charge documented (the first send bills as normal input)1/50 of a cache missAutomatic (on by default, no code change)— (prefix units, automatic)

Sources: OpenAI prompt caching ・ Anthropic prompt caching (multipliers and minimum lengths) ・ Gemini context caching (implicit caching on by default for Gemini 2.5 and newer; minimum token counts) ・ DeepSeek context caching (all checked 2026-09-17). Anthropic's headline multipliers are 5-minute writes at 1.25x, 1-hour writes at 2x, and reads (hits and refreshes) at 0.1x; Fable 5.1 and Mythos 5.1 are documented at 0.025x for reads.

Real rates (USD per 1M tokens)

ModelStandard inputcache hitcache write
GPT-6 Astra$10.00$1.00$12.50
Claude Fable 5.1$10.00$0.25$12.50 (5m) / $20 (1h)
Gemini 3.8 Flash (through 12/31)$0.75$0.075storage $0.50 / 1M / hour
DeepSeek V4.1 Flash (off-peak)$0.15$0.003none (first send bills as input)
Grok 4.6$2.00$0.50—
GLM-5.3-Flash$0.15$0.03storage free for a limited time
MAI-Thinking-1 (Azure Global)$2.00$0.20—

Values come from each provider's official pricing page (verified on this site; full source URLs are on the pricing comparison). Gemini rates shown are the introductory rates available through December 31, 2026; from January 1, 2027 they rise to $1.50 / $0.15.

What 10 requests actually cost

Assumptions: an 8,000-token shared prefix, sent in the same order across 10 requests. Output ignored. The write happens once, on request 1; requests 2–10 all hit the cache.

ModelNo cacheWith cacheDifference
GPT-6 Astra ($10 / $1.00 / $12.50)$0.80$0.172~78% less
Claude Fable 5.1 ($10 / $0.25 / $12.50)$0.80$0.118~85% less
DeepSeek V4.1 Flash ($0.15 / $0.003, off-peak)$0.012$0.0014~88% less
  • The arithmetic (GPT-6 Astra): without caching, 8,000 x 10 = 80,000 tokens at $10 per 1M = $0.80. With caching, the first write costs 8,000 x $12.50 per 1M = $0.10, and the nine reads cost 72,000 x $1.00 per 1M = $0.072 — $0.172 in total.
  • The 1.25x write premium is earned back on the second request. With two or more reuses, the write premium (0.25x) is more than offset by the read discount (0.9x). Conversely, if you send it once, the write is pure loss at $12.50 per 1M.
  • A bigger percentage cut does not mean a bigger saving. DeepSeek cuts 88% and still totals $0.0014; GPT-6 Astra cuts 78% and totals $0.172 — over 100x more. Compare totals, not discount rates.
  • Output is never discounted. Only input can be cached, so on output-heavy workloads the real-world saving is smaller than the table suggests. For totals, use the token calculator (it computes standard rates without caching).

Rates are the published figures from each provider; the arithmetic is ours. Your actual bill depends on hit rates, TTL expiry and output tokens.

Four common misconceptions

1. “It is automatic, so it will just get cheaper.” It is automatic — OpenAI enables it by default, Gemini from 2.5 onward, DeepSeek for all users. But whether you get a hit depends entirely on your prefix matching. Put a timestamp or a random value at the front and you will miss on every request, automatic or not.
2. “Writes are free; only reads are discounted.” The reverse. Writes cost 1.25x (2x on Anthropic's 1-hour TTL) while reads cost 0.1x. OpenAI's own documentation frames the trade-off: pay the write when you know the prefix will be reused. A one-shot giant context remains the most expensive way to call a model.
3. “Any prompt can be cached.” Below the minimum length, nothing is cached at all. Anthropic documents 512 / 2,048 / 4,096 tokens depending on model; Gemini 3.x Flash requires 4,096. Reusing a short instruction block generates no cache hits.
4. “A cache hit cuts my bill by the same percentage.” Only the cached part of the input is discounted; output rates do not move. And once the TTL expires the next request is a miss again (5 minutes by default on Anthropic, 30 minutes on OpenAI GPT-5.6 and later). Design with the TTL and your request interval in mind, together.

Summary: three things to check in a price table

When you see “cache hit”

  • The multiplier vs standard input — 0.1x, 0.025x or 1/50. That is the value of caching.
  • The write multiplier and the TTL — 1.25x or 2x; 5 minutes, 30 minutes or 1 hour. That decides whether it pays off.
  • The minimum prompt length — 512 to 4,096 tokens. Below it, the saving is zero.

Prompt caching rewards architectures that resend the same prefix. Put differently, how you assemble the prompt visibly moves the bill. Agentic loops that resend the same system prompt, tool definitions and reference documents every turn benefit the most.

Next

See every rate in the pricing comparison, run your own numbers in the token calculator, or check worked examples on DeepSeek V4.1 Flash and GPT-6 Astra.

📊 Open the pricing tables →

Glossary ・ How to read benchmark numbers ・ Token calculator ・ Release timeline ・ MS AI (MAI)

⚠️ Disclaimer

  • Prices and specifications were checked against each provider's official pages and documentation on 2026-09-17. They change without notice — always confirm on the vendor's site before you commit.
  • The cost examples are simple arithmetic on published rates and do not guarantee your actual bill. Real cache hit rates depend on your workload.
  • Caching behaviour, defaults and multipliers can change; check each provider's current documentation before implementing.