One term per page. We do not stop at definitions — we go into the rates, conditions and boundaries behind them.
🧭
What is a benchmark? Sorting them by who, what and how
DeepSWE, GDPval-AA, Terminal-bench, OSWorld — too many to make sense of. So instead of memorising names, this entry hands you a procedure built on three questions: who operates it, what it measures, and how it measures it (harness and settings). Six benchmarks are traced to their operator's own repository or paper, including why GDPval and GDPval-AA share a dataset yet are not the same number, and the name collision between the DeepSWE benchmark and the same-named 32B model.
Read the entry →
📏
What is a 1M-token context window? What fits, and what it costs
A million tokens is roughly 750,000 English words, or about 11 hours of audio — but the three big vendors bill long input very differently. OpenAI doubles the entire request past 272K (2x input, 1.5x output); Google moves every token to the higher rate past 200K; Anthropic stays at standard pricing all the way to 1M. Includes the arithmetic (272,000 tokens $2.72 versus 272,100 tokens $5.44), a ten-turn example where 245,000 input tokens carry only 38,000 tokens of new information, and Chroma's July 2025 “Context Rot” findings across 18 models.
Read the entry →
🧰
What is an agent harness? Why the same model scores differently
A harness is everything that is not the model: system prompt, tools, memory, sandbox, retries and scoring. That is why the same model on the same benchmark can score differently. On ARC-AGI-3 it was 62.7% under ARC Prize's neutral harness versus 99.9% under OpenAI's Provider Adapter harness. Includes DeepSeek Harness's own line “Agent = Model + Harness” and its Minimal mode, which the vendor describes as being for benchmarking models in a minimal environment — all with sources.
Read the entry →
🗄️
What is prompt caching? Reading “cache hit” in dollars
The same input can cost 50x less as a cache hit. How prefix matching works, why writes bill at 1.25x while reads bill at 0.1x, the TTLs (5 to 30 minutes) and the minimum cacheable lengths (512 to 4,096 tokens) across OpenAI, Anthropic, Google and DeepSeek — plus what 8,000 tokens sent 10 times really costs ($0.80 versus $0.172).
Read the entry →
🧪
GPT-6 Astra: how to read "ARC-AGI-3 99.9%, ExploitBench 100%"
The same benchmark scores 62.7% under ARC Prize's neutral harness. A perfect score means the test can no longer separate models. Why the unsaturated Terminal-Bench Science 64.6% carries the most information, plus the 272K-token boundary of the 1.05M context window (past it, the entire request is billed at 2x input and 1.5x output) and what a full 128K output actually costs.
Read the entry →
🏛️
Who operates the benchmarks — a four-step way to tell
DeepSWE is Datacurve, OSWorld is XLANG Lab, GDPval is OpenAI, GDPval-AA is Artificial Analysis, AutomationBench is Zapier — this entry fixes a four-step procedure for getting from a benchmark name to the organisation behind it: (1) the paper or official repository, (2) how to classify the operator (vendor, independent evaluator, research lab, business-software company), (3) who reported the score, (4) whose harness ran it. Includes a table of twelve benchmarks and three worked examples of why “a third party operates it” does not mean “fair”.
Read the entry →
🔒
Zero data retention (ZDR) — three catch points
ZDR means “not stored by default”, not “always on”. Three things trip people up: ① there is an exception list of models that require retention, ② a guarantee needs an explicit setting, and ③ “not stored”, ZOA (operators cannot see it) and “not used for training” are three separate promises. Includes Bedrock's retention modes (none / default / inherit / aws_review) and their scope, the published exception lists (OpenAI models and Claude Fable 5 / 5.1), and whether OpenAI, Anthropic, Google and Microsoft give you ZDR by default or only by approval — quoted from their own docs. Glossary #7.
Read the entry →
🛡️
New The four data-protection requirements — "not stored", "not seen", "not learned from", "not sent abroad"
"ZDR supported" is not the end of the story. Telling conditions of an audit: ① not stored (ZDR) ② not seen (ZOA) ③ not learned from (training terms) ④ not sent abroad (data residency) are four separate promises, and almost no provider satisfies all four by default. Across AWS Bedrock, Microsoft Foundry, OpenAI, Anthropic and Google Cloud, this entry lines up which requirement is the default and which needs an application or approval, quoted from their own documentation, then adds a five-step pre-contract check and four myths. Glossary entry 8.
Read the entry →
🔐
New What is zero operator access (ZOA)? — "not retained" and "not viewable" are different promises
If ZDR is about storage, ZOA is about people: the three layers that block access (process, keys, runtime), the three stages of key management (CMK / BYOK / HYOK), and the approval workflows (Microsoft Customer Lockbox, Google Access Approval with Key Access Justifications, OpenAI EKM, Anthropic's controlled access path) — all quoted from official AWS, Microsoft, Google Cloud, OpenAI and Anthropic documentation, including the emergency exceptions. Glossary #9.
Read the entry →
🚫
New No-training explained — why one vendor gives two answers
"ChatGPT trains on your chats but the API does not" is true, and the reason is the layer, not the vendor: consumer app, free tier, or API/commercial contract. The same Google treats free-tier Gemini API traffic as "used to improve products, with human review" while paid quota is not used to improve products at all. Quoted from OpenAI, Anthropic, Google, Microsoft and AWS, including the two exceptions that survive an opt-out (content you send as feedback, and content flagged for safety review) and why a free tier is decided by the billing account. Glossary entry 10.
Read the entry →
🌍
New Data residency explained — “not sent abroad” is two promises
Keeping data in country takes two separate settings: where it is stored at rest, and where inference runs. OpenAI publishes Storage and Processing separately per region — Japan is storage yes / processing no. AWS decides with inference profiles, Google Cloud with endpoints, Microsoft with three deployment types and Anthropic with two settings. Quoted from each vendor's own documents, including why “available” is not “pinned”, the excluded services, and the 1.1x US-only inference surcharge. Glossary entry 11.
Read the entry →