The four data-protection requirements for LLMs — checking "not stored", "not seen", "not learned from" and "not sent abroad" separately

Glossary entry 8. After our zero data retention (ZDR) explainer, readers kept asking the same thing: "If a vendor supports ZDR, that means their staff can't see my data, it isn't used for training, and it never leaves the country — right?" The answer is no. Those are four separate promises, and satisfying one does not satisfy the others. This page splits them into ① not stored (ZDR) ② not seen (ZOA) ③ not learned from (training terms) ④ not sent abroad (data residency), then lines up five providers' own wording to show which requirement is the default and which has to be requested or approved. It is written to be usable as a pre-contract checklist.

"Not stored", "not seen by operators", "not used for training" and "never leaves the region" are four separate promises.

It is rare for one provider to satisfy all four by default. AWS describes ① and ② as a "data security model" while keeping a list of models that require retention. Microsoft documents ③ explicitly, yet retains abuse-monitoring logs for up to 30 days by default. Anthropic delivers ④ through a per-request setting called inference_geo, and its default is global. The work is checking which of the four is the default and which must be requested — one at a time.

● Last updated:  |  Policy: only content we could confirm in the vendors' own documents; anything unconfirmed is labelled as such

Why split them into four

Because different providers satisfy different requirements by default. Technically the four map to different stages of the same request.

① Not stored / ZDR

Zero Data Retention — inputs and outputs are not kept

After inference finishes, the prompt and the response are not retained. This is about the retention window: zero days or thirty.

② Not seen / ZOA

Zero Operator Access — service operators cannot read it

Whether or not anything is stored, no human on the provider's side can access the content. This is independent of retention.

③ Not learned from

Training restriction — not used to train or improve models

A contractual term that the data is not used to train the provider's models. It lives in the contract, not in the retention setting.

④ Not sent abroad

Data residency — where data is stored and processed

Storage location and processing location are different things, decided by the endpoint or deployment type you choose.

When you read "ZDR supported", the question to ask is which of the four it refers to. Mixing them up produces the worst outcome: believing you have covered an audit requirement when you have not.

① Not stored — zero data retention (ZDR)

Start with whether it is the default. ZDR means "not stored by default", not "always on" (see entry 7 on ZDR).

"uses a zero data retention (ZDR) data security model. This means that by default, Amazon Bedrock does not store model inputs or outputs." — AWS documentation (Amazon Bedrock data security model)
  • Default retention differs by provider. AWS offers retention modes (none / default / inherit / aws_review) set per account and per project (new projects inherit). Microsoft retains abuse-monitoring logs for up to 30 days by default and drops them entirely once modified abuse monitoring is approved. OpenAI's ZDR is application-only; the default is up to 30 days of logs.
  • Some models are exceptions. AWS publishes a list of models that require retention: OpenAI-based models retain only traffic flagged by a classifier, for up to 30 days, while Claude Fable 5 / 5.1 retain all traffic for 30 days. The list changes (updated 2026-09-26) — ZDR does not mean zero everywhere.
  • "Not stored" and "no logs" are different. Even with no content retention, request metadata (who called what, when) remains in operational logs. Content not being stored is not the same as calls not being recorded.

② Not seen — zero operator access (ZOA)

ZDR is "not retained"; ZOA is "no human can read it". Independent of whether anything is stored.

"zero operator access (ZOA) ... no operators of the service can access model input or output" — AWS documentation (Amazon Bedrock)
  • Human access is decided by the abuse-monitoring design. At Microsoft, automated review (an LLM classifier) is on by default, and that data is described as not stored by the system and not used to train AI models. Human review happens only when the abuse-monitoring system flags content, and reviewers are based in the EEA — that is how the access surface is limited.
  • There is an enterprise opt-out. Once Microsoft's modified abuse monitoring is approved, neither storage nor human review takes place (automated review remains).
  • Some providers do not publish ZOA wording at all. When that is the case we treat it as unverified rather than assuming "ZDR implies no access".

→ Deep dive now available: #9: Zero operator access covers the three layers that block access, CMK vs BYOK vs HYOK, the approval workflows (Customer Lockbox, Access Approval, Key Access Justifications, EKM) and the emergency exceptions, all quoted from primary sources.

③ Not learned from — training restrictions

A separate clause from retention. "Stored but never used for training" is a real configuration.

"Customer Data and Customer Content is ... NOT used to train any generative AI foundation models without your permission or instruction." — Microsoft (Azure / Microsoft Foundry data and privacy documentation)
"By default, Anthropic does not use customer data from commercial deployments to train models." — Anthropic (regional compliance documentation)
  • "Not passed to the model provider" is yet another clause. Microsoft states that data for models sold by Azure is not made available to OpenAI or other providers, and that those models do not interact with services operated by the providers. That sentence matters when you buy a resold model through Azure.
  • Some services do train on your data by default. DeepSeek is the clearest example: its privacy policy covers using data to improve services and to train machine learning models and algorithms, with a right to opt out. However, whether the same terms apply to the API, and whether an opt-out alone satisfies ③, is not confirmed in the primary sources (the policy also states personal data is processed and stored in China). Review the contract terms before relying on it.
  • Free tiers and business/API terms differ. The same company often has different training terms for consumer UIs and for API or commercial contracts. Everything on this page is based on API and commercial documentation.

④ Not sent abroad — data residency

This is where most confusion lives. Storage location and processing location are separate, and choosing one does not guarantee the other.

"Data-at-rest (storage) ... remains physically stored in the specific Google Cloud location you chose. This residency is maintained regardless of which endpoint you use." — Google Cloud (generative AI data residency documentation)

In other words, on Google Cloud storage follows the location you pick and processing follows the endpoint you pick — and endpoints come in three tiers.

  • jurisdictional multi-region endpoints (US / EU) — processing stays inside the jurisdiction. Use this when data must not leave the region.
  • global endpoints — explicitly documented as possibly processing in any Google location, and as not providing regional isolation or data residency guarantees.
  • AWS cross-Region inference — storage stays "in the AWS Region where you are using Amazon Bedrock", but inference can spread across a geographic profile (US / EU / APAC) or the global profile. The global profile's data residency is "any supported AWS commercial Region worldwide", and AWS itself recommends the geographic option.
  • Azure decides by deployment type. Standard / Regional stays within the customer-selected Azure geography (operationally it may be processed in another region inside that geography), DataZone is US / EU / APAC (the EU Data Zone includes EFTA countries), and Global can process in any Azure region. Choosing Global Standard means you are not choosing a processing location.
  • Anthropic (Claude API) is per request. inference_geo takes us or global, and the default is global. Workspace geo is US-only and cannot be changed after creation. US-only inference costs 1.1× on every token, and passing inference_geo to Opus 4.5, Sonnet 4.5, Haiku 4.5 or earlier returns a 400 error. inference_geo here refers to using Anthropic's API directly; regional handling on partner platforms (Bedrock, Vertex AI, Foundry) is a separate setting to verify.

"You can pick a region" does not mean "your data never leaves the country". Pick an option that pins the geography and back it with the contract.

Five-provider matrix — what is default and what must be requested

Where the vendors' own documentation had it, as of 2026-09-29. "Requested" means it may not be available before you sign.

Data-protection requirements across five providers
Provider① Not stored② Not seen③ Not learned from④ Not sent abroad
AWS Bedrock No storage by default (set via retention mode); models that require retention are exceptions ZOA documented as part of the data security model (exception models may involve human review) Documented as not used to train models by default Stored in the region you use; choose Geographic for cross-Region inference
Microsoft Foundry Abuse-monitoring logs kept up to 30 days by default; modified abuse monitoring removes storage Automated review by default; human review only when flagged, by EEA-based reviewers "NOT used to train any generative AI foundation models without your permission or instruction" + not available to the provider Standard / DataZone / Global; Global processes in any region, Standard stays inside the selected geography (may cross regions within it)
OpenAI API ZDR is application-based (approval plus a modified retention amendment); default is up to 30 days Not confirmed in the vendor's docs (worth re-checking on the current official page, as terms change often) Documented as not training on API / business data Selected per project; outside the US requires approval and a modified retention amendment, UAE adds a further approval
Anthropic ZDR on request Not confirmed in the vendor's docs (worth re-checking on the current official page, as terms change often) By default, commercial data is not used to train models inference_geo (the setting that picks which country processes the request: us/global, default global); US-only costs 1.1×
Google Cloud Handled by contract and configuration (no single default confirmed) — AI/ML Privacy Commitment says customer data is not used to train foundation models (summary; we have not captured the full wording) Location plus endpoint; jurisdictional (US/EU) keeps processing in-region

← Scrolls horizontally. A blank cell means "we could not verify it from primary sources", not "it does not exist".

A five-step check before you sign

  1. Decide which requirements are non-negotiable for you. Regulation, customer contracts and internal policy will rank the four. Treating all four as mandatory stalls the selection.
  2. Look up the provider's defaults. The matrix above is the starting point: the default retention window (0 or 30 days) and whether exception models exist.
  3. Configure or request whatever the default does not cover. Bedrock retention modes, the Azure deployment type, Anthropic's inference_geo, ZDR or modified abuse monitoring applications. This is where "requested" becomes a blocker.
  4. Pin it down in the contract and DPA. Items that require approval are not guaranteed by a console toggle. Get them in writing.
  5. Record it and re-check when you add a model. Keep a note of which requirement is covered how. Exception lists get revised (AWS updated theirs on 2026-09-26), so review on every new model. This record is the first thing lost when staff change.

Four myths

Myth 1

"With ZDR, operators cannot see it either"

ZDR is ① and ZOA is ② — different promises. Enabling ZDR says nothing about whether provider staff can access content.

Myth 2

"If it is not stored, it is not used for training"

Training use is governed by contract terms (③). A retention setting can be perfect while ③ is undocumented.

Myth 3

"I picked a region, so nothing leaves the country"

Global options (Azure Global, Bedrock global profile, inference_geo=global) explicitly process in any region. Choose the geography-pinned option instead.

Myth 4

"It is just a resold model, so it is safer"

On a resale route the reseller's terms apply (no training, not passed to the provider). But deployment type and endpoint still decide where processing happens — see DeepSeek via Azure.

FAQ

If a vendor says "ZDR supported", am I safe?
First check which of the four it refers to. It is usually ① only, leaving ②–④ to verify separately. Even ① has exception models, and that list is revised over time.
How do I pin the region?
Pick the option that names the geography: Geographic cross-Region inference on Bedrock, DataZone / Regional on Azure, inference_geo: us on Anthropic, or jurisdictional multi-region (US / EU) endpoints on Google Cloud. Global options are documented as processing in any region, so they cannot be used to pin.
Does the same table apply to free tiers and consumer apps?
No. Consumer UIs and API / commercial contracts often carry different terms, and training clauses differ most. This page is based on API and commercial documentation; check the specific plan's terms separately.
Why not just say "zero retention" and be done?
Because configurations such as "stored, but never used for training, unreadable by staff, processed in-country" are real — and so is "not retained, but passed to the provider and used for training". Audits ask requirement by requirement, so collapsing them into "ZDR" leaves you unable to answer.

Sources