The four data-protection requirements for LLMs — checking "not stored", "not seen", "not learned from" and "not sent abroad" separately
Glossary entry 8. After our zero data retention (ZDR) explainer, readers kept asking the same thing: "If a vendor supports ZDR, that means their staff can't see my data, it isn't used for training, and it never leaves the country — right?" The answer is no. Those are four separate promises, and satisfying one does not satisfy the others. This page splits them into ① not stored (ZDR) ② not seen (ZOA) ③ not learned from (training terms) ④ not sent abroad (data residency), then lines up five providers' own wording to show which requirement is the default and which has to be requested or approved. It is written to be usable as a pre-contract checklist.
"Not stored", "not seen by operators", "not used for training" and "never leaves the region" are four separate promises.
Why split them into four
Because different providers satisfy different requirements by default. Technically the four map to different stages of the same request.
① Not stored / ZDR
Zero Data Retention — inputs and outputs are not kept
After inference finishes, the prompt and the response are not retained. This is about the retention window: zero days or thirty.
② Not seen / ZOA
Zero Operator Access — service operators cannot read it
Whether or not anything is stored, no human on the provider's side can access the content. This is independent of retention.
③ Not learned from
Training restriction — not used to train or improve models
A contractual term that the data is not used to train the provider's models. It lives in the contract, not in the retention setting.
④ Not sent abroad
Data residency — where data is stored and processed
Storage location and processing location are different things, decided by the endpoint or deployment type you choose.
When you read "ZDR supported", the question to ask is which of the four it refers to. Mixing them up produces the worst outcome: believing you have covered an audit requirement when you have not.
① Not stored — zero data retention (ZDR)
Start with whether it is the default. ZDR means "not stored by default", not "always on" (see entry 7 on ZDR).
- Default retention differs by provider. AWS offers retention modes (
none/default/inherit/aws_review) set per account and per project (new projects inherit). Microsoft retains abuse-monitoring logs for up to 30 days by default and drops them entirely once modified abuse monitoring is approved. OpenAI's ZDR is application-only; the default is up to 30 days of logs. - Some models are exceptions. AWS publishes a list of models that require retention: OpenAI-based models retain only traffic flagged by a classifier, for up to 30 days, while Claude Fable 5 / 5.1 retain all traffic for 30 days. The list changes (updated 2026-09-26) — ZDR does not mean zero everywhere.
- "Not stored" and "no logs" are different. Even with no content retention, request metadata (who called what, when) remains in operational logs. Content not being stored is not the same as calls not being recorded.
② Not seen — zero operator access (ZOA)
ZDR is "not retained"; ZOA is "no human can read it". Independent of whether anything is stored.
- Human access is decided by the abuse-monitoring design. At Microsoft, automated review (an LLM classifier) is on by default, and that data is described as not stored by the system and not used to train AI models. Human review happens only when the abuse-monitoring system flags content, and reviewers are based in the EEA — that is how the access surface is limited.
- There is an enterprise opt-out. Once Microsoft's modified abuse monitoring is approved, neither storage nor human review takes place (automated review remains).
- Some providers do not publish ZOA wording at all. When that is the case we treat it as unverified rather than assuming "ZDR implies no access".
→ Deep dive now available: #9: Zero operator access covers the three layers that block access, CMK vs BYOK vs HYOK, the approval workflows (Customer Lockbox, Access Approval, Key Access Justifications, EKM) and the emergency exceptions, all quoted from primary sources.
③ Not learned from — training restrictions
A separate clause from retention. "Stored but never used for training" is a real configuration.
- "Not passed to the model provider" is yet another clause. Microsoft states that data for models sold by Azure is not made available to OpenAI or other providers, and that those models do not interact with services operated by the providers. That sentence matters when you buy a resold model through Azure.
- Some services do train on your data by default. DeepSeek is the clearest example: its privacy policy covers using data to improve services and to train machine learning models and algorithms, with a right to opt out. However, whether the same terms apply to the API, and whether an opt-out alone satisfies ③, is not confirmed in the primary sources (the policy also states personal data is processed and stored in China). Review the contract terms before relying on it.
- Free tiers and business/API terms differ. The same company often has different training terms for consumer UIs and for API or commercial contracts. Everything on this page is based on API and commercial documentation.
④ Not sent abroad — data residency
This is where most confusion lives. Storage location and processing location are separate, and choosing one does not guarantee the other.
In other words, on Google Cloud storage follows the location you pick and processing follows the endpoint you pick — and endpoints come in three tiers.
- jurisdictional multi-region endpoints (US / EU) — processing stays inside the jurisdiction. Use this when data must not leave the region.
- global endpoints — explicitly documented as possibly processing in any Google location, and as not providing regional isolation or data residency guarantees.
- AWS cross-Region inference — storage stays "in the AWS Region where you are using Amazon Bedrock", but inference can spread across a geographic profile (US / EU / APAC) or the global profile. The global profile's data residency is "any supported AWS commercial Region worldwide", and AWS itself recommends the geographic option.
- Azure decides by deployment type. Standard / Regional stays within the customer-selected Azure geography (operationally it may be processed in another region inside that geography), DataZone is US / EU / APAC (the EU Data Zone includes EFTA countries), and Global can process in any Azure region. Choosing Global Standard means you are not choosing a processing location.
- Anthropic (Claude API) is per request.
inference_geotakesusorglobal, and the default isglobal. Workspace geo is US-only and cannot be changed after creation. US-only inference costs 1.1× on every token, and passinginference_geoto Opus 4.5, Sonnet 4.5, Haiku 4.5 or earlier returns a 400 error.inference_geohere refers to using Anthropic's API directly; regional handling on partner platforms (Bedrock, Vertex AI, Foundry) is a separate setting to verify.
"You can pick a region" does not mean "your data never leaves the country". Pick an option that pins the geography and back it with the contract.
Five-provider matrix — what is default and what must be requested
Where the vendors' own documentation had it, as of 2026-09-29. "Requested" means it may not be available before you sign.
| Provider | ① Not stored | ② Not seen | ③ Not learned from | ④ Not sent abroad |
|---|---|---|---|---|
| AWS Bedrock | No storage by default (set via retention mode); models that require retention are exceptions | ZOA documented as part of the data security model (exception models may involve human review) | Documented as not used to train models by default | Stored in the region you use; choose Geographic for cross-Region inference |
| Microsoft Foundry | Abuse-monitoring logs kept up to 30 days by default; modified abuse monitoring removes storage | Automated review by default; human review only when flagged, by EEA-based reviewers | "NOT used to train any generative AI foundation models without your permission or instruction" + not available to the provider | Standard / DataZone / Global; Global processes in any region, Standard stays inside the selected geography (may cross regions within it) |
| OpenAI API | ZDR is application-based (approval plus a modified retention amendment); default is up to 30 days | Not confirmed in the vendor's docs (worth re-checking on the current official page, as terms change often) | Documented as not training on API / business data | Selected per project; outside the US requires approval and a modified retention amendment, UAE adds a further approval |
| Anthropic | ZDR on request | Not confirmed in the vendor's docs (worth re-checking on the current official page, as terms change often) | By default, commercial data is not used to train models | inference_geo (the setting that picks which country processes the request: us/global, default global); US-only costs 1.1× |
| Google Cloud | Handled by contract and configuration (no single default confirmed) | — | AI/ML Privacy Commitment says customer data is not used to train foundation models (summary; we have not captured the full wording) | Location plus endpoint; jurisdictional (US/EU) keeps processing in-region |
← Scrolls horizontally. A blank cell means "we could not verify it from primary sources", not "it does not exist".
A five-step check before you sign
- Decide which requirements are non-negotiable for you. Regulation, customer contracts and internal policy will rank the four. Treating all four as mandatory stalls the selection.
- Look up the provider's defaults. The matrix above is the starting point: the default retention window (0 or 30 days) and whether exception models exist.
- Configure or request whatever the default does not cover. Bedrock retention modes, the Azure deployment type, Anthropic's
inference_geo, ZDR or modified abuse monitoring applications. This is where "requested" becomes a blocker. - Pin it down in the contract and DPA. Items that require approval are not guaranteed by a console toggle. Get them in writing.
- Record it and re-check when you add a model. Keep a note of which requirement is covered how. Exception lists get revised (AWS updated theirs on 2026-09-26), so review on every new model. This record is the first thing lost when staff change.
Four myths
Myth 1
"With ZDR, operators cannot see it either"
ZDR is ① and ZOA is ② — different promises. Enabling ZDR says nothing about whether provider staff can access content.
Myth 2
"If it is not stored, it is not used for training"
Training use is governed by contract terms (③). A retention setting can be perfect while ③ is undocumented.
Myth 3
"I picked a region, so nothing leaves the country"
Global options (Azure Global, Bedrock global profile, inference_geo=global) explicitly process in any region. Choose the geography-pinned option instead.
Myth 4
"It is just a resold model, so it is safer"
On a resale route the reseller's terms apply (no training, not passed to the provider). But deployment type and endpoint still decide where processing happens — see DeepSeek via Azure.
FAQ
- If a vendor says "ZDR supported", am I safe?
- First check which of the four it refers to. It is usually ① only, leaving ②–④ to verify separately. Even ① has exception models, and that list is revised over time.
- How do I pin the region?
- Pick the option that names the geography: Geographic cross-Region inference on Bedrock, DataZone / Regional on Azure,
inference_geo: uson Anthropic, or jurisdictional multi-region (US / EU) endpoints on Google Cloud. Global options are documented as processing in any region, so they cannot be used to pin. - Does the same table apply to free tiers and consumer apps?
- No. Consumer UIs and API / commercial contracts often carry different terms, and training clauses differ most. This page is based on API and commercial documentation; check the specific plan's terms separately.
- Why not just say "zero retention" and be done?
- Because configurations such as "stored, but never used for training, unreadable by staff, processed in-country" are real — and so is "not retained, but passed to the provider and used for training". Audits ask requirement by requirement, so collapsing them into "ZDR" leaves you unable to answer.
Sources
- AWS — Data protection in Amazon Bedrock (ZDR, ZOA, retention modes, regional storage; retrieved 2026-09-27)
- AWS — Model data retention (exception model list; retrieved 2026-09-27)
- Microsoft — Data, privacy, and security for Azure direct models and Abuse monitoring (training terms, not available to the provider, 30-day retention; retrieved 2026-09-27)
- Microsoft — Azure deployment types and data residency documentation (DataZone / Global, EFTA in the EU Data Zone; retrieved 2026-09-27)
- Anthropic — Regional compliance (
inference_geo, default global, 1.1× pricing, 400 error on older models; retrieved 2026-09-27) - OpenAI — ZDR, abuse monitoring controls and data residency (application process, conditions outside the US, additional UAE approval; retrieved 2026-09-27)
- Google Cloud — Data governance and generative AI and AI/ML Privacy Commitment (storage vs processing, the three endpoint tiers, jurisdictional guarantees; retrieved 2026-09-29)
- Related: zero data retention (ZDR) · China AI · DeepSeek V4.1 Flash (Azure walkthrough) · all glossary entries