AI data residency is a governance requirement that defines the approved geographic locations and access conditions for data, models, metadata, and operational copies. A residency decision is not complete until an organization can identify every copy, the systems that create it, the people who can reach it, and the evidence that confirms those boundaries.
AI workloads complicate residency because prompts, training data, embeddings, checkpoints, logs, backups, and support exports can move through different infrastructure layers. Use the checklist below to turn a location promise into testable technical, operational, and contractual controls with named owners and retained evidence. Include deletion and incident workflows.
Define the Residency Scope Before Evaluating Providers
Start with the governing requirement, not a preferred data-center location. Legal, privacy, security, records, and business teams should define which data classes are restricted, which locations are allowed, whether remote access counts as a transfer, and what exceptions require approval. The answer can differ for personal data, regulated records, source documents, model artifacts, and telemetry.

Data residency is not the same as data sovereignty, localization, security, or regulatory compliance. Residency addresses where data is stored or processed. Sovereignty concerns which laws and authorities apply. Security addresses protection. Compliance combines these with sector-specific duties, risk management, contracts, and operational safeguards.
Map Every AI Data Asset and Derived Copy
| Asset | Locations to record | Evidence to retain |
| Source and training data | Landing, preprocessing, active storage, replicas, archives | Data-flow map, storage inventory, policy configuration |
| Prompts and inference data | Gateway, logs, queues, caches, model service, analytics | Request flow, logging configuration, retention schedule |
| Embeddings and vector indexes | Index nodes, replicas, snapshots, rebuild source | Index topology, tenant boundary, backup location |
| Models and checkpoints | Training storage, registries, deployment nodes, recovery copies | Artifact inventory, promotion record, deletion evidence |
| Control-plane metadata | Identity, orchestration, monitoring, ticketing, support tools | System list, fields collected, provider locations |
| Backups and disaster recovery | Primary, secondary, offline, and cross-site copies | Backup policy, restore test, failover data-flow map |
Checklist 1: Approved Locations and Processing Paths
- Record allowed countries, regions, facilities, and cloud regions. Avoid a broad label such as domestic when a contract or risk decision requires a specific boundary.
- Document storage and processing separately. Data can remain stored in one location while a model, administrator, or service processes it from another.
- Trace ingestion and egress. Include uploads, exports, APIs, data labeling, evaluation, and downstream applications.
- Identify temporary copies. Caches, local GPU disks, staging areas, failed uploads, and job workspaces need the same residency decision.
- Approve failover paths. Confirm where data moves during backup restoration, capacity expansion, or disaster recovery.
Checklist 2: Administrative and Support Access
Location controls must include people and tools. Record where provider administrators, customer engineers, security analysts, and subcontractors can access systems. Determine whether remote access from another jurisdiction is permitted and whether support data enters ticketing, chat, or screen-sharing platforms.
Require named accounts, strong authentication, least privilege, time-bounded elevation, session logging, approval, and periodic review. For emergency access, document the break-glass process and notification. A provider's data center can be in the approved location while its global support workflow creates an unreviewed access path.
Checklist 3: Encryption and Key Jurisdiction
Verify encryption for storage, network transport, backups, local disks, model artifacts, and management traffic. Then document who controls keys, where keys and recovery material are stored, who can request decryption, how rotation works, and how key access is logged.
Customer-controlled keys can reduce provider access but do not by themselves prove residency or compliance. The provider still operates systems that may maintain encrypted data, metadata, or backups. For HIPAA-regulated cloud use, encryption without the provider holding the key does not automatically remove business-associate obligations.
Checklist 4: Subprocessors and External Services
List every subprocessor and external service that can receive data or operational metadata. Common examples include backup platforms, monitoring, incident management, support systems, identity services, model APIs, and security tools. Record service purpose, data fields, location, retention, access, and notification requirements for changes.
Contractual commitments should require the provider to maintain the approved boundary and disclose relevant changes before they affect the workload. The customer should have a review process for new subprocessors rather than discovering them during an audit or incident.
Checklist 5: Retention, Export, and Verified Deletion
Define retention by asset and purpose. Training data, prompts, checkpoints, logs, and backups can require different periods. State what happens when a project ends, a person exercises a data right, or the provider contract terminates. Confirm export format and the time needed to retrieve large datasets and model artifacts.
Deletion should cover primary data, replicas, caches, failed jobs, snapshots, backup expiry, local disks, and support attachments. Ask what evidence is available and which copies are subject to delayed deletion or legal hold. Test deletion in a non-production project so the process is known before a sensitive request arrives.
Checklist 6: Monitoring and Change Evidence
Maintain an asset and location register linked to configuration evidence. Alert on new storage locations, cross-region replication, unapproved endpoints, backup-policy changes, external transfers, and administrative access from unexpected locations. Review the register after architecture changes, provider releases, incident recovery, and capacity expansion.
Evidence should be usable by auditors and operators: system inventory, data-flow diagrams, access logs, key records, backup reports, subprocessor schedules, deletion records, and approved exceptions. A policy statement without operational evidence cannot show that the environment stayed inside its approved boundary.
Private infrastructure can make the data path easier to define because compute, network, storage, and operations can be designed around a dedicated environment. It does not remove the need to trace remote support, control-plane metadata, backups, and external integrations.
OneSource Cloud's Private AI Infrastructure supports dedicated environments and U.S.-based data-location planning. Teams should align those infrastructure choices with AI Storage Architecture for checkpoints, indexes, logs, and retention, and with their own legal and risk requirements.
FAQ
Does HIPAA require healthcare AI data to stay in the United States?
HIPAA does not impose a blanket U.S.-only storage rule for cloud data. However, geographic location can change risk and enforceability, so covered entities and business associates must include it in risk analysis and management. A business associate agreement and the applicable safeguards remain necessary when a provider maintains ePHI.
Do embeddings and vector indexes count as resident data?
They should be included in the residency inventory because they are derived from source data and may retain sensitive meaning or identifiers. Record where indexes, replicas, snapshots, metadata, and rebuild inputs live, who can access them, and how they are deleted or reconstructed.
Can remote support violate a data residency policy?
It can, depending on the policy, contract, data exposed, and applicable law. A system may be hosted in an approved location while an administrator accesses data or logs from elsewhere. Define whether remote access is permitted, limit the data visible, and retain session and approval evidence.
How often should an AI residency map be reviewed?
Review it at a regular risk-based cadence and whenever the workload, provider, subprocessor, backup design, model API, support process, or disaster-recovery path changes. Automated configuration monitoring can identify drift between formal reviews, but a person should approve changes that alter the residency boundary.
Summary
Prove AI data residency by mapping every asset and copy, controlling access and keys, reviewing subprocessors, testing recovery and deletion, and retaining operational evidence. A OneSource Cloud architecture review can help align private AI compute, storage, networking, and operations with an approved data-location boundary.