When an AI agent executes code or fires a tool call, something physical happens — files change, APIs receive requests, money moves. Sandboxing is the discipline of making those happenings containable: a boundary around execution that turns an agent's mistakes and compromises into events instead of catastrophes. This page treats the sandbox as a compliance control, in four parts: what it covers, who owns each layer, what evidence proves it, and the risk it honestly does not absorb.
Scope: What the Isolation Boundary Covers
Agent sandboxing places tool and code execution inside an isolation boundary: agent-generated code runs where a compromised or mistaken agent cannot reach the host, filesystem access is scoped to per-session ephemeral state, and tool side effects happen against scoped resources rather than production credentials — the boundary exists so the agent's mistakes and compromises stay containable events.
| Surface | Inside the boundary | Without the boundary |
| Agent-generated code | Executes in an isolated runtime, host unreachable | Generated code runs with the caller's full privileges |
| Filesystem | Per-session ephemeral state, destroyed after the session | Agent reads and writes durable host state |
| Tool side effects | Act against scoped resources with proxied credentials | Production credentials flow into every tool call |
| Network | Explicit allowlist, everything else refused | Whatever the host's default routing permits |
Scope should be drawn per workload shape: a code-heavy agent that runs its own generated programs needs the full execution boundary, while an API-only agent that never executes code needs the egress and credential halves more than the compute isolation half. Deployment guidance for agent sandboxes treats scoped access, network controls, storage, and lifecycle automation as the verifiable configuration surface — the boundary is a set of settings someone can audit, not a vibe of safety.
Shared Responsibility: Who Owns Which Layer

The boundary is layered responsibility: the agent framework owns the tool interfaces and their contracts, the sandbox runtime owns isolation, egress filtering, and credential proxying — secrets delivered through a proxy rather than into the sandbox, network restricted to explicit allows — and the infrastructure provider owns host tenancy and physical isolation; a boundary with an unowned layer is a boundary with a hole.
- Agent framework layer: tool interfaces and contracts — what parameters tools accept, what they may return, how errors surface.
- Sandbox runtime layer: isolation technology, egress filtering, and credential proxying — secrets reach tools through the proxy, never into the sandbox itself, minimizing exfiltration risk from a compromised session.
- Infrastructure layer: host tenancy, physical isolation, and the guarantees the compute substrate underneath provides.
The layer map earns its keep during incidents and audits, when the first question is whose control failed. An egress violation with a documented runtime-layer owner is a fix; the same violation with three teams each assuming another had it is a finding that recurs. The isolation technology choice also lives at the runtime layer, and documented comparisons of agent sandbox approaches — scored on isolation model, cold start, density, and cost — make it conditional on workload shape rather than a default.
Evidence Artifacts: Proving the Boundary Holds
The sandbox proves itself through artifacts: the isolation configuration as deployed (not as designed), egress logs demonstrating default-deny with explicit allows, credential-proxy records showing secrets never entered the sandbox, and per-session state destruction records — the evidence an assessor or incident review reads instead of trusting the architecture diagram.
- Isolation configuration as deployed: the runtime settings actually in effect, captured from the environment — diagrams age, configs are the truth.
- Egress logs: refused connections outnumbering allowed ones is normal and healthy; the allows should map one-to-one to the documented tool endpoints.
- Credential-proxy records: which secrets were delivered to which tools, through the proxy — and the absence of any direct secret delivery into sandbox memory.
- State destruction records: per-session ephemeral state actually destroyed at session end, proving the filesystem scoping is not advisory.
Artifact depth should track the sensitivity of what the agent touches: an agent reading public reference data needs the set at review level, while an agent acting on customer systems needs all four, retained, and someone whose job includes reading them. The artifacts are the difference between "we run agents in sandboxes" and "we can show you, for any past session, exactly what the boundary did."
Residual Risk: What the Sandbox Does Not Stop
The documented residual risk is the reasoning layer: sandboxes isolate execution, not persuasion — an injected instruction that steers an authenticated, correctly-scoped agent to a permitted action passes every boundary this page describes, which is why sandboxing composes with input controls and agent identity rather than replacing them, and why the residual risk gets written down instead of absorbed by optimism.
Writing the residual down has an operational purpose: it tells the incident-response team what the sandbox's clean logs do not mean. A perfectly contained session can still have performed a damaging permitted action; a perfect allowlist still permits the listed tools. Teams that document this residual staff the other layers — input filtering, permission scoping, behavior monitoring — instead of discovering their absence during an incident that passed through every control they had.
FAQ
Does a sandbox stop prompt injection?
No — and treating it as if it does is the dangerous half-truth: a sandbox contains what an agent can touch, while injection steers an authenticated agent toward actions it is already permitted to take; input filtering and scoped permissions shrink that surface, and the sandbox bounds the blast — layered controls, no single layer absolved.
What should agent egress rules allow?
Default-deny with explicit allows: the tool endpoints each agent genuinely needs, proxied and logged, with everything else — including arbitrary DNS, IP literals, and surprise metadata endpoints — refused; egress is written as an allowlist because a compromised agent's first move is a network call, and the allowlist is the move it cannot make.
Should agent sandboxes use containers or microVMs?
By workload shape: containers win on density and cold start for high-volume, short-lived tool calls; microVMs win on isolation strength for code-heavy agents running untrusted output; documented comparisons score isolation model, cold start, density, and cost — pick the weakest dimension your workload actually stresses.