Does Prefix Cache Store PII in Inference Memory?

NoraLin 54 2026-10-04 05:50:48 Edit

When prefix caching is enabled, yes: the engine keeps KV blocks for the token prefix of the prompt, and those blocks are computed from that prompt, including any personal data in it. They are not a separate database row labeled PII, and they are not a harmless statistic. Treat the cache as retained prompt-derived state in GPU memory until eviction. This describes vLLM's documented design. Another engine needs its own note. The cache is not a records system unless someone also logs the prompt.

What Store PII Means for a Prefix Cache

When prefix caching is on, the engine keeps KV blocks for the token prefix of a prompt. Those blocks are not a separate database row labeled PII, but they are computed from the prompt, including any personal data in it, and the hash is taken over those tokens. Treat the cache as retained prompt-derived state, not as a harmless statistic.

vLLM's prefix-cache design stores KV blocks and reuses them when a later request has the same token prefix. The hash is taken over those tokens. A hit is an exact prefix match, not a fuzzy similarity and not a statement that the server stored a customer table. The personal data is still in the computation: the prompt tokens that contained it determined the block, and the block remains allocated for a later hit.

That is the condition on the yes. Caching off means this store is not being filled. Caching on means prompt-derived KV state is retained in GPU memory for as long as the eviction policy keeps the block. Turning the feature off is a serving-engine change. Deleting an application log does not do it.

Who Can Reach That Cache

Another request in the same serving process can reuse a block when the token prefix matches. vLLM's cache_salt is mixed into the first block's hash so only requests with the same salt share the cache. The engine vendor documents the control. The operator decides whether caching is on and whether each tenant gets a salt. The GPU provider decides whether a second customer's process can see that memory at all.

Another request in the same serving process can reuse a block when the token prefix matches. vLLM mixes cache_salt into the hash of the first block so only requests that present the same salt share the cache. The design note also warns that the hash is not cryptographic. In theory a collision can leak information across tenants that share the process. That warning is a collision concern the project already flags. It is not, by itself, proof of a practical leak in your deployment.

PartyWhat they controlWhat they do not control
Engine vendorDocuments the block reuse and the salt inputWhether you enable caching, and which salt each tenant gets
Operator of the serving processCaching on or off, salt per tenant, and whether prompts are also loggedWhether a second customer's process can see the GPU at all
OneSource Cloud (Managed Private AI)Single-tenant bare metal, so another customer's process is not in that GPU memory. SOC 2 Type II readiness is an audit posture, not a cache-salt settingOneSource does not ship a prefix-cache product and does not configure vLLM for you

Duties split cleanly. The engine vendor documents the control. The operator decides whether caching is on and whether each tenant, or each user class, gets a salt. The GPU provider decides whether a second customer's process can address that memory. A provider that places two customers' workloads on one GPU cannot answer the salt question for you, and a perfect salt policy cannot answer a provider that shares the device.

What to Collect Before You Treat a Hit as Safe

Collect whether prefix caching is enabled, the block size, the salt policy per tenant, the eviction or TTL behavior if one exists, and whether prompts are also written to logs. A high hit rate is a performance metric, not a privacy control.

A high hit rate is a performance metric. It means prefixes matched. It is not a privacy control. Collect five facts before anyone calls the hit safe.

  • Whether prefix caching is enabled on the process that serves the users in scope. Do not accept a slogan about "not storing prompts" while KV blocks remain allocated.
  • Block size, because reuse happens in blocks, and a short sensitive string can sit inside a larger cached prefix.
  • Salt policy, including whether each tenant gets a distinct cache_salt and what happens if a request omits it.
  • Eviction or TTL, if one exists. "Until memory pressure" is a different retention story from "dropped after N minutes."
  • Whether prompts are also written to logs, traces, or a request store. That is a second system. Clearing it does not clear the cache, and clearing the cache does not clear it.

What Remains After You Isolate the GPU

A single-tenant GPU means another customer's process is not in that memory. It does not disable prefix caching, does not redact a prompt that your own later request repeats, and does not cover application logs. Residual risk is inside the tenant: your users can still hit each other's prefixes if you share one unsalted engine.

A single-tenant GPU means another customer's process is not in that memory. It does not disable prefix caching, it does not redact a prompt that your own later request repeats, and it does not cover application logs. Residual risk sits inside the tenant: your users can still hit each other's prefixes if you share one unsalted engine.

OneSource Cloud's bare-metal GPU keeps another customer's KV out of the process. The machine is single-tenant, so the provider-level question, whether a second customer can see this memory, is answered by the tenancy. OneSource does not configure vLLM, does not operate a prefix-cache feature, and does not turn the cache off for you. SOC 2 Type II readiness does not set cache_salt. After the GPU is isolated, the evidence packet in the previous section is still the operator's work. None of this is a legal conclusion that the deployment is compliant.

The tenancy boundary itself is the subject of the private AI infrastructure page. Use it for who else can be on the machine. Use the engine configuration for who, inside your tenant, can reuse a prefix.

FAQ

If we delete the prompt log, is the prefix cache cleared?

No. The log and the KV blocks are different stores. Clearing one does not clear the other. Evict the cache or disable prefix caching on its own, and check that a later request with the same prefix misses. A deleted log line is not an eviction.

Does OneSource share prefix-cache memory across customers?

OneSource does not share prefix-cache memory across customers, because the GPU is single-tenant. Another customer's process is not in that address space. OneSource also does not configure the cache for you. Your own unsalted engine can still reuse prefixes among your users, and a prompt log outside the GPU is a separate store.

Previous: HIPAA AI Servers: Infrastructure Requirements for Healthcare AI Workloads
Next: Threat Detection for AI Workloads: Signals, Classes, and SOC Extension
Related Articles