Cold starts are the tax, warm pools are the prepayment. Keeping model replicas loaded and idle is the only way to guarantee instant answers — and it is also pure cost while the requests are away. The engineering question is not whether to stay warm; it is how warm, for which traffic, and paid how. This page is that arithmetic.
Cost Drivers: What Standing Warm Costs
Warm capacity prices the guarantee that a preloaded replica answers instantly: every pool GPU holds weights in memory and burns availability while waiting — idle standby is the entire cost line — and the tiered pattern prices it in steps: warm tier first (preloaded, instant), cold tier second (provisioned but not loaded, pays the model-load delay), on-demand last (pays the full cold start), so the real question is never whether to be warm but what fraction of traffic earns the warm premium.
| Tier | What it holds | What requests pay |
| Warm | Preloaded replicas, weights resident | Nothing in latency; idle cost always |
| Cold | Provisioned instances, models unloaded | Model-load delay on first hit |
| On-demand | Nothing standing | Full cold start: provision plus load |

The tiers turn a binary into a dial: production guidance for bursty workloads describes exactly this pattern — standby buffers hit in order, with on-demand touched only after warm and cold exhaust — and the dial setting is a business question (whose waiting costs what) answered with an economics method.
Assumptions: What the Pool Size Depends On
Pool size rests on three assumptions worth writing down: the burst profile (how fast traffic can arrive — bursty agent-style workloads need deeper standby buffers than steady chat), the load time it is preventing (model-load minutes × arrival rate is the queue the pool absorbs), and the accuracy of prediction — platform practice now maintains warm pools with predictive algorithms, pre-provisioned and pre-loaded ahead of demand, because a static pool sized to last month's bursts either starves or bleeds; snapshots add the third lever, hibernating instances so resume replaces reload.
- Burst profile: the fastest realistic ramp your traffic can produce — agents and batch jobs arrive in walls, not slopes.
- Load-time exposure: big models on slow storage mean each cold hit queues everyone behind it — the cost the warm tier prevents.
- Predictive maintenance of the pool: pre-provisioning ahead of demand beats a pool sized once and forgotten.
- Snapshot hibernation: resume-instead-of-reload for the middle tier — most of the latency win at a fraction of the idle bill.
Research on serverless inference adds the boundary condition: pools are finite, and pre-loading under capacity limits is itself a scheduling problem — another reason the warm fraction is a decision, not a default.
Decision Framework: When Readiness Pays
Run the lead-time math per service tier: warm capacity pays when cold-start time multiplied by the value of waiting requests exceeds the idle cost of the pool that prevents it — interactive products with real users waiting compute one number, internal batch tools compute a much smaller one — and let the tiers carry the answer: warm only the tier whose waiting is expensive, snapshot the middle, on-demand the rest; the framework's output is a percentage of the fleet standing warm, revisited as traffic patterns drift.
Procurement shape quietly decides the final term of that equation: on committed flat-rate capacity, a warm replica burns hours already paid for — dedicated environments such as OneSource Cloud's price the fleet as a whole — while on metered serverless, every warm hour is a new invoice line; identical architecture, materially different readiness economics.
FAQ
How large should a warm pool be?
Deep enough to absorb arrival between cold-start and scale-out: size against your burst profile (how many concurrent preloaded replicas the fastest realistic ramp needs), validate against the load-time math (minutes of model load times arrival rate is the queue the pool exists to prevent), and let predictive sizing adjust it — static pools sized once are photographs of last quarter's traffic.
Are snapshots a real alternative to fully warm GPUs?
For the middle tier, yes: snapshot-based standby hibernates provisioned instances so a resume replaces a full model load — most of the latency win at a fraction of the idle cost — which is why the tier pattern works: fully warm for the traffic whose waiting is expensive, snapshotted for the middle, on-demand for the long tail; snapshots do not beat true warmth for the hottest tier, they make the tiers below it cheap.
Do warm pools cost less on committed capacity than serverless?
At the margin, yes: idle-warm on metered serverless bills per hour of standby at premium rates, while warm replicas on committed flat-rate capacity — the OneSource Cloud model — burn hours already paid for, so the readiness premium shrinks to the opportunity cost of the hours rather than a new invoice line; the warm-pool math is one more place where the procurement shape quietly decides the architecture economics.