How Much Spare GPU Capacity Inference Needs for Failover

NoraLin 13 2026-10-07 01:18:13 Edit

Serving fleets size for peak load and then discover the second axis: failure. A node dies at the worst moment — statistically, the worst moment is exactly peak — and the question that decides whether users notice is how much capacity you hold that exists for no traffic at all. Spares are capacity with a job description, and like all capacity they can be computed rather than guessed. This page is the computation, the failover design that spends it, and the drill that proves both.

Prerequisites: Failure Rate, Repair Window, Target

Three numbers feed the spare decision: your fleet's failure rate (nodes per thousand per month, from your own history or vendor RMA baselines), the repair window (how long a failed node is out — drain, replace, restore), and the availability target your SLO actually promises — because spare capacity exists to cover simultaneous-failure-plus-repair-window events, and any number computed without all three is a guess wearing decimals.

InputTypical sourceWhat it drives
Failure rateYour fleet history; vendor RMA baselinesExpected concurrent failures
Repair windowDrain + replace + restore timeHow long each failure bites
Availability targetThe SLO you actually promisedHow much of the bite users may feel

The multiplication is unforgiving: a fleet with two expected concurrent failures and a four-week repair window needs its spares to bridge weeks, not minutes — which is why the repair window belongs in the spare conversation even though it feels like a maintenance topic. Availability-target practice treats redundancy requirements as derived from the target itself, not from comfort.

Size and Design: Spare Count and Failover Pattern

Sizing and design run together: compute the spare count the way inference rack calculators do — target throughput, measured per-node performance, utilization, and explicit N+1 redundancy on top — then choose the failover pattern that spends it: health-monitor-driven node cordoning and replacement for most fleets, active-passive for stateful tiers, active-active when the availability target genuinely demands it, with redundant clusters or regional failover reserved for the workloads whose pause is intolerable.

  1. Compute the count: from target TPS, measured per-node performance, utilization, and headroom — with explicit N+1 redundancy layered on top, the pattern serving-capacity calculators implement directly.
  2. Pattern for stateless tiers: health monitoring that detects, cordons the failed node, and replaces its replicas — the default production mechanism set alongside pod replacement and cluster-level redundancy.
  3. Pattern for stateful tiers: active-passive with promotion for caches, routers, and anything holding connections or state.
  4. Pattern for the intolerable: active-active, and for regional SLOs a second region's worth — the expensive tiers, bought only where the SLO demands.

Each step up the pattern ladder costs capacity, and the discipline is spending it only where the target requires: an estate that holds regional spares for a batch scoring tier is paying active-active prices for delay-tolerant work, while an estate with no spares at all on its interactive tier has an availability policy of optimism.

Verify: The Failover Drill

The design proves itself one way: at target load, kill a serving node on purpose and measure whether latency SLOs hold through detection, cordon, and replacement — because spare capacity that has never absorbed a real failure is inventory with a theory, and the drill is where the health-monitor thresholds, the drain behavior, and the spare count are all tested in the only conditions that matter: the ones your users are also experiencing.

  • At target load: the drill means something at peak, not at idle — failover that works at 20% load proves nothing about the peak that motivated the spares.
  • Measure the SLO through the event: p95 during detection-to-replacement is the number the design lives or dies by.
  • Check the thresholds: detection that fires too slowly or drain that cascades shows up here, cheaply.
  • Re-run on change: new serving stack, new topology, new node count — the drill follows the fleet.

Holding capacity idle-for-failure has real cost, which is the honest argument for pricing models that make it rational: flat-rate dedicated environments such as OneSource Cloud's price the fleet as a whole, so the spare GPUs held against failure cost what any other GPU costs — no per-hour meter ticking on machines whose job is to wait.

FAQ

How much spare GPU capacity should a serving fleet hold?

Enough to cover simultaneous failures plus the repair window, computed not guessed: your failure rate times fleet size gives expected concurrent outages, the repair window extends how long each one bites, and the availability target decides how much of that bite users may feel — the N+1 pattern that sizing calculators implement is the common landing point, one spare per failure-capable group, and fleets with regional SLOs hold a second region's worth.

Which failover pattern should inference use?

By tier, not fleet-wide: stateless serving replicas take health-monitor-driven cordoning and replacement as the default; stateful tiers (caches, routers with connections) take active-passive with promotion; and only tiers whose pause violates the SLO justify active-active or a regional second — each step up the pattern ladder costs capacity you should spend only where the SLO actually demands it.

Is spare capacity the same as headroom?

No — they absorb different events and need separate budgets: headroom absorbs load spikes (your fleet running hot at peak), spare absorbs failures (a node vanishing at peak); a fleet with full headroom and no spare fails its SLO the moment hardware dies during a peak, and a fleet with spares but no headroom fails it on the traffic curve alone — size both, and let the flat-rate economics of dedicated capacity such as OneSource Cloud's make holding idle-for-failure GPUs rational rather than heresy.

Previous: Flat Rate Billing for AI GPU Cloud
Next: GPU Hardware Spares and RMA: Fleet Logistics and Downtime Risk
Related Articles