Multi-Model GPU Serving: Capacity Planning When Models Share a Fleet

NoraLin 10 2026-10-08 22:28:04 Edit

One model per fleet was the easy era. Modern AI platforms serve families of models — size variants, fine-tunes, task-specific specialists, embedding models — on shared GPU fleets, and the capacity question changes shape entirely: the constraint stops being total compute and becomes which models are resident where, at which moment, with whose headroom. This page is the planning method for that shape.

Prerequisites: Memory Budgets Per Model

One constraint frames everything — memory: each model's resident footprint (weights plus activation and KV headroom at your batch profile) set against the cards that must host it, because several models sharing one GPU cannot all stay loaded at once, and hands-on practice makes exactly this the core problem; the prerequisite inventory is a table of models × footprints × traffic, from which every placement decision follows.

Inventory columnWhat it capturesWhy it decides placement
Resident footprintWeights + activations + KV headroom at profileHow many models fit a card together
Traffic per modelRequests per second, burst profileWhich models earn permanent residency
Load timeMinutes to load from storageThe swap penalty for long-tail models
VarianceHow spiky the model's usage isHeadroom the co-resident set must hold

The inventory is the whole planning discipline in miniature: research on serving heterogeneous models frames the problem as exploiting memory capacity across co-located models, and none of the strategies below work without the numbers it contains.

Plan the Fleet: Pack, Swap, and Schedule

Three strategies compose: model packing (co-resident small models within a memory budget, the documented pattern for fleets of small models), dynamic loading and unloading (swapping models on demand at the cost of load latency, the managed-endpoint pattern), and two-level scheduling — global placement deciding which models live on which nodes while local allocation manages memory within each — with research on cost-efficient multi-LLM serving validating exactly that two-level structure; per-model headroom lands inside the resident set, sized to KV growth and batch variance, because the aggregate, not any single model, is what fills the card.

  1. Tier by traffic: hot models earn permanent residency; the long tail swaps on demand — the traffic table makes the split mechanical.
  2. Pack the hot set: co-resident models sized so the aggregate footprint plus per-model headroom fits the card with room for the spikiest member.
  3. Place at two levels: global placement maps models to nodes; local allocation manages memory within each — the structure multi-LLM serving research validates.
  4. Write the swap policy: which models may evict which, load-time budgets, and the anti-thrash rule that stops two popular models from evicting each other.

Orchestration platforms matter at exactly this layer: schedulers such as the OnePlus AI Orchestration Platform handle placement, quota, and queueing across a fleet as their core function, which makes the two-level structure an operated feature rather than a home-built scheduler.

Verify: Test Under the Request Mix

Verification runs the production request mix — the traffic pattern across models, not the model list, decides whether the plan holds: replay a representative mix (which models get hit, at what concurrency, with what request sizes) and watch for the failure modes the plan promised to prevent: load storms when a popular model cold-loads, headroom breaches when two spiky models co-reside, and queueing where placement forced hot models to share — because a plan that holds per-model but fails the mix is a plan for a fleet you do not operate.

  • Mix fidelity: real request distribution across models, real concurrency, real sizes — synthetic uniform traffic proves nothing.
  • Load-storm watch: popularity shifts triggering mass cold-loads are the plan's signature failure.
  • Headroom breach watch: co-resident spiky models breaching together — the aggregate rule being tested.
  • Re-run on change: new model, new traffic pattern — the mix test follows the fleet.

This is also where dedicated capacity closes its argument: a reserved fleet whose memory the platform plans deliberately — the model OneSource Cloud operates on — turns placement from a nightly negotiation with noisy neighbors into an owned design, and the mix test becomes confirmation rather than discovery.

FAQ

Should models be packed onto shared GPUs or swapped on demand?

By tier: high-traffic models earn permanent residency (packed or dedicated), long-tail models swap on demand and eat the load latency only when their rare requests arrive — the mistake is uniform policy, either packing everything (memory fills with rarely-used weights) or swapping everything (load storms on every popularity shift); the traffic table from step one makes the tier split mechanical.

How much headroom does each model need on a shared card?

Enough for its own variance, inside a set that fits: KV state grows with context and batch, so headroom is planned per model at your traffic profile — and the binding rule is the resident set's aggregate never approaches the card, because two spiky models sharing a card breach together and take each other's headroom down; fleet-level headroom policy beats per-model heroics.

How does model routing interact with multi-model capacity?

Routing chooses the model per request; capacity decides whether that model is resident — so the router's policy either respects placement (route only to loaded models, with a load-and-wait path for the long tail) or fights it (routing to cold models and manufacturing load storms), which is why the routing table and the placement plan get written by the same team, referencing the same traffic model.

Previous: AI Orchestration: Streamline GPU Operations and Scale AI
Related Articles