hybrid inference capacity planning

Spot and dedicated GPU capacity are two ways to source inference compute: dedicated capacity is reserved and continuously available with predictable latency, while spot capacity is acquired at a lower