LLM autoscaling
-
Autoscaling for LLM Inference Serving and Cold Starts
LLM autoscaling should watch queue depth and KV-cache pressure, not GPU busy percent. Plan warm pool
- 1
LLM autoscaling should watch queue depth and KV-cache pressure, not GPU busy percent. Plan warm pool