continuous batching

Continuous batching — also called iteration-level batching — admits and evicts LLM inference requests mid-generation rather than waiting for an entire batch to finish, keeping the GPU saturated as var