LLM inference latency

LLM inference latency monitoring is an observability practice that measures and explains the time a model-serving request spends in each stage from admission through token delivery. A single end-to-en