LLM inference cost

LLM inference cost is a total-cost measure that captures the infrastructure and operations required to deliver accepted model requests at a defined quality, latency, and availability level. GPU price