inference capacity

H100 inference capacity is a workload envelope that defines how much concurrent LLM serving work the system sustains within memory, throughput, and latency objectives. Context length changes that capa