The latency ticket arrives, the budget does not. That constraint is more workable than it feels: most serving stacks leave large latency gains on the table in software — batching policy, cache behavior, precision, scheduling — before a single accelerator is added, and the performance-engineering literature says so explicitly, noting that scaling to more GPUs has diminishing returns for latency. The skill is selection: which technique moves the latency you actually measured, at what cost, in what order. This page is the selector.
The Families: What Exists Before Hardware
Five families cut latency on the GPUs you already own: continuous batching (requests share weights in flight, the systems-level optimization with the largest documented load-condition gains), KV-cache management (prefix reuse and compression moving TTFT and memory pressure), quantization (lower precision moving per-token time), request scheduling (ordering and admission moving tail latency), and speculative decoding (draft-then-verify moving end-to-end generation) — a menu the technique guides converge on from different directions.
| Family | Mechanism | Documented character |
| Continuous batching | Requests join and leave batches in flight | Largest load-condition gains; config-level in modern engines |
| KV-cache management | Prefix reuse, compression, memory control | Moves TTFT and memory-bound behavior |
| Quantization | Lower-precision arithmetic | Moves per-token time; quality needs verification |
| Request scheduling | Ordering, admission, class policy | Moves tail latency under mixed load |
| Speculative decoding | Draft fast, verify in parallel | Moves generation time; acceptance-rate dependent |

The families are individually well documented — batching improves utilization because concurrent requests share the same model weights, quantization and speculative decoding cut latency without hardware upgrades in practitioner guides — what most teams lack is the map from symptom to family.
The Tradeoffs: Segment, Throughput, Quality, Effort
The families differ in which latency they move and what they charge: batching and scheduling buy tail-latency stability under load at little quality cost but demand systems work; KV management moves TTFT with memory as its currency; quantization moves per-token time and pays in quality risk that needs per-workload verification; speculative decoding moves generation time but pays in acceptance-rate dependency and extra compute — and hardware scaling itself shows diminishing returns for latency, which is why the software families come first.
- Segment mapping: p95 under load lives with batching and scheduling; TTFT lives with KV management; per-token time lives with quantization; end-to-end generation lives with speculative decoding.
- Quality column: quantization and speculative decoding carry real quality risk — verify on your traffic, not vendor claims.
- Throughput interaction: some latency cuts cost tokens per second; know which before you trade.
- Effort column: batching policy is config-level; scheduling policy is engineering; speculative decoding is architecture.
The segment map is the selector's core: a team that measures slow TTFT and responds with quantization has treated the wrong segment, and the fix that finally lands is a benchmark that shows where the milliseconds actually live.
The Conditional Verdict: Sequence by Symptom
Sequence by symptom: p95 blowout under load starts with continuous batching and scheduling admission; slow first tokens start with KV management and prompt-prefix reuse; slow generation starts with quantization verified on your traffic, then speculative decoding; all-of-the-above starts with the cheap config wins before the architecture work — and measure each step against the same fixed benchmark, because latency work without a fixed yardstick optimizes the anecdotes.
- Fix the measurement first: one benchmark, fixed traffic, fixed metrics (TTFT, per-token, p95) — the yardstick every step is judged against.
- Take the config wins: continuous batching on, sensible scheduling defaults — cheap, reversible, usually the largest single gain under load.
- Match the family to the remaining symptom: KV work for TTFT, quantization for per-token (with quality verification), speculative decoding for generation-bound paths.
- Re-measure per step: keep what the yardstick approves, roll back what it does not, and only then discuss hardware — with the benchmark record as evidence of what software already paid for.
The sequence ends, when it must, at capacity — and at that point the fixed benchmark does double duty as the capacity case: the same numbers that exhausted the software families are the workload profile a dedicated serving environment such as OneSource Cloud's is sized against.
FAQ
What is the first latency technique to try?
Whichever is cheap in your stack and moves your measured symptom: most serving deployments start with continuous batching (config-level in modern engines, largest load-condition gains) plus scheduling admission if p95 under load is the complaint; if the complaint is time-to-first-token, start with KV and prefix reuse before touching anything heavier — the first move is always the one that addresses the segment you actually measured slow.
Do latency techniques stack, or do they conflict?
They stack with order and interaction: batching and scheduling compose cleanly, KV compression interacts with prefix reuse (compressed caches reuse differently), quantization shifts the arithmetic speculative decoding drafts against, and the composed result is not the product of the parts — so introduce families one at a time against a fixed benchmark, keeping whichever step the numbers approve and rolling back whichever they do not.
When is buying more GPUs actually the answer?
After the software families are exhausted and the constraint is proven compute: if batching is continuous, scheduling is sane, KV is managed, and quality-safe quantization is in, yet the SLO misses at peak — the capacity is genuinely short, and the honest options are more accelerators or a smaller model, because performance engineering has diminishing returns documented even for tensor-parallel scaling.