Not every LLM call needs an answer now. Document scoring, dataset labeling, archive re-processing, nightly enrichment — a large share of enterprise inference is delay-tolerant, and treating it like chat traffic is the most expensive default in the stack: interactive serving idles GPUs between requests while the batch alternative fills them. This page gives the deciding axis, the cost mechanics behind the difference, and the routing rule that settles each workload in one question.
The Modes: One Axis Decides
One axis separates the modes — how long the answer may wait: real-time inference answers within the interaction (chat, search assist, live decisions), batch inference collects delay-tolerant work (document scoring, dataset labeling, nightly enrichment) and processes it in utilization-filled windows — and the practitioner rule is exactly that blunt: if business outcomes depend on immediate action, real-time; if insights can wait, batch is more cost-efficient.
| Mode | Latency envelope | Typical workloads |
| Real-time | Sub-second to seconds | Chat, search assist, live decisions, copilots |
| Batch, minutes-tier | Minutes with a freshness SLA | Email triage, assistive scoring, moderate queues |
| Batch, overnight-tier | Hours to a nightly window | Corpus enrichment, labeling, archive re-scoring |

The tiers inside batch matter as much as the batch-realtime split: minutes-tier work carries a freshness promise someone has to own, while overnight work costs nothing the business notices — and the routing differences between the two show up in scheduling, monitoring, and what breaks at 2 a.m.
The Tradeoffs: Utilization Is the Lever
The cost difference is mechanical: interactive traffic idles GPUs between requests — and a GPU at 10% utilization costs roughly ten times more per useful token than one at 100% — while batch fills the machine, so providers price batch tiers at steep discounts and private fleets see the same arithmetic in their own utilization graphs; the tricky parts are composition (batches grouped by similar cost and length minimize padding waste, not raw count) and engine reality (an online engine with continuous batching can beat an offline batch mode on identical hardware, so mode labels are not performance guarantees).
- The idle tax: utilization is the cost lever — ten percent utilization means roughly ten times the cost per useful token, whether a provider's meter or your own fleet's depreciation is counting.
- Batch-tier pricing: providers discount delay-tolerant tiers steeply because filled GPUs are cheaper per token — the market passes the utilization arithmetic through.
- Composition matters: similar-cost batching — grouping requests by compute cost and length — minimizes padding waste and matters as much as model size.
- Label caution: offline batch mode is not automatically faster than a well-batched online engine on the same hardware; measure the engine, not the label.
The utilization framing also explains when batching saves little: traffic that is already continuous and uniform fills the GPUs without help, and the batch win there is thin. The big savings sit exactly where most enterprises have them — bursty interactive traffic sharing a fleet with mountains of delay-tolerant work.
The Conditional Verdict: Route by Delay Tolerance
Route by delay tolerance with a written rule: sub-second interactive products go real-time, end of story; work measured in minutes (email triage, assistive scoring) goes batch with a freshness SLA; work measured in hours or overnight (corpus enrichment, labeling, re-scoring archives) goes batch windows without ceremony — and the estate lands hybrid, both modes sharing the same models and capacity, with the routing rule written down so the next workload's placement is a lookup rather than a debate.
- Write the rule: one line per tier — delay tolerance, mode, freshness SLA if any — posted where workload owners can read it.
- Route on intake: every new workload answers the staleness question before it answers the capacity question.
- Review the mix quarterly: workloads drift toward interactivity as products evolve; the routing rule absorbs it or gets amended.
The hybrid estate is also where capacity economics get simple: dedicated flat-rate environments such as OneSource Cloud's let the overnight batch windows burn hours you already own, so the batch savings land as utilization on your own fleet rather than as a discount someone else grants.
FAQ
How do I decide between batch and real-time inference?
Ask one question per workload: how stale can the answer get before anyone is harmed? Sub-second staleness harms the product — real-time; minutes cost a freshness promise — batch with an SLA; overnight costs nothing the business notices — batch windows; the axis is delay tolerance, everything else (cost, utilization, pricing tiers) follows from where the workload lands on it.
How much does batch inference actually save?
On the order of the utilization gap plus the provider's pass-through: providers discount batch tiers steeply (fifty-percent discounts are common in the market) because filled GPUs are cheaper per token, and private fleets see the same arithmetic without the label — a fleet whose interactive traffic runs GPUs at 10% and whose batch windows run them full is paying multiples per useful token for the interactive share; measure your own utilization curve and the savings number falls out of it.
Can a pipeline mix both modes?
Yes, and most enterprises should: the same models and capacity serve interactive traffic in real-time while delay-tolerant scoring runs in batch windows — with the caveat that mode labels are not engine guarantees (online engines with continuous batching can beat offline modes on the same hardware), so route by delay tolerance and then benchmark the engine mode that serves each tier; dedicated flat-rate capacity such as OneSource Cloud's makes the mixed estate simple to cost, since the batch windows burn hours you already own.