Batching groups requests or token work for more efficient execution. Continuous batching changes the active request set at generation scheduling boundaries as requests finish and new work can be admitted. It improves opportunities for utilization; it does not remove queueing or guarantee good latency at unlimited load.
Distinguish three scheduling choices
| Strategy | Membership | Useful setting | Tradeoff |
|---|---|---|---|
| Static batch | Fixed for a batch execution | Known offline work or compatible fixed shapes | Finished sequences may leave unused capacity |
| Dynamic request batching | Collect arrivals within a bounded window | Amortizing execution across nearby requests | Waiting window adds latency |
| Continuous batching | Admit/remove work between generation iterations | Variable-length autoregressive requests | Requires state, memory, and token-budget scheduling |
Inputs need not have identical original lengths; padding, bucketing, or variable-length kernels can handle differences. Returning a completed response early is separate from whether its execution slot can immediately be reused.
Trace the benefit with four requests
Assume a simplified two-slot decoder and requests requiring 2, 6, 2, and 2 output steps. Ignore prefill and admission overhead for this illustration.
| Step range | Fixed batch A+B | Continuous schedule |
|---|---|---|
| 1–2 | A and B active | A and B active |
| 3–4 | A's slot unused, B active | C replaces A; B active |
| 5–6 | A's slot unused, B active | D replaces C; B active |
| 7–8 | C and D active | All complete |
Continuous membership finishes this idealized work in six steps rather than eight. Real requests consume different cache and prompt-processing resources, so “two slots” is a teaching simplification. A scheduler needs token and memory budgets, not only a request count.
Read diagram source
flowchart LR
A[Arrivals] --> B[Bounded queue with deadlines]
B --> C[Admission and token budget]
C --> D[Prefill and decode scheduling]
D --> E[Execute iteration]
E --> F[Complete or cancel requests]
F --> C
In-Flight Batching
“In-flight batching” is terminology used by some runtimes for scheduling work while other requests are already running. Supported implementations can mix prefill and decode work. Check the engine's exact scheduling semantics rather than treating the name as a distinct universal algorithm.
Prefill and decode share accelerator resources. Mixing them may improve utilization, but a large prompt can also delay tokens for existing streams. It is not guaranteed simultaneous use of otherwise idle compute and memory bandwidth.
Chunked Prefill
Chunked prefill divides prompt processing into smaller pieces that the scheduler can interleave with other work. Suppose a long prompt needs 32,000 input tokens processed while several short chats are decoding. Processing all of it at once may create a visible pause. Smaller chunks create more opportunities to serve ongoing decodes, at the possible cost of longer TTFT for the long request.
The current vLLM optimization guide documents chunked-prefill and token-budget behavior. Defaults and feature combinations are version-specific. Record the runtime release and measure the configured policy rather than assuming every engine uses the same priority order.
| Tuning choice | Potential benefit | Regression to watch |
|---|---|---|
| Smaller prefill chunks | Shorter interference with active decoding | More scheduling overhead and slower new long requests |
| Larger token batch budget | More work per iteration | Inter-token latency and cache pressure |
| Larger admission queue | Absorb a brief burst | Stale requests and long tail waits |
| Separate workload pools | Isolate long/batch work | Lower sharing efficiency and extra reserved capacity |
| Tenant-aware quotas | Limit a flooding customer | Requires trusted identity and consistent accounting |
Failure, repair, and load testing
Continuous batching can keep a GPU busy even while the service is overloaded. Growing queue age, missed deadlines, and client retries are evidence of an unhealthy operating point. Bound queue length and waiting time, propagate cancellation, and reject excess load with a retry policy that avoids amplification.
Test a mix of prompt lengths, output lengths, arrivals, and cancellations. Include cold caches and bursts. Compare throughput while meeting latency and success targets. A benchmark that reports only peak tokens per second can reward configurations users would find unusable.
Tip: Keep application routing separate from runtime scheduling. A gateway can enforce quotas, but it does not automatically control which sequence an inference engine advances at each iteration.
Interview practice
- Why does continuous batching help variable output lengths? Finished work can release resources for queued requests before the longest original sequence finishes.
- Does dynamic batching mean continuous batching? Not necessarily. Dynamic collection can still execute a fixed set once formed.
- What does chunked prefill trade? It can reduce disruption to active decodes while changing the new request's TTFT and scheduling overhead.
- Why is batch size alone insufficient? Prompt length, retained cache, output budget, and model shape determine resource use.
- What happens when offered load exceeds capacity? Queueing grows unless the system rejects, defers, or scales work; scheduling cannot create unlimited capacity.
- How do you test fairness? Flood one tenant with long prompts while measuring other tenants' latency, admission, and completion rates.
Recall card and closing
Arrivals → admission → token budget → iteration → completion. Explain who waits, which resource is bounded, and how the policy behaves during a burst. Continue to PagedAttention for cache allocation.