Learnastra AI SYSTEM DESIGNAnup Rai

Concept · Understand the mechanism

Batching: schedule useful work without hiding the wait

By Anup Rai4 min readReviewed September 2026

Batching groups requests or token work for more efficient execution. Continuous batching changes the active request set at generation scheduling boundaries as requests finish and new work can be admitted. It improves opportunities for utilization; it does not remove queueing or guarantee good latency at unlimited load.

Distinguish three scheduling choices

Strategy Membership Useful setting Tradeoff
Static batch Fixed for a batch execution Known offline work or compatible fixed shapes Finished sequences may leave unused capacity
Dynamic request batching Collect arrivals within a bounded window Amortizing execution across nearby requests Waiting window adds latency
Continuous batching Admit/remove work between generation iterations Variable-length autoregressive requests Requires state, memory, and token-budget scheduling

Inputs need not have identical original lengths; padding, bucketing, or variable-length kernels can handle differences. Returning a completed response early is separate from whether its execution slot can immediately be reused.

Trace the benefit with four requests

Assume a simplified two-slot decoder and requests requiring 2, 6, 2, and 2 output steps. Ignore prefill and admission overhead for this illustration.

Step range Fixed batch A+B Continuous schedule
1–2 A and B active A and B active
3–4 A's slot unused, B active C replaces A; B active
5–6 A's slot unused, B active D replaces C; B active
7–8 C and D active All complete

Continuous membership finishes this idealized work in six steps rather than eight. Real requests consume different cache and prompt-processing resources, so “two slots” is a teaching simplification. A scheduler needs token and memory budgets, not only a request count.

Architecture / visual model
flowchart LR A[Arrivals] --> B[Bounded queue with deadlines] B --> C[Admission and token budget] C --> D[Prefill and decode scheduling] D --> E[Execute iteration] E --> F[Complete or cancel requests] F --> C
Read diagram source
flowchart LR
  A[Arrivals] --> B[Bounded queue with deadlines]
  B --> C[Admission and token budget]
  C --> D[Prefill and decode scheduling]
  D --> E[Execute iteration]
  E --> F[Complete or cancel requests]
  F --> C

In-Flight Batching

“In-flight batching” is terminology used by some runtimes for scheduling work while other requests are already running. Supported implementations can mix prefill and decode work. Check the engine's exact scheduling semantics rather than treating the name as a distinct universal algorithm.

Prefill and decode share accelerator resources. Mixing them may improve utilization, but a large prompt can also delay tokens for existing streams. It is not guaranteed simultaneous use of otherwise idle compute and memory bandwidth.

Chunked Prefill

Chunked prefill divides prompt processing into smaller pieces that the scheduler can interleave with other work. Suppose a long prompt needs 32,000 input tokens processed while several short chats are decoding. Processing all of it at once may create a visible pause. Smaller chunks create more opportunities to serve ongoing decodes, at the possible cost of longer TTFT for the long request.

The current vLLM optimization guide documents chunked-prefill and token-budget behavior. Defaults and feature combinations are version-specific. Record the runtime release and measure the configured policy rather than assuming every engine uses the same priority order.

Tuning choice Potential benefit Regression to watch
Smaller prefill chunks Shorter interference with active decoding More scheduling overhead and slower new long requests
Larger token batch budget More work per iteration Inter-token latency and cache pressure
Larger admission queue Absorb a brief burst Stale requests and long tail waits
Separate workload pools Isolate long/batch work Lower sharing efficiency and extra reserved capacity
Tenant-aware quotas Limit a flooding customer Requires trusted identity and consistent accounting

Failure, repair, and load testing

Continuous batching can keep a GPU busy even while the service is overloaded. Growing queue age, missed deadlines, and client retries are evidence of an unhealthy operating point. Bound queue length and waiting time, propagate cancellation, and reject excess load with a retry policy that avoids amplification.

Test a mix of prompt lengths, output lengths, arrivals, and cancellations. Include cold caches and bursts. Compare throughput while meeting latency and success targets. A benchmark that reports only peak tokens per second can reward configurations users would find unusable.

Tip: Keep application routing separate from runtime scheduling. A gateway can enforce quotas, but it does not automatically control which sequence an inference engine advances at each iteration.

Interview practice

  1. Why does continuous batching help variable output lengths? Finished work can release resources for queued requests before the longest original sequence finishes.
  2. Does dynamic batching mean continuous batching? Not necessarily. Dynamic collection can still execute a fixed set once formed.
  3. What does chunked prefill trade? It can reduce disruption to active decodes while changing the new request's TTFT and scheduling overhead.
  4. Why is batch size alone insufficient? Prompt length, retained cache, output budget, and model shape determine resource use.
  5. What happens when offered load exceeds capacity? Queueing grows unless the system rejects, defers, or scales work; scheduling cannot create unlimited capacity.
  6. How do you test fairness? Flood one tenant with long prompts while measuring other tenants' latency, admission, and completion rates.

Recall card and closing

Arrivals → admission → token budget → iteration → completion. Explain who waits, which resource is bounded, and how the policy behaves during a burst. Continue to PagedAttention for cache allocation.

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← Speculative decoding: propose cheaply, verify correctly
NEXT LESSONPagedAttention: allocate the cache as the sequence grows →

Explore the diagram