Inference is using a trained model to compute outputs from new inputs. This chapter covers the serving pipeline for a decoder-only autoregressive language model: how a request becomes generated text, how the next token is chosen, and how the service manages latency, memory and concurrent users.
An LLM produces scores for the next token. A decoding algorithm turns those scores into a choice. Ordinary inference does not update the model's learned weights. The application can separately retrieve information, call tools or update conversation state.
The most useful distinction is: temperature changes relative probabilities; top-k limits candidate count; top-p limits the candidate set by cumulative probability mass. They operate at each token-generation step, not once for the whole answer.
Table of Contents
- Generation Basics
- Prefill and Decode Phases
- Sampling Strategies
- Temperature: Reshape the Distribution
- Why Temperature Zero Does Not Divide by Zero
- Top-K: Limit the Candidate Count
- Top-P: Choose a Probability-Mass Cutoff
- Using the Controls Independently
- Using the Controls Together
- A Complete Sampling Implementation
- Tuning and Related Controls
- Stopping Conditions
- Speculative Decoding
- Latency Metrics
- Memory and Compute Requirements
- Continuous Batching and Prefix Caching
- Multi-LoRA Serving
- Streaming
- Production Considerations
- Interview exercise: serve a multi-tenant text assistant
- Interview Questions
- Final revision cards
- References
Generation Basics
For a decoder-only autoregressive model, the next-token distribution is conditional on the prompt and all tokens generated so far:
Here, is the prompt and is an output token. A token may be a word, part of a word, punctuation, or a special symbol.
- Tokenize the prompt using the model's tokenizer and chat template.
- Run the model to obtain next-token logits: one raw score per vocabulary entry.
- Apply any token constraints or penalties, then choose the next token.
- Append that token to the context. If it is a stop token, finish.
- Otherwise, run the next model step and choose again.
Sampling does not change the learned weights. It changes which continuation the system selects from those weights' predictions. Once a different token is selected, the next context changes, so later logits can diverge substantially.
Read diagram source
flowchart LR
R[Authenticated request] --> A[Authorize model and bound input/output]
A --> T[Tokenizer and chat template]
T --> Q[Admission and scheduling]
Q --> P[Prefill uncached prompt positions]
P --> L[Next-token logits]
L --> S[Constraints and token selection]
S --> E{Stop condition?}
E -->|No| O[Detokenize and deliver text delta]
O --> D[Decode selected token using KV state]
D --> L
E -->|Yes| F[Finalize permitted text, finish reason and usage]
This is the logical model loop; execution and delivery can overlap. A final ordinary token at an output limit can still be returned, while an EOS marker usually is not displayed. Stop-string handling may retain a suffix so prohibited trailing text is not emitted before the matcher recognizes it. See tokenization for the distinction between token IDs and displayed text.
Prefill and Decode Phases
| Phase | What the model processes | What it produces |
|---|---|---|
| Prefill | Prompt positions in parallel, with a causal attention mask | Cached keys/values for the prompt and logits used to select the first output token |
| Decode | The most recently selected token, using cached earlier keys/values | Its cache entries and logits used to select the next token |
Causal means a prompt position can attend to itself and earlier positions, never future positions. Parallel computation does not remove that mask.
Prefill often makes good use of large GPU matrix operations and can be compute-bound. Small-batch decode often becomes memory-bandwidth-bound because weights and an expanding KV cache must be read for relatively little work per step. These are workload-dependent tendencies: batch size, context length, architecture, and kernels can shift the bottleneck.
The KV cache avoids recomputing earlier positions' keys and values. It does not eliminate attention over those positions. Consequently, decode time can grow with context length; it is not universally constant. Dense prefill attention has a quadratic sequence-length component, although other model operations scale linearly in prompt length.
Sampling Strategies
Start with logits, not percentages
Use this deliberately tiny vocabulary throughout the examples:
| Token | Logit | Probability at |
|---|---|---|
| cat | 4 | 64.391% |
| dog | 3 | 23.688% |
| bird | 2 | 8.714% |
| car | 1 | 3.206% |
Logits are unrestricted scores, not probabilities. A negative logit is valid. Softmax converts them into positive probabilities that sum to one. The percentages above are rounded; calculations below use unrounded values.
Sampling means drawing one token with those probabilities. At , cat is selected about 64.4% of the time across repeated draws from this same context and distribution. A single draw can still select car. This is not a 64.4% probability that “cat” is factually correct.
Greedy Decoding
Greedy decoding selects the highest-scoring eligible token:
For our example, that is always cat. Softmax is unnecessary because it preserves the ordering of logits.
Greedy selection is deterministic for identical scores and a fixed tie-breaking rule. That does not guarantee identical results across different hardware, model revisions, or serving implementations. It also does not guarantee truth, valid JSON, or the most probable whole sequence: a locally best choice can lead to weaker later choices.
Temperature: Reshape the Distribution
For strictly positive temperature:
Temperature changes the gaps between scores before converting them to probabilities:
| Token | |||
|---|---|---|---|
| cat | 86.495% | 64.391% | 45.505% |
| dog | 11.706% | 23.688% | 27.600% |
| bird | 1.584% | 8.714% | 16.741% |
| car | 0.214% | 3.206% | 10.154% |
- : sharpens the distribution. At , logits become
[8, 6, 4, 2]; the favorite has a larger advantage. - : ordinary softmax; no temperature adjustment.
- : flattens the distribution. At , logits become
[2, 1.5, 1, 0.5]; alternatives gain probability.
A precise way to see the effect is the probability ratio:
Cat's score exceeds dog's by 1. Their probability ratio is at , at , and at .
Positive temperature preserves rank. Cat remains first at every positive temperature. Temperature alone does not intentionally discard tokens: with finite logits and exact arithmetic every token keeps positive probability. In implementations, masked logits of remain excluded and very small probabilities can underflow to zero.
As , the distribution approaches uniform over the finite, eligible logits. High temperature does not make better ideas more likely by definition; it gives lower-ranked alternatives more opportunity.
If you already have baseline probabilities , the equivalent transformation is:
Simply dividing probabilities by and renormalizing does nothing: the common factor cancels. Temperature must act on logits, or through the power transformation above.
Why Temperature Zero Does Not Divide by Zero
Substituting into is undefined. A sampler that supports zero must branch before performing the division:
The branch is easiest to remember as control-flow pseudocode; the complete Python sampler appears later in this chapter.
If temperature is zero:
select the eligible token with the largest logit
Otherwise, for positive temperature:
divide logits by temperature, apply the configured filters, and sample
This is an API convention, not a new rule of arithmetic. For example, vLLM documents zero as greedy sampling. Other interfaces expose a separate greedy mode or reject zero in their temperature processor. In Transformers, ordinary greedy generation uses do_sample=False, num_beams=1. Consult the engine rather than assuming all APIs accept the same settings. vLLM sampling parameters, Transformers generation configuration.
The limit explains the convention
Subtract the largest logit without changing softmax:
For a unique winning token, its numerator is . Every other token has , so its numerator tends to zero as . Therefore the winner's probability tends to one.
With [4, 3, 2, 1] and , cat already has about 99.99546% probability. It is still sampling; a positive temperature does not logically guarantee greedy behavior.
An implementation may also switch sufficiently small positive values into its greedy branch or enforce a minimum temperature. That is an engine convention, distinct from the mathematical limit. Inspect the effective settings when comparing deployments.
Tie caveat: with logits [4, 4, 2, 1], the limit splits probability equally between the two tied maxima. Greedy argmax instead chooses one according to its tie-breaking rule. The one-hot limit requires a unique maximum.
Subtracting the maximum also improves numerical stability by avoiding exponentiation of large positive values. It still does not make division by zero valid.
Top-K: Limit the Candidate Count
Top-k retains the highest-scoring candidates, removes the rest, and renormalizes before sampling.
For k=2 at :
| Token | Before filtering | After top-k and normalization |
|---|---|---|
| cat | 64.391% | 73.106% |
| dog | 23.688% | 26.894% |
| bird | 8.714% | 0% |
| car | 3.206% | 0% |
For example, cat becomes .
Renormalization means dividing each surviving probability by the total surviving mass so the new distribution sums to one. Filtering changes absolute probabilities but preserves probability ratios among survivors.
k=1leaves one token, so selection is effectively greedy, regardless of positive temperature.kequal to or larger than the vocabulary size imposes no extra restriction.- “Disabled” is an API setting, not a request to keep zero tokens. The example implementation below uses
None; engine conventions vary. - A strict rank implementation keeps exactly entries. Implementations based on a score threshold can keep extra tokens tied at the boundary.
Top-k uses a fixed count even when confidence changes. If the favorite has 98% probability, k=50 can still admit weak alternatives. If many continuations are reasonable, k=5 may be unnecessarily restrictive. Neither case means the model knows which tokens are actually correct.
Top-P: Choose a Probability-Mass Cutoff
Top-p, also called nucleus sampling, sorts candidates by probability and retains the smallest leading group whose cumulative probability is at least , where .
At with p=0.8:
| Rank | Token | Probability | Cumulative probability | Keep? |
|---|---|---|---|---|
| 1 | cat | 64.391% | 64.391% | Yes |
| 2 | dog | 23.688% | 88.080% | Yes: reaches 80% |
| 3 | bird | 8.714% | 96.794% | No |
| 4 | car | 3.206% | 100% | No |
The nucleus contains cat and dog. After renormalization their probabilities are 73.106% and 26.894%, just as for k=2 in this particular example.
Keep the token that reaches or crosses the threshold. The retained mass can exceed because tokens are indivisible. Top-p does not mean “keep each token whose individual probability exceeds ,” “keep percent of the vocabulary,” or “the final answer is correct with probability .”
Why the candidate count adapts
For p=0.9:
- Distribution
[0.98, 0.01, 0.01]: the nucleus has one token. - Ten equally likely tokens: the nucleus has nine tokens in exact arithmetic.
p=1 disables nucleus truncation. A small positive can leave only the best token. p=0 is outside the convention used here; do not rely on it as a portable greedy setting.
Nucleus sampling was introduced to remove an unreliable probability tail during open-ended generation. Its benefits depend on the task and model; it is not always superior to top-k. Holtzman et al., The Curious Case of Neural Text Degeneration.
Using the Controls Independently
Assume random sampling is enabled and all unmentioned filters or penalties are disabled:
| Experiment | Temperature | Top-k | Top-p | Operation |
|---|---|---|---|---|
| Baseline sampling | 1 | Disabled | 1 | Softmax, then draw |
| Temperature only | Positive | Disabled | 1 | Rescale logits, softmax, draw |
| Top-k only | 1 | 1 | Keep highest , normalize, draw | |
| Top-p only | 1 | Disabled | Softmax, keep nucleus, normalize, draw | |
| Greedy | Separate mode or supported zero convention | Irrelevant to ordinary greedy choice | Irrelevant to ordinary greedy choice | Argmax after applicable constraints/penalties |
“Temperature 1” does not disable sampling. It disables only temperature adjustment. Likewise, top_p=1 does not imply greedy selection.
Check the effective configuration: omitting a parameter may inherit a model-specific default instead of disabling it.
Using the Controls Together
Here is the explicit order used in this chapter's implementation. Engines can use a different order.
Read diagram source
flowchart TD
A[Next-token logits] --> B[Apply constraints and penalties]
B --> C{Greedy mode?}
C -->|Yes| D[Select argmax]
C -->|No| E[Apply positive temperature]
E --> F[Keep top-k if enabled]
F --> G[Softmax over survivors]
G --> H[Keep top-p nucleus if enabled]
H --> I[Renormalize and sample one token]
D --> J[Check stop condition]
I --> J
Worked example: all three controls
Start with logits [4, 3, 2, 1], then use T=2, k=3, p=0.8:
| Token | After temperature + softmax | After top-k, normalized | Final distribution after top-p |
|---|---|---|---|
| cat | 45.505% | 50.648% | 62.246% |
| dog | 27.600% | 30.720% | 37.754% |
| bird | 16.741% | 18.632% | 0% |
| car | 10.154% | 0% | 0% |
- Temperature flattens the distribution.
- Top-k removes car. The remaining mass is normalized to 100%.
- Within those three candidates, cat plus dog cover 81.368%, reaching
p=0.8. - Remove bird and normalize again. Sample cat with 62.246% probability or dog with 37.754% probability.
The final probabilities are neither the original probabilities nor a uniform choice between two tokens.
Temperature and top-p interact
Using top-p alone with p=0.8 on our four-token example:
| Temperature | Smallest nucleus reaching 80% |
|---|---|
| 0.5 | cat alone: 86.495% |
| 1 | cat + dog: 88.080% |
| 2 | cat + dog + bird: 89.846% |
Lower temperature can therefore make nucleus filtering much more restrictive. Increasing temperature cannot restore a token already removed by a filter.
Order matters: which distribution does top-p see?
Consider probabilities [0.60, 0.25, 0.10, 0.05], k=2, and p=0.7:
- Top-k first: keep the first two and normalize to approximately
[0.7059, 0.2941]. Top-p now needs only the first token. Final result:[1, 0, 0, 0]. - Top-p first: the original first token has only 0.60, so keep two to reach 0.85. Top-k removes nothing further. Final result: approximately
[0.7059, 0.2941, 0, 0].
These are different distributions. “Top-p 0.7” must be understood relative to the probabilities present at that stage.
Positive temperature and strict top-k commute for fixed logits and consistent tie-breaking: scaling preserves rank. Temperature and top-p generally do not commute because scaling can change how many tokens cross the cumulative threshold. A backend may choose to apply temperature after filtering; then temperature changes the survivors' odds but not that already-chosen nucleus. The order above is a defined example, not a universal law.
A Complete Sampling Implementation
This runnable, standard-library Python example makes the order and edge cases explicit. It handles one vocabulary vector, uses None to disable top-k, accepts -inf for masked tokens, and resolves score ties by original token index. It is teaching code; production GPU engines use optimized kernels.
import math
import random
def next_token_distribution(logits, temperature=1.0, top_k=None, top_p=1.0):
"""Temperature -> strict top-k -> normalized top-p -> final probabilities."""
scores = [float(z) for z in logits]
if not scores or any(math.isnan(z) or z == math.inf for z in scores):
raise ValueError("Use a nonempty vector with finite scores or -inf")
if not math.isfinite(temperature) or temperature < 0:
raise ValueError("temperature must be finite and >= 0")
if not math.isfinite(top_p) or not 0 < top_p <= 1:
raise ValueError("top_p must be in (0, 1]")
if top_k is not None and (type(top_k) is not int or top_k < 1):
raise ValueError("top_k must be None or a positive integer")
eligible = [i for i, z in enumerate(scores) if z != -math.inf]
if not eligible:
raise ValueError("At least one token must be eligible")
ranked = sorted(eligible, key=lambda i: (-scores[i], i))
result = [0.0] * len(scores)
if temperature == 0:
result[ranked[0]] = 1.0 # Branch BEFORE division.
return result
# Scaling by positive T preserves rank, so rank before scaling.
if top_k is not None:
ranked = ranked[:top_k]
maximum = scores[ranked[0]]
weights = [math.exp((scores[i] - maximum) / temperature) for i in ranked]
total = math.fsum(weights)
probabilities = [w / total for w in weights]
if top_p < 1:
cumulative = 0.0
for count, probability in enumerate(probabilities, start=1):
cumulative += probability
if cumulative >= top_p:
break # Includes the token that crosses the cutoff.
ranked = ranked[:count]
probabilities = probabilities[:count]
retained_mass = math.fsum(probabilities)
for i, probability in zip(ranked, probabilities):
result[i] = probability / retained_mass
return result
def sample_next_token(logits, temperature=1.0, top_k=None, top_p=1.0, rng=None):
probabilities = next_token_distribution(logits, temperature, top_k, top_p)
survivors = [i for i, p in enumerate(probabilities) if p > 0]
if len(survivors) == 1:
return survivors[0] # No random draw needed.
rng = rng if rng is not None else random
return rng.choices(range(len(probabilities)), weights=probabilities, k=1)[0]
logits = [4, 3, 2, 1]
print(next_token_distribution(logits, temperature=2, top_k=3, top_p=0.8))
# Approximately [0.622459, 0.377541, 0.0, 0.0]
print(sample_next_token(logits, temperature=0))
# 0: the token index for cat
Floating-point arithmetic can affect a cutoff at an exact boundary. Different tie policies, precision, and minimum-candidate rules can also explain small discrepancies across implementations.
Tuning and Related Controls
A practical tuning method
- Start with the model's documented generation configuration and a representative evaluation set.
- Hold filters fixed while varying temperature; assess correctness, diversity, repetition, and format validity.
- Then vary a truncation control while holding temperature fixed. Change both together only when you can explain the intended effect.
- Evaluate multiple outputs for stochastic configurations. One attractive answer is weak evidence.
- Record the model revision, prompt/template, effective settings, seed where supported, and engine version.
For a basic experiment, compare temperatures 0.5, 1, 2 with filters disabled; then compare top-p 0.8, 0.95, 1 at fixed temperature. These are illustrative experiments, not universal production recommendations. Model-specific reasoning behavior can make generic “low temperature for code” recipes unreliable.
Repetition, presence, and frequency penalties
These modify token scores based on history; they are distinct from temperature and truncation. A common multiplicative repetition penalty applies to previously seen tokens as follows:
Dividing a negative score by would increase it, accidentally rewarding the token. For example, -2 / 1.2 = -1.667 is higher than -2; multiplying gives -2.4, which lowers its probability relative to unchanged tokens.
An additive presence penalty subtracts a fixed amount if a token has appeared; a frequency penalty subtracts an amount proportional to its count. The exact history scope and formula are backend-specific. Excessive penalties can discourage legitimate repetition, such as a variable name in code.
Other controls worth distinguishing
| Control | Purpose |
|---|---|
| Random seed | Controls the pseudorandom stream; does not ensure reproducibility across all serving environments |
| Logit bias | Raises or lowers selected token scores |
| Grammar/schema constraints | Exclude tokens that would violate the allowed structure; still do not guarantee factual correctness |
| Min-p | Filters relative to the most probable token, rather than cumulative mass |
| Beam search | Keeps several partial sequences and scores continuations; is not the same as top-k token sampling |
| Output-token limit | Caps generation length; does not control next-token randomness |
Stopping Conditions
Generation can stop when an EOS/end-of-turn token is selected, a configured stop string is matched, the output-token budget is exhausted, or the request is cancelled or times out.
- EOS is a model/tokenizer-specific token, not necessarily a visible string.
- Stop strings can span multiple tokens or streamed chunks; a robust matcher retains enough trailing text to detect them.
- A token budget can end an answer mid-sentence or mid-JSON object. Check the returned finish reason. Provider output limits may include reasoning tokens as well as visible text; use the selected model’s contract.
- Prompt tokens plus generated tokens must fit the supported context policy; output budget alone does not describe total context use.
- Minimum-length constraints may temporarily mask EOS. Apply eligibility constraints before selecting a token, including in greedy mode.
Speculative Decoding
A draft mechanism proposes several tokens. The target model verifies them in parallel, potentially reducing expensive sequential target-model calls.
For exact speculative sampling, verification uses acceptance probabilities and a correction distribution on rejection. It is not simply “accept if both models' top tokens match.” With the appropriate algorithm, the target sampling distribution is preserved. Chen et al., Accelerating Large Language Model Decoding with Speculative Sampling.
The speedup depends on draft cost, acceptance rate, verification overhead, batch size, and hardware. Accepted tokens still require computation; there is no universal 2–3× gain. Draft models, additional prediction heads, and prompt lookup provide different ways to propose candidates.
For example, suppose ordinary target decoding takes 20 ms per new token. A speculative cycle spends 16 ms drafting and 24 ms verifying/correcting, and advances the sequence by an average of three tokens. Its rate is 40/3 ≈ 13.3 ms per token, or 1.5× faster for this measured workload. If it advances only one token, the same cycle takes 40 ms per token and is twice as slow. Count draft, verification, correction and cache-management work; acceptance rate alone is insufficient.
Exact sampling must verify the target distribution after the intended constraints and sampling transforms. “Lossless” refers to that algorithmic distribution, subject to implementation numerics; it does not mean arbitrary proposal acceptance preserves correctness. See the deeper speculative decoding chapter.
Latency Metrics
Time to First Token (TTFT)
TTFT is the time from request submission until the first output token arrives. It can include network transit, queuing, tokenization, prefill, first-token selection, and transport buffering. Always state the measurement boundary.
Tokens Per Second (TPS)
For a response with tokens, if is first-token arrival and is last-token arrival:
The corresponding average time per output token after the first is . Inspect inter-token latency percentiles as well: an average can hide visible stalls.
For zero- or one-token responses this rate is undefined. Client chunks are not necessarily individual tokens, and the first event may be metadata or a heartbeat. Record first text separately from first event, and use token-aware engine measurements when transport batching hides individual arrival times. For reasoning models, distinguish time to visible answer from any earlier internal generation.
Total Latency
With a constant average decode rate:
For 100 output tokens, TTFT of 0.2 seconds, and 50 tokens/second after the first: seconds. The first token is already counted in TTFT.
Throughput
Aggregate output tokens/second and requests/second measure serving capacity. Per-request TPS measures the experience of one user. Batching can improve aggregate throughput while reducing each user's token rate. Compare systems under the same prompt lengths, output lengths, concurrency, and latency targets.
Memory and Compute Requirements
Model Weights
Raw weight storage is approximately parameter count times bytes per parameter. A dense 70-billion-parameter model needs about 140 GB at FP16 or 35 GB at ideal packed INT4. Quantization scales, metadata, unquantized tensors, and runtime allocations add overhead. GB here means decimal gigabytes.
KV Cache
For ordinary full-context attention with a uniform cache layout:
The factors are keys plus values, layer count, KV head count, head dimension, and bytes per cache element. Multiply by the cached sequence length and number of requests for a simple batch estimate.
For an illustrative model with 80 layers, 8 KV heads, head dimension 128, and two-byte cache elements:
Per token = 2 × 80 × 8 × 128 × 2 = 327,680 bytes = 320 KiB
At 8,192 cached tokens = 2.5 GiB per request
For four such requests = 10 GiB of raw KV cache
Using 64 KV heads instead would require eight times as much cache: 20 GiB per request at that length. Do not substitute the number of query heads when the model uses grouped-query attention. Check the actual model configuration. Cache paging, quantization, sliding windows, and prefix sharing can change allocation requirements.
Total GPU Memory
Budget weights + KV cache + activations/workspaces + runtime overhead and headroom. Check how each is sharded or replicated across devices; summed VRAM alone is not proof that a deployment fits or performs well.
FLOPs per Token
A rough dense-model estimate for parameter matrix operations is floating-point operations per token, excluding additional context-dependent attention work:
P = 70 billion
2P = 140 billion FLOPs = 140 GFLOPs per token
At 40 tokens/second: about 5.6 TFLOPs/second for this component
140 billion FLOPs is 140 GFLOPs, not 140 TFLOPs. This operation estimate does not predict latency on its own: memory bandwidth, batch size, communication, and kernel utilization matter. Mixture-of-experts models also require distinguishing active parameters from total stored parameters.
Continuous Batching and Prefix Caching
Continuous batching schedules work at generation iterations: finished requests leave and waiting requests can enter without waiting for the longest response in a static batch. KV-cache availability, prompt processing, and scheduling policies constrain admission.
Prefix caching reuses cached state for an identical compatible token prefix, such as a repeated system prompt. It can reduce prefill work. It is not a cache of the final answer and it does not eliminate decoding. Tokenization, model/adapter identity, and relevant cache configuration must match; visually similar text is insufficient.
Measure cache-hit rate and latency under your workload instead of assuming a fixed percentage gain. vLLM automatic prefix caching.
PagedAttention maps a sequence's logical KV blocks to physical memory blocks so the whole sequence need not occupy one contiguous allocation. This reduces fragmentation and supports sharing; it does not remove the attention computation or make memory unlimited. The PagedAttention paper explains this distinction.
Chunked prefill splits a long prompt's processing across scheduling iterations so it can be interleaved with decoding. Smaller chunks can reduce stalls for existing streams while delaying the new prompt's first token. Disaggregated prefill/decode uses separate workers for the phases; it adds KV-transfer and routing costs. These are different choices from continuous batching and prefix reuse. Compare them under the required latency distribution, not a fixed speedup claim. See vLLM scheduling guidance.
Treat prefix state as private computation, not public content. Assign cache-sharing scope at an authenticated gateway; do not let an untrusted caller select another tenant's scope. Current vLLM supports a cache_salt to partition reuse. It can reduce cache-based timing leakage across trust groups, but it is not a replacement for authorization or broader isolation. vLLM prefix-cache design.
Multi-LoRA Serving
Multiple LoRA adapters can share one base model while contributing small, request-specific parameter updates. The server loads or schedules the appropriate adapter and must associate cached state with the correct model/adapter combination. Adapter memory, transfer time, compatible batch execution, and kernel support determine the practical concurrency limit.
Streaming
Streaming delivers incremental output, often using server-sent events. It improves perceived responsiveness but does not inherently make model computation faster. A streamed chunk may contain part of a token's decoded text or several tokens, depending on the API and buffering.
Handle cancellation, disconnects, backpressure, stop-string boundaries, partial structured output, and final usage/finish metadata. Do not treat the arrival of the first chunk as completion of the answer.
Production Considerations
- Scheduling: use admission control, bounded queues, and fair priorities; an endless high-priority workload should not starve other requests.
- Timeouts and cancellation: stop unnecessary generation and release request state promptly. Return a clear status for partial output.
- Fallbacks: evaluate any smaller-model or alternate-provider fallback against quality and format requirements.
- Observability: record TTFT, inter-token latency, throughput, errors, finish reasons, queue time, and cache utilization under realistic load.
- Cost: use actual input/output usage and the deployed model's applicable rates; distinguish cached-input or other billing categories when relevant.
- Quality: evaluate grounding, factuality, and schema validity separately from sampling settings. A sharper distribution can make an incorrect answer more consistent.
Follow one request through loading, streaming, and cancellation
Consider a request for tenant A using adapter finance-v7, with a five-second end-to-end deadline. The gateway authenticates the tenant, resolves an immutable base-model/adapter pair, estimates token demand, and admits the request only if its queue and budget allow completion. The worker tokenizes with the pinned chat template, checks compatible prefix state, performs prefill, and streams decode output. Record queue time, TTFT, inter-token latency, completion status, and actual usage separately.
If no first token arrives before the allowed fallback boundary, cancel the first attempt and confirm cancellation as far as the interface supports it before launching an evaluated fallback within the remaining deadline. After partial output has reached the user, silently splicing in another model's answer can produce contradictions. Prefer an explicit interrupted status or a clearly restarted answer. Closing the browser connection is not proof that an upstream provider stopped generation; propagate cancellation and account for any residual billed work.
An adapter cache miss follows a concrete sequence:
- Resolve the authorized adapter digest and verify its base-model compatibility and trusted artifact origin.
- Fetch bytes into a bounded CPU cache; validate the digest before use. Never let a user-supplied path load arbitrary model code.
- Reserve GPU adapter capacity. Evict only an adapter with no in-flight references, or wait within the deadline.
- Transfer the adapter tensors to GPU memory and complete the transfer before scheduling its request. Measure this cold-load delay separately from model computation.
- Bind the request to that adapter version for its lifetime. Prefix-cache identity must include all relevant model/adapter state; reuse across incompatible adapters changes results.
- On completion or cancellation, release the request's KV allocation and adapter reference. Retain reusable artifacts only under the cache's capacity and isolation policy.
These are application responsibilities around runtime-specific capabilities. The vLLM LoRA documentation describes adapter serving and limits; benchmark the exact configuration instead of assuming a fixed number of simultaneous adapters or a fixed latency penalty.
Keep administrative adapter-loading endpoints separate from customer generation endpoints. A trusted control plane can publish vetted versions; an ordinary user should select only an authorized model alias. Runtime support for loading a path is not permission to expose that operation to tenants.
Interview exercise: serve a multi-tenant text assistant
Prompt: “Design the inference service for an enterprise assistant used by several customer organizations.” Focus this exercise on serving authorized prompts; retrieval quality and business-tool execution are separate services with their own contracts.
1. Functional requirements
- Authenticate each request and authorize its tenant, model alias and optional adapter.
- Validate the model-specific chat template, input length, output allowance and supported generation settings.
- Stream text progressively with a clear completion, interruption or failure state.
- Support cancellation, remaining-deadline-aware recovery and per-tenant usage records.
- Release new model/adapter configurations through evaluation and rollback.
2. Non-functional requirements
- Latency: target p95 first visible text below 800 ms and p95 completion below 15 seconds for the agreed request class; validate both under peak load.
- Capacity: 100,000 requests/day over 16 active hours, with an eight-times-average peak and a 5% additional-attempt allowance.
- Input/output: mean 2,000 input and 250 output tokens; enforce maximums of 4,096 input and 512 output tokens for this class.
- Isolation: no unauthorized model, adapter, cache state, prompt or result crosses a tenant boundary.
- Availability: retain enough serving capacity for one replica failure; reject excess work clearly rather than building an unbounded queue.
- Cost and observability: track actual attempts, cache categories, completed requests and quality review; do not equate a valid stream with a useful answer.
A replica here means a complete serving deployment capable of answering independently. It may contain multiple GPUs. It does not mean a single accelerator or a single model process in every architecture.
3. Baseline and flaws
Start with one authenticated gateway and one model server. A bounded queue is sufficient for the first load test. Inspect failure behavior before adding more services.
| Flaw found in the baseline | Consequence | Change and cost |
|---|---|---|
| Admission counts requests only | A burst of long contexts exhausts KV capacity | Budget tokens and memory as well as request count; conservative limits can reduce utilization |
| Large prefills block existing decoding | Already-open streams visibly stall | Test chunked prefill; new requests may wait longer for their first token |
| All tenants share scheduling priority | One customer monopolizes capacity | Tenant quotas and fair scheduling; less ability to borrow spare capacity without rules |
| The cache key ignores adapters or trust scope | Incompatible state or cross-tenant timing exposure | Bind model/adapter identity and cache-sharing scope; fewer reusable entries |
| Client disconnect does not reach the engine | Unnecessary GPU work and possible provider charges | Propagate cancellation and settle residual usage; cancellation may be best-effort upstream |
| Every timeout launches a second model | Duplicate work and incoherent partial answers | Bound attempts, use remaining deadlines and make restarts explicit |
| Logs store every prompt forever | Sensitive content accumulates unnecessarily | Minimize/redact records and apply retention; debugging may require approved samples |
4. Detailed serving architecture
Read diagram source
flowchart TB
C[Client] --> G[Gateway: identity, model policy, deadline]
G --> A[Admission: tenant quotas, token and spend limits]
A --> Q[Bounded fair scheduler]
Q --> R[Replica router]
R --> S[Serving replica: prefill and decode]
REG[Trusted model and adapter registry] --> S
S <--> K[(Scoped KV and prefix blocks)]
S --> V[Output checks and stream adapter]
V --> C
G --> X[Cancellation controller]
X --> Q
X --> S
S --> M[Usage, latency and failure telemetry]
V --> M
M --> P[Capacity and release decisions]
Use immutable model and adapter versions for each request. Route with cache locality when it helps, but balance that against queue length and load; sending every shared prefix to one hot replica can erase the cache benefit. A new replica is ready only after required artifacts and kernels are loaded and health checks pass.
Output policy determines buffering. A high-risk structured answer may require complete validation before release, while a lower-risk conversational answer can stream checked segments. Partial JSON is not a completed validated object. Streaming itself does not make output safe.
If the model server restarts, either resume through a deliberately implemented recovery protocol or tell the client the attempt was interrupted. SSE delivery does not, by itself, make model state durable or support exactly-once display after reconnection.
5. Capacity arithmetic
Assume an eight-second mean active service time for the measured request mix. These are planning inputs, not specifications for a named model.
| Estimate | Calculation | Result |
|---|---|---|
| Average original arrival rate | 100,000 / (16 × 3,600) | 1.74 requests/s |
| Peak rate including extra attempts | 1.74 × 8 × 1.05, using unrounded values | 14.58 requests/s |
| Input-token arrival at peak | 14.58 × 2,000, using unrounded values | 29,167 tokens/s |
| Output-token production at peak | 14.58 × 250, using unrounded values | 3,646 tokens/s |
| Mean active requests at peak | 14.58 × 8, using unrounded values | About 117 |
Suppose one complete replica sustains four requests/second at the required latency for this same input/output distribution. At a chosen 70% planning utilization, budget 2.8 requests/second per replica. ceil(14.58 / 2.8) = 6 serving replicas; keep seven to retain that six-replica capacity after one fails. Correlated host or zone failures need a separate placement and recovery plan.
Do not call the 70% rule a universal latency guarantee. Bursts, long-tail sequence lengths, adapter loads and KV pressure require load testing. A throughput number measured with 32-token outputs cannot size this 250-token workload without another benchmark.
Memory admission must also respect the maximum request. Using the earlier illustrative 80-layer, 8-KV-head layout, 4,096 input plus 512 reserved output tokens require up to 1.40625 GiB of raw KV capacity per request. Forty-eight such reservations total 67.5 GiB, before weights and workspaces. Sharing or incremental allocation may reduce typical usage, but the scheduler must prevent future growth from exhausting memory.
6. Compare complete cost
Assume thirty days/month: 3 million original requests. The 250-token mean represents all billable output in this example; a reasoning model can require a different allowance. These rates are editable interview assumptions, not current vendor quotes. Compare equivalent quality, privacy and service levels before deciding between hosted and owned serving.
Hosted option: input $0.50/million tokens, cached input $0.05/million, output $2/million. Assume 30% of input tokens qualify for cached billing, and the extra 5% of attempts have the same usage mix.
| Monthly cost | Calculation | Estimate |
|---|---|---|
| Uncached input | 4.2 billion / 1 million × $0.50 | $2,100 |
| Cached input | 1.8 billion / 1 million × $0.05 | $90 |
| Output | 750 million / 1 million × $2 | $1,500 |
| Additional attempts | 5% × $3,690 | $184.50 |
| Gateway, logs and application infrastructure | Assumed allocation | $600 |
| Maintenance | Assumed allocation | $1,500 |
| Quality review | 1% × 3 million × 1 minute / 60 × $35/hour | $17,500 |
| Total | Listed costs | $23,474.50 |
That is $7.82 per 1,000 original requests. If 97% satisfy the useful-completion criteria, cost is $8.07 per 1,000 useful completions. The 500 review hours/month are substantial even though the model bill is small. Review rate is an assumption to validate, not a standard staffing rule.
Owned option: assume the seven complete replicas cost $6/hour each, all day, for 720 hours/month: $30,240. Add the same $600 application, $1,500 maintenance and $17,500 review allowances to get $49,840, or $16.61 per 1,000 original requests. There is no additional token-API charge for inference done on these owned replicas. Hardware rates, maintenance effort, billing for idle capacity and model quality can change the comparison; the equal operational allowances are only a simplifying assumption.
The hosted option is cheaper under these inputs. Owned serving can still be justified by deployment control, data requirements, custom kernels or sustained utilization. Test those requirements and the provider's actual quotas instead of declaring either approach universally better.
7. Close the exercise
“I would begin with the evaluated serving option that meets the required privacy and latency. I would use token-aware admission, fair scheduling, immutable model/adapter identities and explicit stream completion. I would scale from measured prompt/output distributions and inspect both queue and KV pressure. My next experiments would test long-context bursts, a replica failure, cancellation after partial output and the cost of quality review.”
Interview Questions
Q: How do temperature, top-p, and top-k differ?
Strong answer: Temperature rescales logits and changes relative probabilities. Top-k keeps a fixed maximum candidate count; top-p keeps the smallest highest-probability group reaching a cumulative threshold. We renormalize and sample among survivors. They act on every next-token distribution. Their composition depends on order, especially on the distribution used for top-p.
Q: Why does temperature zero not cause division by zero?
Strong answer: The temperature-softmax formula requires positive temperature. An API supporting zero implements a separate greedy argmax branch. With a unique maximum, the positive-temperature distribution converges to that token as temperature approaches zero. Exact ties split the mathematical limit, whereas greedy decoding uses a tie rule.
Q: Can top-p and top-k produce the same result?
Yes, when they keep the same candidates at the same temperature. For logits [4, 3, 2, 1] at , both k=2 alone and p=0.8 alone retain cat and dog. They need not match at the next token position because the distribution changes.
Q: Does increasing temperature change the top-k set?
For positive temperature and a fixed set of logits, no: rank is preserved. It changes probabilities within that set. It can change the top-p set if nucleus selection happens after temperature scaling.
Q: Why can two engines disagree with identical numeric settings?
They can differ in processing order, default filters, tie handling, zero-temperature conventions, precision, model/template revisions, or random-number generation. Inspect effective configuration and implementation behavior rather than comparing just three numbers.
Q: Explain prefill versus decode and their performance implications.
Prefill processes prompt positions in parallel under a causal mask and produces the first next-token distribution. Decode processes newly selected tokens using earlier cached keys/values. Prefill is often compute-heavy; small-batch decode is often bandwidth-limited. Prompt length, cache length, batching, and scheduling determine which optimizations help.
Q: How do you estimate serving capacity?
Estimate weight and KV-cache memory using the actual KV head count and cache dtype, then add runtime headroom. Benchmark the intended hardware with representative input/output lengths and concurrency. Report both per-request latency and aggregate throughput; fitting in memory alone does not establish capacity.
Q: Can speculative decoding reduce latency while preserving sampling behavior?
Yes, with the appropriate target-distribution acceptance and correction algorithm. Measure the total work per advanced token. Low acceptance or an expensive draft can make the system slower, and an arbitrary “accept a plausible token” rule does not preserve the target distribution.
Q: A prefix-cache hit made the first token faster, but the full answer barely changed. Is the cache broken?
Not necessarily. Prefix reuse saves compatible prefill work; it does not remove generation of a long answer. Compare saved prompt work with decode time, transport buffering and any other latency components.
Q: Why is a chunk timestamp not automatically an inter-token timestamp?
The transport may group tokens, split text, buffer output or emit non-text events. Define first visible text and completion at the client, and obtain token-level timing from an engine that exposes it. Report unavailable measurements as unavailable.
Q: A 70B model's weights fit across two GPUs. Is that enough evidence to deploy?
No. Add per-device KV allocation, workspaces, runtime memory and headroom, then verify how tensors are sharded or replicated. Benchmark communication and service-level performance. Aggregate VRAM is only an initial feasibility check.
Q: Should all tenants with the same system prompt share prefix-cache entries?
Only within the intended trust and compatibility scope. Shared timing can reveal information about previous requests. Assign scope in the authenticated service, keep adapter identity in cache matching, and assess whether additional isolation is required.
Q: What should happen after a client disconnects?
Propagate cancellation, stop queued work, release safe-to-release request state and settle actual usage. A remote provider may continue briefly or expose no hard cancellation guarantee. A closed client socket does not establish that computation or billing stopped.
Q: Why can increasing batch size improve throughput while hurting the user experience?
More requests share weight loads and use the accelerator efficiently, but each iteration can take longer and queueing can increase. Compare aggregate work per second with per-request TTFT and inter-token latency at the required load.
Q: Can lowering temperature fix hallucinations or guarantee valid structured output?
No. It concentrates probability on higher-scoring choices, including incorrect ones. Grounding, authorized evidence, schema constraints and application validation address different failure modes. Do not substitute sampling settings for those controls.
Final revision cards
| Concept | Recall statement | Interview mistake to avoid |
|---|---|---|
| Inference | Apply trained parameters to new inputs | Calling ordinary generation another training step |
| Prefill / decode | Process known prompt positions / extend from selected tokens | Removing the causal mask or assuming decode cost is constant |
| Temperature | Rescale scores for positive temperature | Dividing by zero or interpreting token probability as truth |
| Top-k / top-p | Candidate count / cumulative mass | Confusing nucleus mass with vocabulary percentage |
| Speculation | Propose, verify and correct | Counting accepted tokens while ignoring drafting overhead |
| TTFT / TPS | First token boundary / subsequent token rate | Counting the first token twice or timing metadata as text |
| KV sizing | 2 × layers × KV heads × head dimension × bytes | Substituting query heads for KV heads in GQA |
| Continuous batching | Admit work at scheduling iterations | Promising a fixed speedup for every workload |
| Prefix caching | Reuse compatible prompt state | Treating it as an answer cache or a substitute for authorization |
| Serving cost | Count full operations per useful completion | Comparing only weight fit or token price |
Practice tip: explain one sampling calculation, one memory estimate and one overloaded-request trace without naming a vendor. Then use current documentation to show how your chosen runtime implements those ideas.
References
- Holtzman et al.: The Curious Case of Neural Text Degeneration — nucleus sampling.
- Transformers generation configuration — sampling, greedy mode, and related parameters.
- vLLM sampling parameters — engine-specific parameter conventions.
- Chen et al.: Accelerating Large Language Model Decoding with Speculative Sampling — distribution-preserving speculative sampling.
- Kwon et al.: Efficient Memory Management for Large Language Model Serving with PagedAttention — serving and KV-cache management.
- vLLM automatic prefix caching — reuse of matching prompt prefixes.
Previous: Embeddings and Vector Spaces | Next: Model Taxonomy