System-design interview · Extended interviews
Design an LLM inference platform
Design token-based admission, prefill and decode scheduling, GPU and KV-cache budgets, streaming, cancellation and tenant fairness.
You will learn to
- Explain the separate prompt-processing and token-generation phases.
- Calculate token throughput and attention-memory demand instead of sizing by QPS alone.
- Design fair admission, private prefix reuse, cancellation, and model-version rollout.
Practice in this chapter
8 interview questions with model answers and follow-ups.
Go to interview practiceUseful foundations: Capacity estimation: throughput, latency, concurrency and storage · Message queues, event logs, delivery guarantees, and backpressure · Authentication, authorization, and tenant isolation
Workload and timing examples are interview assumptions.
Dotted concept links open the relevant explanation in a new tab.
01Problem and scope
An LLM inference platform serves versioned models within latency, memory and throughput budgets. Request count alone is insufficient: a 4,000-token prompt with 600 requested output tokens consumes different prefill, decode and KV-cache resources from a short interactive request. Accept a request only when its token and memory costs fit the available budget. Specify which streamed events a client can resume, and return a defined overload error when the budget is exhausted. Training, fine-tuning orchestration and execution of generated tool calls are outside this serving scope.
Define the serving metrics and state
A token is a unit produced by the model's tokenizer, often a word fragment. Prefill processes input tokens and builds attention state. Decode repeatedly produces the next output token using that state and prior output. Time to first token (TTFT) is the delay from request arrival to the first output token; inter-token latency (ITL) is the delay between successive output tokens. A KV cache stores attention keys and values so subsequent decoding does not recompute the entire prefix from scratch.
Clarify training, tools and model scope
Candidate: “Do we train models, execute generated tools, or only serve text?” Interviewer: “Serve versioned text models, stream tokens, support tenants and cancellation.” Candidate: “I will start with one worker and bounded admission, then measure mixed prefill/decode load before adding replicas or parallelism.” Training, fine-tuning orchestration and autonomous tool execution are outside scope. A model-produced tool-call proposal is output data; another authorized application decides whether to execute it.
02Functional requirements
- Admit a generation. Record one generation for the tenant/request identity and reserve its allowed work budget.
- Stream ordered events. Events have generation ID and increasing sequence; no mixed attempts.
- Retry a request. Return existing generation/status for same payload identity.
- Cancel a generation. Persist intent and remove the sequence from future scheduling promptly.
- Change model versions. New requests use an explicit version; active requests keep their pinned version.
- Account for usage. Record cumulative usage counters durably without counting repeated reports twice; state who pays for work completed after the last saved counter if the worker crashes.
Generation lifecycle
A tenant submits a request with an immutable model version, input and maximum output tokens. The service validates limits and either admits a bounded generation or rejects it before promising unlimited queued work. The client receives ordered stream events and a final reason such as completed, output-limit, canceled or interrupted. It can inspect status and cancel; cancellation stops future work at a safe execution boundary rather than undoing already delivered tokens.
Constraints and exclusions
Specify reconnect behavior. This design retains a bounded stream buffer for reconnect while its worker lives; after worker loss, the generation is interrupted unless output events were separately persisted under a stronger tier. Reusing a request ID does not make stochastic computation reproduce identical text. The service does not silently restart and concatenate a new answer after an old partial stream. Tenant quotas include input/output token budgets, concurrency and queue work, not merely request count. Slow clients cannot accumulate an unbounded per-stream buffer in server memory.
03Non-functional requirements
- Interactive latency. One-second p95 time to first token (TTFT) and 50 ms p95 inter-token latency (ITL) for the defined interactive mix. Also measure completion time and successful-output rate.
- Admission and queue bound. Targets apply to admitted requests under tested limits, not arbitrary million-token prompts. Budget queue time; reject or offer a batch tier when predicted waiting consumes most of the TTFT allowance.
- Availability. Target 99.9% availability for authenticated API requests within the documented request and workload limits; continue service after one worker/zone failure using reserved capacity.
- Control durability. Replicate request identity, terminal status and usage so they survive the declared zone failure.
- Execution limits. Bound context length, generated length and total active KV blocks.
- Regional recovery. Include compatible weights, tokenizer/runtime and warm capacity in the recovery plan; cold model reload time contributes to the recovery objective.
Execution invariants
| Boundary | Required guarantee |
|---|---|
| Generation ownership | Only the worker assigned the current execution epoch may publish current results; the epoch is a version number that changes when ownership changes |
| Client stream | Never merge events from different attempts into one apparent stream |
| Prefix reuse | Never share private prefix state across incompatible model or tenant contexts |
| Cancellation | Do not free memory still referenced by an in-flight kernel |
| Model quality | Pin tested weight/tokenizer/template settings; stochastic output can still vary |
Ephemeral execution and measured capacity
Live GPU KV state is ephemeral here: a worker crash interrupts its active streams even though control metadata survives. Hardware, precision, batch mix and model architecture determine capacity. This chapter uses illustrative measurements, not named-GPU performance claims.
04Capacity estimates
At fifty generated tokens/s per stream, 600 tokens take about twelve seconds. Stable arrivals at 100/s imply roughly 100 × 12 = 1,200 active decoders before queued/prefill requests. Longer outputs occupy slots and memory longer even if request QPS stays unchanged.
The KV calculation counts the stored attention state for each token across the model’s layers. A KV head contributes a key vector and a value vector; the head dimension is the number of elements in each vector. Multiplying those counts by bytes per element gives storage per token. Use the stored KV-head count, which need not equal the query-head count.
For an illustrative attention architecture with 32 layers, eight KV heads, head dimension 128 and two-byte elements, KV bytes per token are 2 × 32 × 8 × 128 × 2 = 131,072, or 128 KiB. The initial factor two is for keys and values. A full 4,600-token sequence needs about 575 MiB; 1,200 such fully grown sequences would need about 674 GiB. Average active length is lower, and architectures/parallelism/quantization change the allocation. Weights, activations, communication buffers and runtime reserve are additional.
| Resource illustration | Result | Admission consequence |
|---|---|---|
| 20 GiB usable KV pool | 20 GiB / 575 MiB ≈ 35 maximum-length sequences |
Cannot admit unlimited concurrent contexts |
| 2,000-token prefix at 128 KiB/token | 250 MiB cached state | Reuse saves compute but occupies real memory |
| 100 canceled streams, 575 MiB each | Up to about 56 GiB released eventually | Cancellation latency affects useful capacity |
These estimates explain token and block budgets before choosing a serving framework.
05APIs and contracts
Effective input and request identity
Request A sends POST /v1/generations with {"requestId":"generation-81","model":"summarizer-v4","input":"...","maxOutputTokens":600,"stream":true}. Tenant identity comes from authentication. The gateway applies the exact versioned prompt template and tokenizer before computing limits; counting only the visible user text would miss system/tool-format overhead. A reused request ID with changed effective input/settings returns 409.
Generation interfaces
| Interface | Contract |
|---|---|
| Generation response | generationId:g81, queued/running state and ordered token events |
| Stream event | {generationId:g81,epoch:7,sequence:121,text:"..."} |
| Final event | Terminal reason, durably recorded billable usage and pinned model/template versions |
DELETE /v1/generations/g81 |
Idempotent cancellation intent; terminal state remains inspectable |
GET /v1/generations/g81 |
Current state and declared reconnect/interruption behavior |
GET /v1/models |
Authorized immutable versions and supported limits |
Errors, resume and usage semantics
Reject invalid settings/context with 400, tenant exhaustion with 429 and unavailable compatible capacity with 503. Retry hints include jitter expectations, and batch requests may use a different latency tier. A bounded retained stream buffer can resume from an event sequence only while those events remain available; if the cursor expired, report it. Do not promise replay of unpersisted output after a worker crash. Usage semantics must distinguish processed input, generated output, delivered output and cached-input discounts if any; these are product accounting choices, not inferred from packet count.
Durable-watermark billing
A usage watermark is a saved cumulative count, such as 100 output tokens processed so far. Later reports advance that count; receiving the same report twice must not double the charge. The crash tail is work completed after the last saved report and lost from accounting when a worker fails.
- Define billable work. For this design, billing uses only durably recorded cumulative work watermarks, not all physically executed work.
- Advance durable counters. Each attempt periodically reports cumulative input/output work; the accounting authority advances each counter monotonically for its (generationId, epoch) and deduplicates event identity.
- Flush final usage. On normal completion or cancellation, flush and acknowledge the final watermark before sending a final usage event.
- Account for the crash tail. If the worker crashes after token 120 while only 100 are durably recorded, the unreported 20-token tail is an unbilled internal cost.
- Disclose interrupted usage. The interrupted result labels its recorded usage accordingly.
- State the throughput/accounting tradeoff. This avoids a durable round trip per token; it accepts some underbilling rather than claiming exact crash-proof accounting from an asynchronous stream.
Seal terminal accounting
- Seal at terminalization. When the attempt completes, is canceled or is declared interrupted, the accounting transaction marks its billable totals final. This is called sealing the attempt.
- Record final totals and release reservation. It atomically records the final billable watermarks and releases unused reservation; later stale reports cannot reopen or increase that sealed invoice.
- Accept pre-seal reports. Before sealing, a delayed authenticated report can advance a watermark monotonically.
- Reconcile post-seal work internally. After an interrupted attempt is sealed, a late usage report contributes only to the operator’s estimate of actual compute consumed; it cannot add a new user charge.
- Keep the crash-tail policy stable. This makes the stated unbilled crash-tail policy stable across delayed messages.
- Separate physical recompute from logical billing. Here internal preemption/recompute and retransmitted stream events are not new billable logical tokens; report that physical work separately for capacity analysis.
06Data model and access patterns
Durable entities and registry
- Generation identity.
Generation(tenantId,requestId,generationId,payloadHash,modelVersion,state,executionEpoch,cancelRequested,budgetReservation)owns request identity. - Execution attempt.
Attempt(generationId,epoch,workerId,startedAt,terminalReason)identifies execution. - Usage identity.
UsageEvent(generationId,epoch,eventSequence,cumulativeInput,cumulativeOutput)has a unique identity. - Accounting rule. The accounting authority takes monotonic per-attempt watermarks, rather than adding cumulative counters as if they were independent deltas.
- Compatible model version.
ModelVersionrecords weight checksum, tokenizer, template, adapter and runtime compatibility. - Worker registry. A worker registry advertises loaded compatible versions, health and approximate available token/block capacity.
Ephemeral worker state and epochs
A KV block is a fixed-capacity allocation for cached key/value elements. A block table tells the worker which physical blocks hold a sequence’s logical token positions. This separation lets a growing sequence use available blocks without requiring one large contiguous allocation; the worker must still track every live reference before reusing a block.
Worker-local state contains tokenized input, active sequence positions, KV block tables, scheduler queues and bounded stream buffers. This state is not made durable just by saving a Generation row. On worker failure, the control plane marks the attempt interrupted and either leaves retry to the client or creates an explicitly separate attempt under a documented policy. A time-limited ownership lease and its epoch number let the control store reject terminal-state updates from a replaced worker; the stream gateway also rejects output carrying an old epoch.
Prefix identity and model artifacts
A prefix is the leading token sequence shared by two requests, such as the same instructions and source document. Prefix caching retains the prefill state computed for those tokens so a compatible request can start from it. It reuses prior input computation; it does not supply the new request’s generated answer.
Prefix cache keys incorporate exact effective tokens, compatible model/weight/tokenizer/adapter settings and a server-controlled tenant or trust-group scope. A prefix cache reuses internal tensors; an answer cache returns text and requires a different correctness policy. Raw private prompts and tensors are not shared through an unscoped key. Durable weights live in an artifact store with checksums and access control; loading a similarly named model without checking version compatibility would break both output consistency and cache safety.
07Basic working design
One model and one worker
Start with one authenticated API and one inference worker hosting a model that fits its hardware. Tokenize request A's effective prompt, validate the 4,000-plus-600 token bound, and admit only if the queue and memory budget can support it. The worker performs prefill, then decodes tokens one iteration at a time and streams numbered events. Request B waits behind request A in a simple first-in queue; that is inefficient but understandable.
Lifecycle, cancellation and worker loss
The API records g81 before scheduling and returns a terminal outcome only when completion/cancellation/interruption is known. A client disconnect triggers cancellation under a stated grace policy. The worker releases KV memory after it is no longer used by execution, and the accounting path bills durably recorded logical token work rather than maximum reserved tokens; unrecorded crash-tail work remains an internal cost. If the worker crashes, g81 is interrupted; the service does not claim its saved metadata can recreate the lost KV cache.
Measure before adding complexity
For a small internal service this can be the right starting point. Load-test prompt/output distributions, measure TTFT and ITL separately, and retain a small set of known quality prompts. A faster throughput benchmark that allows multi-second token gaps does not validate the interactive target. The baseline gives us the data needed to justify batching, prefix reuse or multiple workers, instead of guessing capacity from the GPU's memory size alone.
The baseline measures real prefill/decode behavior, streams tokens and reports interruption honestly if its worker fails.
Read each connection in order
- sync1. Generate / cancelRequest A and request B clients → Authenticated generation API
- sync2. Reserve request identityAuthenticated generation API → Generation control records
- sync3. Admit bounded token workAuthenticated generation API → Single inference worker and KV memory
- control4. Load compatible modelVersioned model artifacts → Single inference worker and KV memory
- sync5. Stream ordered tokensSingle inference worker and KV memory → Request A and request B clients
08Find the baseline flaws
| Failure test | What breaks and what must follow |
|---|---|
| Long prefill blocks short requests | Request A's 4,000-token prefill monopolizes the worker while request B's short question waits. A fixed batch can create another inefficiency: all requests start together, but short completions leave empty slots until the longest finishes if the scheduler cannot add new work. The problem is scheduling at the token-iteration level, not simply insufficient HTTP threads. |
| KV allocation and utilization | Suppose the worker has 20 GiB available for KV after weights/reserve. Admitting 100 sequences that can grow to 575 MiB requires about 56 GiB, exceeding the pool even though all input requests fit in CPU memory. “There are only 100 requests” is not a memory estimate. Paged allocation reduces waste but cannot make those live tokens free. The service needs an explicit strategy: conservative reservation, preemption/recomputation, bounded swapping or rejection, with corresponding latency consequences. |
| Prefix privacy and abandoned computation | A third counterexample is cross-tenant prefix reuse. If a private document's cached prefix makes a guessed request noticeably faster for another tenant, timing can reveal information about cache residency. Raw tensor bytes need not be returned for a side channel to matter. Cache scope must follow trusted identity, not a client-chosen salt that can impersonate another tenant. Finally, canceling only the HTTP connection leaves abandoned generation consuming GPU blocks unless cancellation reaches the execution scheduler. |
09Improve the design, step by step
1. Add token-based admission and fair bounded queues
- Trigger: Long-prompt overload triggers quotas on input/output work, context length and concurrent reserved blocks.
- Mechanism: Weighted tenant scheduling and queue deadlines protect interactive users.
- Benefit, cost and alternative: This improves predictable TTFT and isolation but rejects some requests that an unbounded queue would accept and later time out. Separate batch queues are preferable for workloads that can trade delay for utilization; request-QPS-only limits remain insufficient.
2. Use continuous batching with chunked prefill
- Trigger: Completed requests leave unused slots in a fixed batch, while a long prompt can delay tokens for existing streams.
- Mechanism: Between iterations, replace completed requests with new ones and process long prompts in bounded pieces. Existing streams keep opportunities to decode while spare capacity handles new inputs.
- Benefit, cost and alternative: Benefits are higher utilization and smoother output; costs are scheduling overhead, tuning and possible slower TTFT for long prompts. A simpler fixed batch suits offline homogeneous jobs. vLLM documents these mechanisms, but configuration must match the tested model/workload.
3. Manage KV memory in blocks and reuse authorized prefixes
- Trigger: Fragmentation and repeated common prompts trigger paged allocation plus compatible prefix caching.
- Mechanism: As sequences grow, the worker allocates blocks and counts which active sequences still reference each block. Matching prefixes within an authorized cache scope avoid repeated prefill.
- Benefit, cost and alternative: Benefits are less wasted memory and input computation. Costs include metadata, eviction, recomputation and security scope. Full worst-case reservation is simpler but may waste capacity; unscoped reuse is rejected because it crosses privacy boundaries.
4. Add compatible replicas and controlled model parallelism
- Trigger: Aggregate demand or model size triggers more workers.
- Mechanism and tradeoff: Replicas scale independent requests when the model fits; tensor parallelism splits layer computation across devices when needed, adding communication. Separate prefill/decode pools are a later measured alternative, with large KV transfers and new failure modes. Use separate pools only when measured benefits justify that transfer cost. Every rollout pins immutable versions and drains active streams before retiring old workers.
Each step is judged against both useful token throughput and latency, not GPU utilization alone.
10Detailed architecture
Admission and compatible routing
The gateway authenticates, applies the pinned template/tokenizer, validates budgets and records generation identity in a replicated control store. Admission selects an allowed latency tier and reserves tenant work. The router selects a worker with the right model using readiness and estimated capacity. The worker must then reserve actual KV blocks: its registry report may already be out of date.
Worker-owned scheduling and memory
Each inference worker owns its scheduler, active sequences, KV block manager, isolated prefix cache and stream output. A worker may be one device or a coordinated group running parts of the same model. State whether losing one device interrupts the whole group, and count memory across that group. Weights/tokenizer artifacts are checksum-verified before readiness. The diagram keeps prefix cache inside this execution boundary rather than presenting it as a generic shared answer cache.
Stream and accounting boundaries
A stream gateway forwards only events with the generation's current epoch and bounds slow-consumer buffers. It propagates cancellation and disconnect policy to the owning scheduler. Usage events flow asynchronously to an idempotent accounting aggregator; critical state transitions update the control authority. Metrics report queueing, TTFT, ITL, memory and fairness independently.
Independent scaling and rollout
The main synchronous path is admission through token production, while model deployment and usage aggregation are background work. Control-store replication protects request/status identity but does not checkpoint GPU tensors. A stale worker must not publish current terminal state, and the scheduler must prevent canceled work from starting another iteration before returning its blocks to the reuse pool.
The worker owns actual KV allocation. Durable control state is distinct from ephemeral tensors and stream buffers.
Read each connection in order
- sync1. Submit generation-81 / cancel g81Tenant clients → Authenticated API and tokenizer
- sync2. Reserve / cancel generation under epochAuthenticated API and tokenizer → Replicated generation / epoch authority
- sync3. Admit token/memory budgetAuthenticated API and tokenizer → Token admission and fair queues
- async4. Schedule compatible workToken admission and fair queues → Compatible-model worker router
- sync5. Assign work; check local capacityCompatible-model worker router → Inference scheduler and execution
- sync5b. Claim current generation epochInference scheduler and execution → Replicated generation / epoch authority
- sync6. Reserve / reuse isolated blocksInference scheduler and execution → KV block manager and scoped prefix cache
- sync7. Emit g81 epoch 7 eventsInference scheduler and execution → Epoch-checked stream gateway
- sync7b. Check current generation epochEpoch-checked stream gateway → Replicated generation / epoch authority
- sync8. Stream bounded ordered outputEpoch-checked stream gateway → Tenant clients
- control9. Propagate current cancellationAuthenticated API and tokenizer → Inference scheduler and execution
- sync10. Guard terminal state by current epochInference scheduler and execution → Replicated generation / epoch authority
- control11. Load checked model versionImmutable weights and tokenizer store → Inference scheduler and execution
- control12. Route only ready versionsModel readiness and rollout control → Compatible-model worker router
- async13. Save cumulative token-usage countsInference scheduler and execution → Idempotent usage aggregation
- async14. Finalize budget usageIdempotent usage aggregation → Replicated generation / epoch authority
- async15. Report TTFT, ITL and blocksInference scheduler and execution → Latency, memory and fairness metrics
11Write path and acknowledgement
Before scheduling, reserve the allowed prompt/output work and memory and fix the model version. Bill only from usage counts saved under the declared durable-watermark policy.
Numbered admission and generation trace
- Tokenize the exact effective request. Request A authenticates under tenant t9. The API applies summarizer-v4's exact template and tokenizer, measures 4,000 input tokens, validates maxOutputTokens 600, and hashes the effective request settings.
- Reserve one logical generation. Atomically reserve generation-81 as g81 with its tenant budget and execution policy. A duplicate identity returns the same g81; a different payload conflicts. Admission checks queue deadline and predicted token/memory demand.
- Reserve worker-local capacity. Route to a ready compatible worker. The worker atomically reserves its local sequence/block budget before acknowledging execution epoch 7, so simultaneous gateway decisions cannot overcommit the same remaining slots.
- Reuse only compatible scoped prefixes. Look up a compatible prefix under t9's server-controlled cache scope. Reuse only valid blocks and increment references; otherwise schedule prefill. A cache hit changes work, not the model version or allowed output limit.
- Interleave prefill and decode. Process request A's prefill in bounded chunks interleaved with decoding for existing requests. Request B's short request can enter later iterations instead of waiting for a whole fixed batch to finish.
- Stream ordered output and durable usage. Decode outputs, assign stream sequence numbers and send events tagged g81/epoch 7. Stop at model end, output limit, deadline or cancellation. Periodically report cumulative work watermarks with unique usage-event identity; the durable authority advances counters monotonically.
- Drain execution and finalize accounting. On terminal state, stop scheduling, wait for in-flight execution to release references, free/reuse eligible blocks, flush the final usage watermark, and finalize the durable result. A normal final usage event waits for that durable acknowledgement; a crash instead reports the last recorded watermark as interrupted usage. Unused budget reservation is released under the accounting policy.
Registry estimate versus atomic reservation
Two routers may both see the same free memory. The worker’s atomic reservation lets only one claim that remaining capacity.
12Read and delivery path
Streaming obeys backpressure and disconnect policy. Cancellation stops new scheduling before freeing state still used by in-flight GPU work.
Numbered stream and cancellation flow
- Authorize ordered stream delivery. The client of request A subscribes to g81 and receives ordered text events. The stream gateway verifies tenant ownership and execution epoch before forwarding them. Client rendering handles event repetition by sequence where reconnect buffering permits it.
- Resume only retained events. The client can inspect status without creating another generation. If it reconnects within retained buffer limits, it asks after its last event sequence; otherwise it receives an explicit expired/interrupted outcome.
- Observe interactive fairness. The client of request B measures TTFT separately from ITL. The scheduler's fairness policy should keep the short interactive request from waiting behind an unbounded queue of long-document requests.
- Persist cancellation intent. The client cancels request A after token 120. The API durably marks cancelRequested and notifies epoch 7's worker. A repeated cancel is harmless; canceling an already completed request reports its terminal state rather than pretending output was undone.
- Stop scheduling before freeing memory. The scheduler observes cancellation at its next safe boundary, prevents new decode/prefill work for g81 and marks its stream canceled. It waits until in-flight kernels no longer reference blocks before freeing them. Some already-computed events may have been in transit; the client knows the cancellation boundary is not retroactive erasure.
- Finalize usage and unused reservation. Actual usage is finalized according to the declared policy, and reserved-but-unused work is released. A bounded slow-consumer policy can similarly pause briefly or cancel instead of allowing unlimited stream-buffer growth.
Worker-loss interruption contract
If a worker dies after token 120, the service reports interrupted. A new model attempt might produce a different continuation even with the same high-level question, so it is not silently appended under g81's old event sequence. Durable output replay or exact continuation would require additional checkpoint/state guarantees beyond this design.
13Correctness deep dive
Paged KV ownership
Paged KV allocation manages fixed-size blocks with ownership/reference counts. It avoids reserving one contiguous region for the maximum possible sequence. Smaller blocks reduce unusable gaps between allocations, called external fragmentation; unused slots inside a partially filled final block are internal fragmentation. It does not reduce the number of logical attention values required for distinct live tokens. Prefix reuse lets compatible requests refer to already computed blocks, while later divergent tokens allocate separate blocks.
Arrows from two requests converge on the same prefix blocks, then lead to different suffix blocks.
Remember: Same prefix can share memory; different continuations need their own state.
Read the diagram
- Trace shared and request-specific KV block ownership.
- Requests A and B reference compatible prefix blocks P1 and P2.
- Each owns different suffix blocks; reuse and release must respect isolation and in-flight GPU work.
Try from memoryCan request A free prefix block P1 as soon as A finishes?
Not if B or in-flight work still uses it. Shared ownership must be accounted for before recycling the block.
Memory transition table
| Event | Required enforcement | Result |
|---|---|---|
| Admit g81 | Worker scheduler reserves within block/token limits | No double admission of the same free capacity |
| Match prefix | Exact compatible key and trusted tenant scope | Increment references to reusable blocks |
| Cancel g81 | Mark sequence unschedulable for current epoch | No future iterations are added |
| Kernel still in flight | Keep references until execution completion | Blocks cannot be reused prematurely |
| Release last reference | Block manager observes zero live references | Return block to pool or permitted prefix cache |
Cancellation race
Prefix security boundary
Prefill savings and memory cost
Prefix reuse primarily saves prefill; six hundred new output tokens still require decode work. An answer cache is separate and must include task permissions, source freshness and generation settings. Neither cache should be described as a proof of deterministic output or universal protection from every hardware side channel.
Shared-prefix copy-on-write
Shared prefix blocks are read-only while referenced by several sequences. A sequence that must append into a shared partial block first allocates and copies a private block, or the implementation shares only complete immutable blocks. It must not write new KV entries into another sequence's shared state. The block manager orders reference updates, eviction and allocation so they cannot race. Eviction releases the cache’s claim, but a block remains allocated while a computation still uses it.
Canonical cache identity
Cache identity uses a collision-resistant hash over canonical model/tokens/scope data, with validated compatibility metadata. A fast unverified hash collision must not substitute another prompt's tensors. Different attention layouts, quantization formats or adapters can change compatibility even if displayed model names match. Treat a serving framework's cache-salt and hash options as version-tested configuration, not a claim that default settings meet every tenant boundary.
A block still referenced by an in-flight kernel cannot be reused for another sequence, even after the client cancels.
Read each connection in order
- syncCancel g81 after token 120Request A client → API authority
- syncPersist cancelRequested for epoch 7API authority → API authority
- syncCancel current g81 epoch 7API authority → Worker scheduler
- syncMark sequence unschedulableWorker scheduler → Worker scheduler
- syncWait for in-flight iteration boundaryWorker scheduler → Execution engine
- returnNo active references for g81Execution engine → Worker scheduler
- syncRelease g81 block referencesWorker scheduler → KV block manager
- returnReuse only zero-reference blocksKV block manager → Worker scheduler
- syncSave final usage; finalize canceled resultWorker scheduler → API authority
- returnCanceled; prior output retainedAPI authority → Request A client
14Failure and recovery
| Failure or condition | Surviving state, response and recovery |
|---|---|
| Worker crash | Worker crash: Active KV state and unpersisted stream buffers disappear. The control plane marks epoch 7 interrupted after its lease/health failure is established. The client of request A retains whatever text it already received and can start an explicit new attempt. Model weights reload from durable artifacts; warm compatible replicas absorb new work within reserved capacity. The durable request and previously recorded usage survive. Physical work after the last usage watermark can be lost from accounting; under our declared policy that crash tail is unbilled, not fabricated as an exact count. |
| Control-authority partition | Control authority partition: A minority cannot create new generation identities or safely change ownership. Existing admitted workers may continue within their bounded lease/policy, but terminal state and cancellation propagation require reconciliation. Epoch checks prevent a stale worker and replacement from both presenting one continuous current stream. In this design, workers stop scheduling new iterations and stream gateways stop admitting further events when their bounded execution/forwarding lease expires. Renewal uses the authority; local timeout checks use a conservative deadline accounting for elapsed request time and clock uncertainty. A replacement is activated only after the old forwarding/worker lease interval is fenced. Already admitted kernel work or network bytes may finish; the service does not claim instantaneous physical cancellation. |
| Long-prompt flood | Long-prompt flood: Token budgets and tenant concurrency limits reject work before memory collapse. Weighted queues reserve interactive capacity; batch work can wait longer. A scheduler may preempt and recompute lower-priority sequences under an explicit policy, trading latency for memory. Repeated preemption indicates over-admission and should not become invisible “free” capacity. |
| Slow or disconnected client | Slow consumer or disconnected client: Bounded buffers and cancellation free execution resources after a grace period. A gateway that drops only the socket but leaves the worker running wastes expensive tokens and blocks. Apply backoff/jitter to retries so an overloaded model is not hit by synchronized repeated prefills. A full-region outage requires capacity elsewhere with the compatible model loaded. Recovering request metadata alone does not load the weights or make another accelerator ready. |
15Operations, security, and cost
Latency, token and memory metrics
Observe input and output tokens/s, queue delay, TTFT, ITL, completion latency, active sequences, occupied/free KV blocks, prefix-hit tokens, preemptions, cancellation lag and per-tenant service share. Measure useful successful tasks alongside raw tokens and hardware utilization. A worker at 99% utilization may produce unacceptable token gaps or spend much of its time recomputing preempted prefixes.
Measured capacity and cost
Cost comparisons use resource units rather than guessed device prices. If a repeated private 2,000-token prefix saves 2,000 prefill tokens, ten reuses avoid 20,000 input-token computations while retaining about 250 MiB in the illustrative architecture. Compare the worker time saved with the memory no longer available to other sequences, and measure how often the prefix is evicted. For disaggregated prefill/decode, moving a 4,000-token KV prefix at 128 KiB/token transfers about 500 MiB per request; at 100 requests/s that is roughly 49 GiB/s before transport overhead. This quantifies the network bandwidth needed between the prefill and decode pools and helps decide where to place them.
Model and tenant security
Protect model/artifact integrity, tenant cache scopes and prompt/output retention. Keep credentials out of prompts and do not let generated text become server code. Tool execution belongs to a separate authorized service. Logs should prefer IDs, lengths and error classes over raw private prompts unless a reviewed debugging policy permits content access.
Version rollout and draining
Roll out immutable weight/tokenizer/template/runtime combinations through quality regression, mixed-load latency tests, a canary and controlled routing. Drain old workers while active requests finish; do not swap weights beneath live KV state. Test cancel-during-kernel, stale epochs, model reload failure, queue saturation and stream reconnect. Verify usage-event deduplication after crashes so repeated reporting does not distort tenant budgets.
Watermarks and billing reconciliation
Track the gap between worker-reported work and durable watermarks, reporting delay and unbilled interrupted tails separately from usage-event duplication. A 100-token watermark followed by a crash at token 120 is a recovery test: charge 100 recorded output tokens, mark the result interrupted, and never add a late duplicate report twice. Monitor actual hardware work separately. Preventing duplicate billing records does not prove that every computed token was recorded.
16Decision ledger and limitations
Decision table
| Decision | Benefit | Cost / consequence | Change trigger |
|---|---|---|---|
| Continuous mixed batching | Reuses slots as requests arrive/finish | Scheduler complexity and contention | Homogeneous offline jobs favor simpler batches |
| Chunked prefill | More decode opportunities during long prompts | Prefill scheduling overhead and tuning | Tight long-prompt TTFT changes the balance |
| Paged KV allocation | Less fragmentation and flexible growth | Block metadata and lifecycle correctness | Simpler fixed workload may tolerate reservation |
| Tenant-scoped prefix reuse | Saves repeated input work | Memory occupancy and isolation policy | Low reuse favors earlier eviction |
| Early work-based rejection | Predictable admitted latency | Explicit client errors during peaks | Batch tier can accept longer deadlines |
| Replicas before phase separation | Simple failure boundaries and no KV network handoff | Mixed-resource interference | Benchmarks justify disaggregated prefill/decode |
Parallelism and disaggregated serving
Tensor parallelism splits model-layer operations across devices and adds communication; pipeline parallelism places successive layers on different devices, which can sit idle while waiting for an earlier stage to produce input. Replicating complete workers is generally simpler when a model already fits and only aggregate throughput is lacking. The correct mix depends on actual weights, memory, network and latency targets, not a universal rule that one strategy is fastest.
Reservation versus utilization
Full maximum-length reservation gives a simple memory bound but may waste unused output capacity. Allocating memory as sequences grow can use space better, but may require pausing and recomputing work or tighter admission limits. It must still prevent uncontrolled out-of-memory failures. Prefix caching and answer caching solve different problems. A high input-cache hit rate does not mean output generation is cheap, and a larger batch can improve throughput while worsening each stream's latency. Report both before claiming an optimization succeeded.
Quantization and speculative decoding
Two further optimizations are worth discussing after the baseline is measured. KV quantization can reduce bytes per cached token but changes numerical behavior and needs compatible kernels, quality tests and scale/metadata accounting; the worked 128 KiB/token calculation deliberately assumes two-byte elements. Speculative decoding uses a cheaper draft process to propose multiple tokens and a target-model verification step to accept/correct them. Exact sampling preservation requires the algorithm's target verification and acceptance rules; blindly accepting draft tokens changes the model distribution. Its benefit depends on draft acceptance, verification cost and traffic mix. Neither optimization removes the tenant, budget or cancellation boundaries, and support varies with the chosen model/runtime.
17Interview closing
Rehearse the architecture and contract
“I designed a versioned multi-tenant text-generation service. A four-thousand-token prompt with six hundred output tokens consumes much more capacity than a short exchange, so I budget input tokens, output tokens and KV memory rather than only requests per second. Our example needs 400,000 input and 60,000 output tokens per second, with about twelve hundred active decoders; isolated throughput bounds are not a mixed-capacity proof.
Defend the critical boundary
“I begin with one bounded worker, then add token-based admission, continuous batching, chunked prefill and paged KV allocation. Prefix reuse is compatible-model and tenant scoped. The worker atomically owns its memory budget, and cancellation reaches the scheduler before blocks are safely released. Compatible replicas scale the service; splitting prefill and decode waits for evidence because KV transfers are large.
State the cost and next measurement
“I persist request and usage identity but explicitly mark active streams interrupted when ephemeral execution state is lost. I do not silently concatenate a different regenerated answer. My next measurement is the mixed-workload latency/memory curve and cancellation recovery under peak load.”
Answer the follow-up
If the interviewer asks for an offline bulk tier, allow longer queues and larger batches under separate capacity/budgets while protecting interactive reservations. The service objective changes; the same scheduler settings should not be assumed optimal for both tiers.
Practise the interview questions
Say your answer aloud before opening the model answer. Then answer the follow-up and compare the reasoning.
Explain prefill and decode to an interviewer.
Reveal a model answer
Prefill processes the input and constructs attention state. Decode generates the next tokens iteratively using that state. A long prompt stresses input processing; a long answer keeps generation and memory active for longer.
Interviewer follow-up
Which latency captures each user experience?
Reveal the follow-up answer
TTFT includes queueing and initial processing before the first token. ITL describes gaps during continued generation; total completion time also depends on output length.
What the answer must demonstrate: Do not collapse every latency into one average.
Can you size this service from 100 QPS?
Reveal a model answer
Not alone. With 4,000 input and 600 output tokens per request, it needs 400,000 input and 60,000 output tokens per second. I also estimate active sequences and KV memory, then benchmark the actual mix under latency targets.
Interviewer follow-up
Can you take the maximum of isolated prefill and decode worker counts?
Reveal the follow-up answer
That is only a lower-bound check. Shared compute and memory mean mixed-workload performance can require more workers and headroom.
What the answer must demonstrate: Do not treat separately measured prefill and decode throughput as capacity simultaneously available on the same worker.
Why use continuous batching instead of waiting for a fixed batch to finish?
Reveal a model answer
Requests have different output lengths. Continuous batching removes finished sequences and admits new work between iterations, reducing idle capacity. The scheduler still limits tokens and memory so larger batches do not ruin streaming latency.
Interviewer follow-up
Why chunk a long prefill?
Reveal the follow-up answer
It creates scheduling opportunities for existing decoders instead of letting one large prompt monopolize a long execution interval. I measure the overhead and latency tradeoff.
What the answer must demonstrate: Explain who waits and why.
Is prefix caching the same as answer caching?
Reveal a model answer
No. Prefix caching reuses compatible internal prompt state and then generates a new continuation. Answer caching returns an existing result and needs additional freshness, permission, and semantic rules.
Interviewer follow-up
What prevents cross-tenant prefix reuse?
Reveal the follow-up answer
A server-controlled isolation scope in the cache identity, compatible model/version settings, and enforced routing. Client-supplied arbitrary salts cannot define trusted tenant identity. Shared blocks remain read-only; diverging continuations allocate private writable state.
What the answer must demonstrate: Treat cache isolation as part of authorization design.
The user closes the tab at token 120. What should happen?
Reveal a model answer
The gateway propagates cancellation to the scheduler, which stops further generation and frees the sequence’s resources when safe. Stream buffers are bounded, and actual usage is recorded under the declared contract.
Interviewer follow-up
What if the disconnect is temporary?
Reveal the follow-up answer
The API specifies whether buffered events can be resumed. I do not keep expensive generation alive indefinitely just because reconnection is possible.
What the answer must demonstrate: Send cancellation to the worker scheduler and verify that it stops work and releases unused memory.
A GPU worker dies halfway through the answer. Can you transparently continue on another worker?
Reveal a model answer
Not without a defined recoverable state/output protocol. Normally I mark the stream interrupted; a fresh attempt may generate different text. I must not append unrelated regenerated text to the old stream silently.
Interviewer follow-up
When would you split prefill and decode across fleets?
Reveal the follow-up answer
When measured isolation or utilization gains exceed KV-transfer cost and added failure complexity. It is an optimization, not the starting architecture.
What the answer must demonstrate: State the recoverability limit of live KV state.
Why do 27 prefill workers and 40 decode workers not prove that 40 mixed workers suffice?
Reveal a model answer
Those are lower bounds from isolated benchmarks. Prefill and decode share compute, bandwidth and KV capacity on the same workers, and the batch mix changes latency. I use them to reject obviously undersized plans, then measure representative mixed traffic with headroom under TTFT and ITL targets.
Interviewer follow-up
Would adding the numbers prove 67 workers are enough?
Reveal the follow-up answer
No. An additive model assumes a particular way resources are time-shared. It may be a conservative planning approximation or still miss memory, networking and tail behavior. The mixed benchmark and failure reserve determine the actual fleet.
What the answer must demonstrate: Do not turn isolated maximum throughput into simultaneous guaranteed capacity.
Why not free request A’s KV blocks as soon as the API receives cancel?
Reveal a model answer
A running kernel may still read those blocks. The API saves and forwards cancellation; the worker stops scheduling new work, waits for the running computation to finish and releases its references. Reusing memory sooner could corrupt another request or expose data.
Interviewer follow-up
What if the worker does not respond to cancellation?
Reveal the follow-up answer
Bound its lease and detect cancellation lag. Interrupt or replace the unhealthy execution group under epoch fencing, mark the stream interrupted/canceled according to outcome, and do not route stale events as current. Capacity is reclaimed only when execution can no longer use it. Local memory must not be reassigned until the old execution is stopped or drained; epoch rejection alone only fences output/state publication.
What the answer must demonstrate: Cancellation acknowledgement, scheduler stop and memory reclamation are distinct moments.
Blank-page exercise · 45 minutes
Build the answer yourself
Design a multi-tenant inference service for request A’s long summary and request B’s short question, then cancel request A and lose a worker mid-stream.
- Define prefill, decode, TTFT, ITL, and KV memory.
- Calculate input/output token rates and a memory estimate.
- Trace token-based admission and continuous batching.
- Explain private prefix reuse and cancellation.
- State retry/stream behavior after worker failure.
Check that each component and design decision follows from your requirements and workload.
Recall the key ideas
Answer from memory before opening each card. Explain why the choice works and what it costs. Revisit missed cards tomorrow.
Design an LLM inference platformWhat is the difference between prefill and decode?Recall first, then reveal
Prefill processes the prompt; decode generates subsequent tokens using retained attention state.
Read the prompt; write the continuation.
Return to lessonDesign an LLM inference platformDoes a prefix cache store the answer?Recall first, then reveal
No. It reuses compatible prompt computation; new output tokens still require generation.
Reuse the beginning, generate the ending.
Return to lessonDesign an LLM inference platformWhy is QPS insufficient?Recall first, then reveal
Requests differ in input tokens, output tokens, duration, and KV memory.
Count tokens and live context.
Return to lessonFinal revision
Summary and interview notes
An inference platform limits token work, waiting and memory before generation starts. Workers schedule prompt processing alongside ongoing token generation and own the KV blocks those computations use. Reuse only compatible authorized prefixes, wait for active computations before freeing memory, and bill from durable usage records.
Remember these points
- Isolated prefill/decode throughput gives lower bounds, not proof of mixed-worker capacity.
- The KV estimate depends on architecture, head count, element size and live sequence length in addition to model weights.
- Prefix sharing requires compatible scoped identity, immutable shared blocks and safe reference lifetimes.
- Cancellation stops future scheduling before in-flight work drains and memory becomes reusable.
- Saving cumulative usage counters prevents repeated reports from duplicating charges. Finalizing the attempt prevents late reports from adding charges for work left unrecorded at the crash.
Interview tips
- Before choosing hardware, calculate input/output tokens per second, active requests as arrival rate × mean service time (Little’s law), and KV bytes per retained token.
- Explain who owns local capacity when two routers both see the last free slot.
- Test cancellation during a kernel and a worker crash between a usage report and the next token.
Important qualifications
- vLLM configuration changes over time; pin and load-test a runtime/model combination rather than relying on rolling-document defaults.
- KV quantization and speculative decoding are workload-dependent optimizations with compatibility and quality requirements.
- Output fencing prevents stale events being accepted; it does not by itself stop a GPU kernel or reclaim its memory.
Technical references
- vLLM optimization and tuningDocuments chunked prefill, scheduling tradeoffs, preemption, and parallelism choices.
- vLLM automatic prefix cachingDocuments compatible block reuse and cache-salt isolation.
- PagedAttention paperPrimary research on efficient attention-memory management for serving.
- vLLM documentationOfficial serving capabilities, including continuous batching and streaming; settings evolve by version.
Practice marks stay in this browser.