LLM infrastructure is the compute, networking, storage, scheduling and operational machinery that serves model-backed application requests. Its job is to meet the application's quality, latency, availability and cost requirements under normal load and defined failures.
Request count is only the beginning. A short classification call and an agent making ten long-context calls create very different workloads. Start with token lengths, concurrency, deadlines and dependencies before choosing a GPU or provider.
Define the terms used in capacity planning
| Term | Meaning | Practical consequence |
|---|---|---|
| Prefill | Process input tokens to prepare generation state | Long prompts can delay generation and consume substantial compute |
| Decode | Generate output tokens iteratively | Long answers keep requests active and consume more capacity |
| Time to first token (TTFT) | Time from the chosen request boundary to first output token | State whether client/network/queue time is included |
| Inter-token latency | Time between successive output tokens/events, according to the metric | Streaming events may contain more than one token |
| Throughput | Completed work per time, with stated units | Requests/s and tokens/s answer different questions |
| Concurrency | Requests in progress at the same time | Influences memory and scheduling pressure |
| Admission control | Decide which work the service can accept | Prevents unlimited queues from becoming a latency failure |
| Scheduling | Decide which accepted work runs next | Balances fairness, deadlines and hardware efficiency |
A fast first token does not guarantee fast completion. Track both. See inference fundamentals and KV caching.
Interview scope: one million coaching requests per day
Assume the service generates interview feedback and study guidance. The numbers below are planning assumptions, not measured Learnastra traffic.
Functional requirements
- Authenticate requests and apply account/module access rules.
- Serve interactive feedback with streaming when appropriate.
- Run document ingestion and offline evaluation as asynchronous jobs.
- Record job status, usage, selected model and result provenance.
- Support cancellation, versioned releases and failure recovery.
Non-functional requirements
- Size for a stated peak and token-length distribution, not only the daily average.
- Define separate latency/availability objectives for interactive and offline work.
- Bound queue age, concurrency, tokens and spend by account and workload class.
- Preserve data-handling policy through tools, logs, caches and fallback paths.
- Keep accepted durable work recoverable within agreed recovery objectives.
Clarify what “a request” means. If it means a user task, translate it into the number and shape of model calls. Do not silently treat a ten-call agent task as one inference request.
Start with a minimal service
A useful baseline is an authenticated application calling one suitable managed model, with bounded timeouts, usage recording and a durable queue for offline jobs. This avoids owning a serving fleet before the product's traffic and quality needs are known.
| Managed model API | Self-hosted model |
|---|---|
| Provider operates model-serving infrastructure | Team operates hardware/instances, runtime and model rollout |
| Team still owns quotas, access, retries, evaluation and incidents | Team also owns memory, utilization, redundancy and serving failures |
| Data handling depends on provider and configuration | Data may still leave through tools, telemetry or remote fallbacks |
| Model choices follow supported provider offerings | Model choices follow available artifacts, licenses and runtime support |
| Pay-per-use or capacity commitments vary | Idle capacity, staffing and upgrades affect effective cost |
There is no universal request-volume crossover or guaranteed “under 100 ms” self-hosted latency. Compare equal quality and service objectives. A hybrid design is useful only if both paths can be operated and evaluated.
Derive the workload
average arrival rate = 1,000,000 / 86,400 ≈ 11.6 requests/s
assumed planning peak = 60 requests/s
assumed mean time in system at stable load = 6 seconds
mean in-flight work ≈ 60 × 6 = 360 requests
This is Little's Law applied to the same system boundary and stable workload. It is not a GPU count or a guarantee during overload. If arrivals continuously exceed capacity, queue growth invalidates a steady-state estimate.
Now suppose each model call averages 2,000 input tokens and 300 output tokens at peak:
input demand = 60 × 2,000 = 120,000 input tokens/s
output demand = 60 × 300 = 18,000 output tokens/s
Measure the distribution, including long tails, rather than sizing from averages alone. API quotas can constrain requests, input tokens, output tokens, concurrent requests or provisioned capacity differently. Inspect the actual contract.
Convert benchmark evidence into a replica estimate
Assume one replica sustains 2,000 output tokens/s on this input/output mix while meeting the latency objective. With a planning limit of 70% of that measured capacity:
normal replicas = ceil(18,000 / (2,000 × 0.70)) = 13
replicas to preserve the same limit after losing one = 14
Also verify prefill throughput, memory, network and shared dependencies. The 70% planning limit is illustrative, not a universal utilization target. If two replicas share a failure domain, “one replica lost” is too weak a resilience assumption.
Evolve the design for load and failures
Read diagram source
flowchart TD
U[Clients] --> G[Authentication, quota and deadline]
G --> A[Admission and workload classification]
A --> Q[Bounded interactive queue]
A --> B[Durable offline queue]
Q --> R[Eligible endpoint or warm replica]
R --> M[Prefill, decode and streaming]
M --> V[Validate completion and record usage]
B --> W[Separately limited workers]
W --> D[Job and artifact stores]
A --> O[Queue, demand and outcome metrics]
M --> O
W --> O
O --> C[Capacity controller and on-call]
A queue absorbs bounded bursts when waiting is acceptable. It does not create compute or provider quota. Reject, defer or degrade work before it becomes impossible to meet the promised deadline. Separate interactive and offline capacity so a backfill cannot starve a live interview session.
Each accepted request carries trusted account scope, workload class, deadline, token/output limits, release revision and operation ID. Cancellation should propagate to downstream work where supported; the system must still reconcile usage or effects that completed before cancellation.
Self-hosting: memory and parallelism
Memory includes model weights, key/value cache, activations/workspaces and runtime overhead. “The weights fit” is not a production concurrency test.
For a conventional attention cache, an illustrative per-sequence estimate is:
KV bytes ≈ 2 × layers × KV heads × head dimension × bytes/value × sequence tokens
= 2 × 32 × 8 × 128 × 2 × 8,192
= 1,073,741,824 bytes = 1 GiB
This example assumes those dimensions and an uncompressed full-sequence cache. Sliding-window attention, latent attention, hybrid architectures, cache quantization and parallel placement can change the formula or allocation. Add runtime overhead and validate actual memory use.
| Technique | Benefit | Cost or limit |
|---|---|---|
| Continuous batching | Reuse active capacity as requests finish | Scheduling/fairness and cache constraints remain |
| Horizontal replicas | Increase independent serving capacity | Replicated weights and possible cold starts |
| Tensor parallelism | Divide layer operations across devices | Communication bandwidth/latency |
| Pipeline parallelism | Divide model layers into stages | Pipeline bubbles and balancing complexity |
| Expert parallelism | Distribute MoE experts | Routing traffic, imbalance and interconnect demand |
| Prefill/decode separation | Tune resource pools for different phases | Cache transfer and extra scheduling/network complexity |
| Quantization | Reduce selected memory/computation costs | Kernel support and quality validation |
Decode at small batches is often limited by memory bandwidth; the bottleneck is workload-dependent. More accelerators can make a poorly partitioned workload slower. See serving infrastructure.
Current runtime and hardware choices
As reviewed in September 2026:
- vLLM and SGLang are serving candidates; check the exact model, hardware backend, precision and feature combination.
- TensorRT-LLM provides NVIDIA-oriented serving paths. Benchmark the supported backend instead of assuming a universal performance winner.
- Text Generation Inference is in maintenance mode. Its maintainers recommend alternatives including vLLM and SGLang for ongoing adoption; a working existing deployment still needs a planned lifecycle decision.
- CPU/edge engines can fit local or smaller workloads, but should be tested against the same product latency and quality requirements.
| Platform family | What to examine before selection |
|---|---|
| NVIDIA GPUs | Supported runtime/kernels, memory, interconnect topology, supply and total cost |
| AMD GPUs and ROCm | Exact device/runtime compatibility, numerical behavior and feature coverage |
| AWS accelerators and Neuron | Model/runtime support, instance availability and migration effort |
| Google Cloud TPU | Supported software path, topology, model fit and regional capacity |
| Specialist platforms, such as Tenstorrent | Actual supported models/operators, programming stack and workload benchmarks |
Headline FLOPS, active parameter counts, vendor funding and announced datacenter capacity do not establish usable application throughput. Distinguish a product announcement from capacity you can provision. A single well-operated platform can be the right choice; multi-vendor operation has its own engineering and evaluation cost.
Autoscaling and observability
Scale from queue age, admitted token demand, active requests, usable memory and measured service capacity. CPU utilization alone may miss a GPU memory or bandwidth bottleneck. Kubernetes HPA supports custom/external metrics through the appropriate metrics APIs/adapters; a YAML metric name does not create the metric pipeline.
Warm-up includes scheduling an accelerator, loading weights and preparing the runtime. Maintain appropriate warm headroom and drain requests before removing replicas. Autoscaling cannot rescue a request whose deadline expires before the new capacity is ready.
vLLM metrics include token distributions and latency measurements. Combine model-server measurements with client/application traces to locate waiting outside the server. Keep high-cardinality request IDs in traces/logs rather than unbounded metric labels.
A representative request might queue for 80 ms, emit its first token at 350 ms, finish generation at 2,400 ms and complete validation at 2,450 ms. These timestamps expose where time went; they are not promised production performance.
Find flaws and define recovery
| Failure | Repair | Tradeoff |
|---|---|---|
| Background jobs consume every slot | Separate quotas/pools and fair scheduling | Some capacity may be less fully utilized |
| Provider throttles requests | Admission/backoff and an eligible tested alternate | More queueing or alternate-model cost |
| Worker dies after a durable job starts | Lease/claim, persisted progress and bounded retry | Must deduplicate external effects |
| Cache returns another account's answer | Scope keys by access and relevant versions | Lower reuse but correct isolation |
| Region or database fails | Tested restore/failover for the relevant state | Replication, storage and operating cost |
| Alternate model changes tool semantics | Capability and outcome checks before enabling fallback | More evaluations and constrained fallback options |
Recovery time objective (RTO) is the target time to restore service; recovery point objective (RPO) is the tolerated data-loss window. Set them for each state class. Losing a reconstructible cache differs from losing accepted jobs or user study plans. Exercise restoration and representative requests, not merely server startup.
Use alerts tied to a service consequence and a response: rising interactive queue age plus sustained token backlog should trigger capacity/overload investigation. Quality, cost and access violations need their own signals. An API failover does not repair a shared identity or storage outage.
Cost and closing decision
Compare provider charges plus application operations with self-hosted compute, idle headroom, staff, storage, networking, support and migration. Include retries and quality-related review. Track cost per successful task, not just dollars per advertised token.
Start with the smallest deployment that meets the measured workload and operating constraints. Add serving complexity when a benchmark or failure analysis identifies the requirement it solves. Record the assumptions that would trigger a different decision.
Interview questions and answer checks
- Why is one million requests/day insufficient sizing information? It omits peaks, token distributions, model calls per task, latency and failure requirements.
- Why did adding GPUs fail to improve latency? The bottleneck may be admission, interconnect, tools, storage, scheduling or an inefficient partition.
- Why can weights fit but serving fail? Concurrent cache/workspace demand and long sequences consume additional memory.
- When does batching hurt? When waiting or shared scheduling pushes interactive requests beyond their latency budget.
- Does a second provider guarantee availability? No; quotas, quality, policy and correlated dependencies must be checked.
- How do you test disaster recovery? Restore the intended state, run representative authorized requests/jobs, and measure actual recovery time and data loss.
- Can a queue absorb permanent overload? No; bound acceptance and add capacity, reduce permitted work or reject/defer requests.
- What makes self-hosting a good decision? Demonstrated model quality, sufficient utilization/control benefit, realistic total cost and a team able to operate it reliably.
Final notes
Remember workload → capacity → admission → observation → recovery. Requests become tokens and tool work; concurrency consumes time and memory. Size and operate that complete path before choosing infrastructure from a product name.