Learnastra AI SYSTEM DESIGNAnup Rai

Concept · Understand the mechanism

KV and prefix caches: reuse computation with clear boundaries

By Anup Rai4 min readReviewed September 2026

A KV cache stores attention keys and values for already processed tokens so autoregressive generation can reuse them. A prefix cache reuses compatible prompt-prefix computation across requests. An answer cache returns a previous result. These save different work and have different correctness requirements.

Calculate what is retained

For a conventional decoder with uniform layer shapes:

KV bytes ≈ 2 × layers × KV_heads × head_dimension
           × retained_tokens × concurrent_sequences × bytes_per_element

The two represents keys and values. For 80 layers, eight KV heads, width 128, 128,000 tokens, and two-byte elements, one sequence uses 41.94 decimal GB, or 39.06 GiB, of ideal cache payload. Add weights, allocation metadata, workspace, and other model state. This is an illustrative architecture, not a named model specification.

The cache is a resource consumed throughout a request. A short prompt followed by a very long output can create pressure later, even if admission looked inexpensive. Account for maximum retained history and concurrency, and decide how to reject, preempt, or shorten work when capacity is exhausted.

Grouped-query attention changes the shape

Multi-head attention (MHA) uses multiple query, key, and value heads. Grouped-query attention (GQA) shares each KV head across a group of query heads. Multi-query attention (MQA) uses one KV head. These are model-architecture choices, not arbitrary serving toggles for any trained checkpoint.

Example architecture Query / KV heads Relative ideal cache
MHA 32 / 32 1
GQA 32 / 8 1/4
MQA 32 / 1 1/32

The ratios assume all other dimensions match. Quality depends on the trained model and task; heads are not assigned universal human roles such as “logic” or “creativity.” See the GQA paper.

Trace a prefix-cache lookup

Architecture / visual model
flowchart TD A[Authorized request and model revision] --> B[Build compatible cache identity] B --> C{Matching prefix available?} C -->|Yes| D[Reuse its KV state] C -->|No| E[Compute and optionally retain prefix] D --> F[Process uncached suffix] E --> F F --> G[Decode new output]
Read diagram source
flowchart TD
  A[Authorized request and model revision] --> B[Build compatible cache identity]
  B --> C{Matching prefix available?}
  C -->|Yes| D[Reuse its KV state]
  C -->|No| E[Compute and optionally retain prefix]
  D --> F[Process uncached suffix]
  E --> F
  F --> G[Decode new output]

Cache compatibility can depend on token IDs, model weights, position handling, adapter, modality processing, and runtime rules. A byte-identical human-readable document is not enough if its tokenization or model context differs. Namespace caches for the isolation requirements; an entry being present never grants access to its content.

Example: three requests have a 2,000-token shared policy prefix and different 100-token questions. A warm prefix cache may avoid repeating the 2,000-token prefill. It does not reuse the different questions or automatically return an old answer. Changing an early system instruction can invalidate reuse for later tokens.

Memory tiers and eviction

Where supported, a cache hierarchy can use GPU memory, CPU host memory, and backing storage. HBM is GPU memory, not a separate tier after “VRAM.” A host or disk hit must transfer data; compare that cost with recomputation. Include finite capacity, eviction, version invalidation, and worker failure. SGLang HiCache documents one concrete hierarchy.

Provider prompt caching is a billing contract too

API providers may implement automatic or explicit caching with model-specific minimum lengths, lifetimes, matching rules, and read/write/storage charges. Track reported cached-input usage instead of assuming all repeated text is discounted. Check the dated pricing reference before making a financial estimate.

With illustrative costs of $1 for an uncached prefix, $1.25 to create its cache entry, and $0.10 for each later hit, n uses cost:

uncached = n
cached = 1.25 + 0.10 × (n − 1)

Caching wins above about 1.277 uses, so two uses suffice under these assumptions. Real storage charges, expirations, misses, and incompatible releases change the threshold. Put stable material first only when doing so preserves the prompt's intended semantics.

Four ways to reduce cache or context cost

Intervention Resource it changes Risk or cost
Paging and prefix sharing Allocation waste and repeated compatible state Metadata, eviction, identity and reference management
Cache quantization Bytes per stored value Quality change, scale metadata, kernel support
Architecture-level compression such as MLA Representation of attention state Requires compatible trained architecture and runtime
Retrieval or summarization Information included in the prompt Missing evidence, lost qualifications, summary errors

None inherently extends the model's validated context window by a guaranteed multiplier. See the MLA explanation for the architecture distinction.

Compare long context with retrieval

For a stable 50,000-token document, cached long context is a reasonable baseline if the model can use it reliably and the user is allowed to see all of it. Retrieval may win when the corpus grows, permissions differ by passage, or selective evidence improves quality. Whole-document inclusion does not guarantee that the model notices every relevant clause.

Evaluate answer correctness, citation support, update frequency, cache hit rate, TTFT, output behavior, and complete cost. Combine retrieval with caching when a stable instruction prefix and changing evidence make that useful.

Interview practice

  1. What does the KV cache avoid? Recomputing earlier keys and values during conventional autoregressive generation; new tokens still require work.
  2. Why use KV-head count in the formula? GQA and MQA share keys and values across query heads, changing storage.
  3. Is a prefix hit an answer hit? No. It reuses prompt processing, while the new output is still generated.
  4. What invalidates reuse? Incompatible tokens, weights, adapters, positions, runtime configuration, or authorization boundaries.
  5. Why might a disk cache lose to recomputation? Transfer and lookup latency can exceed the saved compute for the workload.
  6. When is cached context better than retrieval? Only when the measured task, access, freshness, latency, and cost comparison supports it.

Recall card and closing

Shape → compatibility → locality → lifetime → economics. Explain the saved work and what remains, then show how the cache behaves after a model release, an access change, and a miss.

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← Inference: follow the request before optimizing it
NEXT LESSONSpeculative decoding: propose cheaply, verify correctly →

Explore the diagram