Learnastra AI SYSTEM DESIGNAnup Rai

Concept · Understand the mechanism

Short-Term Context Management

By Anup Rai7 min readReviewed September 2026

Context management is the application's selection, organization and budgeting of information supplied to a model interaction. It includes instructions, messages, tool schemas, retrieved evidence and tool results. Computational caching can make some of that input cheaper to process; it does not decide which information is relevant or safe.

Remember: Context is selected information. A KV cache is saved computation. A session record is persisted application state.

Start with the actual limits

Limit Meaning What to check
Model/API context limit Maximum supported input/output accounting for the selected endpoint Exact model, modality and API contract
Maximum output Allowed generation length Whether reasoning or other internal tokens share this budget
Application input budget Smaller budget chosen for quality, latency and cost Measured workload and required evidence
Tool-result budget Bound on what one tool can add Truncation, pagination and artifact references

Do not describe the published context limit as solely a hardware constant: training, positional methods, model implementation and API restrictions all matter. Do not hard-code a list of “current” model sizes into the architecture; verify the selected deployment.

Suppose an illustrative endpoint allows 32,000 total tokens and we reserve 4,000 for output plus 1,000 of accounting margin. That leaves at most 27,000 input tokens. We choose a lower 20,000-token application input target:

Input allocation Tokens
Instructions and tool schemas 3,000
Current request and task state 2,000
Recent messages 5,000
Retrieved evidence and tool observations 8,000
Summary of earlier context 2,000
Total selected input 20,000

The remaining 7,000 is spare capacity, not a requirement to fill the window. This arithmetic assumes the stated endpoint accounting. Use a compatible tokenizer or provider token-count endpoint for the actual serialized request, including schemas and multimodal content.

Assemble context as a controlled pipeline

Architecture / visual model
flowchart TD R[Current request and authenticated task] --> S[Load permitted state and recent messages] S --> E[Retrieve relevant evidence] E --> V[Validate scope, freshness and message structure] V --> B{Fits application and API budgets?} B -->|No| C[Select, compact or fetch smaller artifacts] C --> V B -->|Yes| M[Call model under output and time limits] M --> T[Record response and tool requests] T -->|Tool work requested| U[Execute permitted tools and retain results] T -->|Final response| D[Return response] U --> S
Read diagram source
flowchart TD
    R[Current request and authenticated task] --> S[Load permitted state and recent messages]
    S --> E[Retrieve relevant evidence]
    E --> V[Validate scope, freshness and message structure]
    V --> B{Fits application and API budgets?}
    B -->|No| C[Select, compact or fetch smaller artifacts]
    C --> V
    B -->|Yes| M[Call model under output and time limits]
    M --> T[Record response and tool requests]
    T -->|Tool work requested| U[Execute permitted tools and retain results]
    T -->|Final response| D[Return response]
    U --> S

An overlong request is not universally handled by automatic eviction. An API may reject it or offer an explicit truncation/compaction feature. Make the application's behavior deliberate instead of relying on an undocumented “oldest tokens disappear” rule.

Choose what to retain

Technique Best use Main failure
Recent-turn window Local conversational continuity Drops an older constraint or decision
Structured task record Exact IDs, constraints, completed work Becomes stale unless updated reliably
Summary Compact account of earlier discussion Loses caveats, negation or source distinctions
Selective retrieval Recover relevant earlier evidence Misses a necessary item
External artifact reference Large code, documents or query outputs Reference inaccessible or revision changes
Hybrid selection Combine exact state, recent turns and selected evidence More selection logic to evaluate

There is no universally best “last ten messages plus summary” policy. Trigger compaction based on projected size and preserved information, not just a magic turn count. A single tool result can exhaust the budget before ten turns; a hundred short turns might fit.

Preserve a compact task record

For a documentation-update task, keep these fields exact:

  1. Goal and approved scope.
  2. Target schema/repository revision.
  3. Decisions and constraints that remain in force.
  4. Completed edits and validation results tied to artifact revisions.
  5. Pending questions and next useful work.
  6. External operation IDs and any uncertain outcomes.

A summary saying “the update is almost finished” is not an adequate substitute. If the schema revision changes, invalidate the affected checks rather than treating all earlier success as current.

Keep message protocols valid

Trimming must preserve the selected provider's required relationships between tool calls and tool results. Do not leave orphaned result IDs or move untrusted tool output into a higher-authority instruction role. If the provider requires opaque continuation items, preserve them according to that contract. Do not invent a universal rule to remove all reasoning-related blocks.

When truncating a result, label it incomplete and provide a way to request the missing range. Silently cutting a stack trace, SQL result or contract clause can change its meaning.

Understand KV memory without confusing it with context selection

During standard autoregressive Transformer inference, keys and values from earlier token computations can be cached for subsequent attention. This avoids recomputing those projections on every generated token. The cache still consumes memory and must be accessed as required by the attention implementation.

PagedAttention manages KV storage in blocks through an indirection mechanism, reducing waste associated with allocation and supporting sharing. It does not summarize a conversation, choose relevant sentences or give a model unlimited context. Paged KV implementations and attention kernels vary across runtimes; do not label every block-based system the same algorithm.

For a simplified decoder with 32 layers, 8 KV heads, head dimension 128 and 2-byte elements:

KV bytes/token = 2 × 32 × 8 × 128 × 2 = 131,072 bytes

The leading 2 accounts for keys and values. At 16,384 tokens this is 2 GiB per sequence, before allocation overhead and implementation-specific sharing or compression. Full multi-head attention with 32 KV heads would be four times as large under these assumptions. KV-cache calculations explain the boundary in more detail.

Do not promise a fixed 60–80% memory saving. Benefits depend on the prior allocator, sequence lengths, cache sharing, model and load.

Prefix caching changes computation, not the answer contract

Prefix caching reuses eligible computations for a matching earlier portion of input. vLLM automatic prefix caching primarily avoids repeated prefill work; it does not eliminate generation of the new answer. Application summaries and retrieval still determine the content.

For eligible workloads, keep stable instructions and tool definitions before changing request data. Check the actual provider's matching rules, cache lifetime, minimum length, authorization isolation and invalidation behavior. Never retain obsolete instructions merely to preserve a cache hit.

Hosted APIs may charge for cache creation and reads. For example, Claude prompt caching documents cache lifetimes and distinct read/write pricing. Thus “you pay only for new tokens” is not a general rule. Cached input also still occupies the model's applicable context accounting.

With illustrative prices, suppose 12,000 reused input tokens cost $0.20 per million to read, 2,000 new tokens cost $2 per million and 1,000 output tokens cost $8 per million. The request costs $0.0024 + $0.004 + $0.008 = $0.0144, excluding any earlier cache-creation charge. An uncached equivalent costs $0.028 + $0.008 = $0.036. These are hypothetical rates, not a vendor quote.

Compression must preserve meaning

Extracting relevant paragraphs, summarizing and learned token/KV compression are different techniques. An application-level rewrite is not guaranteed semantically equivalent because it is 50% shorter. Losing “except,” an identifier, a date or a negative constraint can reverse the answer.

Use exact spans for calculations, citations, schemas and critical conditions. Preserve source references and uncertainty in summaries. Specialized token/KV compression needs evaluation for the supported model/runtime; an ordinary API caller cannot assume arbitrary access to the server's KV tensors.

Long input does not inevitably fail, and short input is not automatically better. The Lost in the Middle study motivates testing evidence placement. Measure your model and task with relevant evidence at different positions, with distractors, and after multiple compaction cycles.

Failure drill and measurement

The first turn says “only change public API documentation.” Forty turns later, compaction omits that constraint and the agent edits internal schemas.

Fix: store active scope in structured task state, preserve it during context assembly and enforce permitted file operations in application code. Test that rephrasing, compaction and context resets cannot broaden authority. A better summary alone is not the complete authorization boundary.

Measure:

  1. Task success and preservation of critical constraints.
  2. Input/output tokens, cache hit tokens and total billed cost.
  3. Time to first token and total task latency separately.
  4. Retrieval misses and summaries that change facts.
  5. Invalid message sequences and context-limit errors.
  6. Recovery after context reset or a changed artifact revision.

Interview practice

Q1: Does prefix caching make a 100,000-token prompt equivalent to a short prompt?

No. It can reduce repeated input processing, but context still contains that information and generation still does work. Quality, attention behavior, cache eligibility and billing must be measured separately.

Q2: Why keep a structured task record in addition to a summary?

Some fields must remain exact: scope, IDs, versions, approvals and uncertain operations. Prose summaries are useful for narrative context but can omit or reinterpret those fields.

Q3: What happens when the next tool result will exceed the budget?

Request a bounded subset, paginate, retain an external artifact or extract the relevant part. Label omissions, preserve provenance and re-count the final request before dispatch.

Q4: Does PagedAttention solve long-context reasoning errors?

No. It addresses KV-memory management. Better utilization does not prove the model selects or uses the right evidence.

Q5: Would you always summarize after ten turns?

No. I would use size, information value and task boundaries, then evaluate semantic preservation. Repeated summaries can compound errors; exact records and retrieval offer ways to recover omitted evidence.

Q6: What belongs in the closing recommendation?

The selected input/output budgets, which information remains exact, how older evidence is recovered, protocol-safe compaction, and measured quality/cost/latency. Then explain which computational caching is available in the chosen deployment.

Final notes

Recall card: Budget → select → preserve exact constraints → validate structure → cache eligible computation → measure the result.

Continue with long-term memory and semantic response caching, which solves a different reuse problem.

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← Memory Architectures
NEXT LESSONLong-Term Memory →

Explore the diagram