Learnastra AI SYSTEM DESIGNAnup Rai

Concept · Understand the mechanism

Contextual Retrieval

By Anup Rai8 min readReviewed September 2026

Contextual retrieval adds information about a passage's surrounding document to its searchable representation. The aim is to make a chunk understandable when separated from its original location. In the approach described by Anthropic, generated context is prepended before both embedding and lexical indexing. This is an ingestion technique; it does not require generating that context for every user query. Anthropic's description.

It addresses a particular failure: a relevant chunk may omit the subject, document scope or effective date needed to retrieve and interpret it. It does not repair an incorrect source or guarantee a correct answer.

Start with the missing context

Consider this illustrative source:

Document: Release approval policy, revision 12
Section: Production database migrations

A migration that removes a column requires a rollback plan.
It must be reviewed by the database owner before deployment.

If a boundary leaves only the second sentence in a chunk, “It” has no clear referent. A search for database-migration approval may miss it or confuse it with another approval rule.

Stored item Example Purpose
Raw chunk It must be reviewed by the database owner before deployment. Exact source evidence
Reliable metadata Release approval policy; revision 12; Production database migrations Subject and version
Candidate generated context This passage describes review of the rollback plan for a column-removing database migration. Accept only after confirming that the source supports this reading
Search representation Metadata/context followed by the raw chunk Input to retrieval indexing

The pronoun could be ambiguous even in the source: does the owner review the plan, the migration, or both? A production parser should not silently invent certainty. Use a larger source span or request editorial clarification when the distinction changes the answer. Context generation can expose source ambiguity; it cannot settle policy on the author's behalf.

Interview tip: First try preserving headings and a coherent paragraph. An LLM enrichment stage is unnecessary if a better boundary and reliable metadata solve the problem.

Define the requirements

Functional requirements:

  1. Preserve raw text, source offsets, document identity and version.
  2. Add only source-supported context useful for retrieval.
  3. Search eligible representations through dense, lexical or hybrid retrieval.
  4. Resolve each hit back to evidence that can support the answer and citation.
  5. Update or remove every affected derivative when its source changes.

Non-functional requirements:

  1. Measured evidence recall and answer quality against a simpler baseline.
  2. Bounded ingestion cost, provider concurrency and publication delay.
  3. Consistent source versions across representations and retrieval indexes.
  4. Permission enforcement on both raw passages and derived context.
  5. Recoverable ingestion with traceable failures and retry-safe writes.

Build the ingestion and query paths

Architecture / visual model
flowchart TD D[Versioned source and permissions] --> P[Parse coherent chunks and headings] P --> R[Preserve raw evidence and offsets] P --> C[Add metadata or generate source context] C --> V[Validate support and access scope] V --> E[Embed searchable representation] V --> L[Lexical index of representation] E --> M[Mark document version ready] L --> M Q[Query and authenticated scope] --> H[Search ready permitted versions] M --> H H --> K[Fuse and rerank candidates] R --> A[Load current permitted source evidence] K --> A A --> G[Answer with source citations]
Read diagram source
flowchart TD
    D[Versioned source and permissions] --> P[Parse coherent chunks and headings]
    P --> R[Preserve raw evidence and offsets]
    P --> C[Add metadata or generate source context]
    C --> V[Validate support and access scope]
    V --> E[Embed searchable representation]
    V --> L[Lexical index of representation]
    E --> M[Mark document version ready]
    L --> M
    Q[Query and authenticated scope] --> H[Search ready permitted versions]
    M --> H
    H --> K[Fuse and rerank candidates]
    R --> A[Load current permitted source evidence]
    K --> A
    A --> G[Answer with source citations]

Keep raw evidence separate from enrichment. A generated prefix is a useful search aid, not an additional independent source. When the answer depends on a fact supplied only in the prefix, fetch and cite the original span that supports it.

For a generation stage, the instruction should ask for the chunk's subject, scope and relevant references using only the supplied source. Require an explicit uncertain outcome when that information is absent. Treat instructions inside source documents as untrusted content. Validation should inspect added claims, source applicability and permission scope; valid JSON alone is insufficient.

Use bounded queues and concurrency based on measured throughput and provider quotas. Retry transient failures with limits. Store a job identity based on source version, chunk boundaries and enrichment configuration so repeated deliveries do not create duplicate publications.

Technique What changes Main tradeoff
Heading/metadata enrichment Search text gains reliable document attributes Depends on parser quality; may not resolve references
Generated contextual prefix Search text gains a source-derived explanation Generation cost and unsupported additions
Parent expansion Retrieve a child, then load its larger source unit More answer-context tokens; permission checks on the parent
Late chunking Encode the longer text before pooling chunk vectors Requires a suitable embedding architecture and input budget
Contextualized chunk embeddings Model encodes chunks with surrounding document context Provider/model contract and representation migration
HyDE Generate an answer-like search representation for a query Query-time latency and hypothetical-premise bias

Late chunking contextualizes token representations before chunk pooling. It does not prepend generated prose. Voyage's current contextualized embedding documentation lists voyage-context-4, with voyage-context-3 as an older available model. Pre-chunked input groups chunks by document; the group provides context for each chunk's embedding. Validate the current input and token limits before integration. Voyage documentation.

Related lessons: chunking strategies, embedding contracts, HyDE.

Interpret the published benchmark correctly

Anthropic's 2024 study reported the following retrieval failure rates, using 1 − recall@20 across its evaluated datasets and configuration:

Configuration Reported failure rate Relative reduction from baseline
Baseline 5.7% —
Contextual embeddings 3.7% About 35%
Contextual embeddings plus contextual BM25 2.9% About 49%
Contextual retrieval plus reranking 1.9% About 67%

These are reported experimental results, not expected gains for every corpus. The change from 5.7% to 2.9% is 2.8 percentage points, approximately 49% relative. Retrieval recall does not directly measure final-answer correctness or the fraction of completely failed user queries. Benchmark and methodology.

For your evaluation, isolate metadata, generated context, lexical fusion and reranking as separate changes. Hold the test questions and source snapshot fixed. Report cases where enrichment hurts, such as a wrong referent or an outdated version in the prefix.

Calculate ingestion costs explicitly

Assume an illustrative document contains 8,000 tokens split into twenty non-overlapping 400-token chunks. Generate a 60-token prefix for each chunk.

  1. If the full document is sent separately for each chunk, repeated document input totals 20 × 8,000 = 160,000 tokens before instructions and chunk-specific input.
  2. Generated prefix output totals 20 × 60 = 1,200 tokens.
  3. Embedding the prefixes and chunks totals 20 × (400 + 60) = 9,200 tokens, a 15% increase over the raw 8,000 tokens.
  4. Lexical indexes also store additional terms; retrieval and storage effects depend on the implementation.

Prompt caching can reduce the cost of repeated document input if the provider's cache rules are satisfied. Include cache-write and cache-read rates, expiration, misses and any minimum prefix size. Concurrent jobs may not all benefit from a cache that has not yet been populated. A “90% cached-input discount” would not imply a 90% reduction in total ingestion cost.

Compare total incremental cost over the document's update interval. Frequently revised documents may need repeated context generation, embedding and index work. Metadata enrichment avoids generation calls but still has parsing, validation and indexing costs.

Publish coherent versions and invalidate dependencies

A chunk's raw text can remain unchanged while its surrounding context changes. Renaming the document, changing the effective date or correcting an earlier definition may invalidate that chunk's generated prefix and embedding.

Track dependencies from source versions to chunks, generated context, vectors, lexical entries and caches. Do not use only the raw chunk hash to decide whether regeneration is needed.

Two separate indexes generally do not share a transaction. A useful publication design is:

  1. Build the new document version in both indexes with stable versioned IDs.
  2. Record successful completion for each required representation.
  3. Mark the version ready in an authoritative manifest after validation.
  4. Make queries select ready versions and reject stale hits against that manifest.
  5. Retire old entries asynchronously, with reconciliation for failed deletions.

For strict freshness or revocation, validate at evidence fetch and before exposure as required by the contract. New content can wait for indexing; revoked access should not wait for a background rebuild.

Diagnose the remaining failures

Symptom Check Possible repair
Wrong product appears in results Added context and original source identity Correct enrichment, regenerate affected entries
Exact identifier is missed Tokenization, lexical field and hard constraints Preserve identifiers; use filters where exact matching is required
Old policy wins Ready-version selection and effective dates Reject stale hits and repair version publication
A public chunk reveals private context Enrichment's source dependencies Restrict the derivative or rebuild from permitted material
Recall rises but answers worsen Duplicates, contradictions and evidence packing Adjust reranking and preserve supporting source spans
Indexing cost grows unexpectedly Reprocessing, cache misses and source churn Fix invalidation granularity and retry behavior

Use risk-based review samples and targeted test cases. A fixed 1% sample is not inherently sufficient, particularly for rare harmful errors. Include ambiguous references, tables, exceptions, mixed permissions and malicious source instructions.

Interview practice

Q1: What is the simplest alternative to generated context?

Preserve coherent source units and add reliable titles and section paths. Measure the remaining retrieval failures before introducing generation. The simpler design has fewer inferred claims and fewer update dependencies.

Q2: Is contextual retrieval the same as late chunking?

No. A generated prefix changes the text being indexed. Late chunking changes how chunk vectors are produced from contextualized token representations. They may address related context loss through different mechanisms and costs.

Q3: Can the answer cite the generated prefix?

It should cite the original supporting source. If the prefix resolves a reference using an earlier paragraph, retrieve that paragraph when needed. Otherwise the answer may rely on an unchecked inference that the visible citation does not support.

Q4: What does a 49% relative failure reduction mean?

For the cited experiment it means the reported failure metric fell from 5.7% to 2.9%, roughly 2.8 points divided by the original 5.7. It does not mean answer accuracy increased by 49 points or that a new application will achieve the same result.

Q5: Why regenerate an unchanged chunk?

Its representation may depend on a changed title, definition, date or neighboring passage. Version and invalidate the full enrichment dependency, not only the chunk's bytes.

Q6: How do you keep dense and lexical retrieval consistent?

Use versioned entries, explicit readiness and a publication policy that verifies required indexes before exposing a version. Reconcile partial writes and filter stale hits. Do not assume two services support a shared atomic commit.

Q7: What is the main security risk of adding context?

The derivative may expose information from a source section the reader cannot access, or repeat instructions embedded in untrusted documents. Preserve provenance, validate access scope and treat generated text as data rather than authority.

Final notes

Recall card: Preserve the source → restore missing context → validate additions → publish coherent versions → measure retrieval and answers separately. Context enrichment is useful when it fixes a demonstrated loss of meaning at chunk boundaries.

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← Advanced Retrieval Patterns
NEXT LESSONLate Interaction and ColBERT →

Explore the diagram