Retrieval enhancements change the query, indexed representation or candidate-selection process to address a specific evidence-retrieval failure. Start by naming that failure. A more elaborate pipeline is useful when it improves evidence coverage or ranking enough to justify additional work.
This chapter distinguishes techniques that are often grouped together even though they operate at different stages.
Map the failure to the intervention
| Observed failure | Candidate technique | Where cost is added |
|---|---|---|
| Query wording differs from the corpus | Query rewriting or multi-query expansion | Query time |
| Question requires several facts | Decomposition and evidence gathering | Query time, possibly dependent rounds |
| Short query is poorly represented for retrieval | A compatible asymmetric encoder or HyDE experiment | Model selection or query time |
| A chunk loses its subject or time context | Headers or contextual enrichment | Ingestion and updates |
| One document can answer several distinct intents | Multiple indexed representations | Ingestion, storage and candidate fusion |
| Candidates are present but badly ordered | Reranking | Query time |
None of these repairs a missing authoritative source by itself.
Multi-query expansion is not decomposition
Multi-query expansion produces alternative formulations of approximately the same information need. For example, “reduce model-serving latency” might also search “improve time to first token” and “reduce decode latency.” Those are related facets, so check that an expansion does not silently narrow or change the original task.
Decomposition splits a compound question into required subquestions. For “Compare the Q3 and Q4 incident totals and explain any change,” retrieve each period's count and evidence for the causes. Do not assume there was a decrease or infer a causal explanation from the totals alone.
- Preserve the user's entities, dates and hard constraints.
- Generate a bounded set of reformulations or subquestions.
- Identify dependencies; run independent searches concurrently.
- Deduplicate by source identity and combine candidates with a defined ranking policy.
- Check coverage of each required fact before answering.
Three searches returning twenty candidates each produce at most sixty distinct candidates, and usually fewer after deduplication. They may add coverage or just repeat the same weak evidence. Measure the marginal benefit of each search.
Hypothetical Document Embeddings (HyDE)
HyDE generates a hypothetical answer-like document, embeds it, and uses that representation to retrieve real documents. The generated text is a search aid. It is not evidence and must not be presented as a source. The original paper studies this as a zero-shot dense-retrieval approach without relevance labels. Gao et al..
Read diagram source
flowchart LR
Q[Original question] --> H[Generate hypothetical passage]
H --> E[Embed search aid]
E --> R[Retrieve actual permitted sources]
Q --> V[Verify evidence against original question]
R --> V
V --> A[Answer from actual sources or abstain]
Suppose a query asks whether a product supports offline synchronization. A hypothetical passage may invent that feature and retrieve documents about a different product that does support it. Preserve the original product constraint and inspect real evidence before answering.
Combining ordinary lexical/dense retrieval with HyDE candidates can reduce dependence on one generated hypothesis. Rank fusion is one option. It does not automatically remove the hypothesis's bias or establish that the feature exists.
Understand asymmetric retrieval
Asymmetric retrieval matches inputs with different roles, such as a short question and an answering passage. This is more than a difference in word count: the desired relationship is “answers this question,” not necessarily “is a paraphrase.”
Models may use shared weights with different query/document instructions, or distinct compatible encoders trained for the relationship. Two arbitrary encoders do not become compatible because one receives short text and the other receives long text.
A suitable retrieval model may already handle this asymmetry without HyDE. Compare the documented query/document formatting, task-specific training and simple query rewriting before adding another generation call.
Enrich documents without inventing evidence
| Representation | Retrieval benefit | Validation requirement |
|---|---|---|
| Document title and section path | Names the subject and scope | Confirm parser metadata and effective version |
| Generated chunk context | Resolves references to surrounding content | Check every added claim against the source |
| Summary | Represents a long source compactly | Preserve a path to detailed evidence and exceptions |
| Synthetic questions | Matches likely user formulations | Ensure the source actually answers each question |
| Entity or relationship metadata | Supports filtering and graph expansion | Validate identity, type and provenance |
Store generated representations separately from original evidence. Multiple representations should map to the same source/version so duplicate hits do not appear to be independent corroboration. When a source changes, invalidate its derived questions, summaries and embeddings as required.
There is no general rule that modern retrieval indexes should embed questions instead of passages. Compare single and multiple representations on the actual query distribution. Extra representations add index entries, generation cost and maintenance work.
See contextual retrieval for enrichment using surrounding document context and chunking for parent expansion and late chunking.
Reranking inside a large context
A model can receive a list of candidate passages and select the relevant ones. This is a form of listwise reranking, sometimes combined with answer generation. A large context window makes larger inputs possible but does not prove that every candidate will be considered accurately.
One hundred passages averaging 1,500 tokens require about 150,000 tokens before instructions, queries, identifiers and output allocation. Check the actual model's limits and complete-request cost. Do not assume every contemporary model offers a million-token window or that a cache makes this equivalent to a small prompt.
Require valid candidate IDs, test position bias and verify that final claims cite the selected source material. A separate ranking stage can simplify diagnosis; a combined stage can reduce calls but entangle selection and generation failures.
Compare costs over a workload
Let C_ingest be the incremental cost of document enrichment over an update interval, Q the number of queries in that interval, and C_query the extra cost per query-side expansion. The direct added costs are C_ingest versus Q × C_query, before any differences in storage, latency and downstream model use.
For an illustrative $200 enrichment pass and $0.002 per expanded query, the arithmetic crosses at 100,000 queries. This is not a quality equivalence or a universal break-even point. Frequent source changes can repeat ingestion work; repeated queries may make caching valuable. Compare supported-answer quality as well as cost.
Interview practice
Q1: How do multi-query and decomposition differ?
Multi-query explores alternative formulations or facets of an information need. Decomposition identifies distinct facts required to answer a compound question. Both need bounded search and deduplication, but decomposition also needs a coverage and dependency plan.
Q2: Why can HyDE help and hurt?
An answer-like passage can give the retriever a more useful representation than the raw query. It can also invent a premise or steer search toward the wrong topic. Use real retrieved evidence for the answer and test false-premise and no-answer cases.
Q3: Is a generated question a valid citation?
No. It is a derived search representation. Resolve it to the source passage and verify that the source answers the user's question. Generated text should not become independent evidence for itself.
Q4: What is the simplest contextual enrichment baseline?
Add reliable document and section metadata to the searchable representation, preserving the raw passage. It may solve subject ambiguity without an LLM call. Compare generated context only where the simpler method leaves meaningful failures.
Q5: Can a large-context model replace a reranking service?
It can perform candidate selection, but I would measure ranking and final-answer quality, latency, token cost and position sensitivity. Capacity alone does not establish that reading a hundred passages is the best design.
Q6: How do you choose the next retrieval enhancement?
Classify failures by source availability, query interpretation, candidate coverage, ranking and answer generation. Change the stage responsible and compare against the baseline. Adding every technique at once makes benefits and regressions difficult to attribute.
Final notes
Recall card: Change a representation to fix a measured mismatch; keep generated search aids separate from source evidence. The closing interview argument is the measured benefit and the failure modes that remain.