A design pattern is a reusable approach to a recurring design problem, with assumptions and consequences. It is not a framework, a mandatory architecture, or a guarantee of quality. In an interview, name the problem first, explain the mechanism, and identify what new failure it introduces.
This chapter uses a hypothetical Learnastra study assistant: it answers questions from course material, suggests exercises, and can propose changes to a learner's study plan. The numbers below are design assumptions, not claims about the deployed product.
1. Establish the contract before choosing patterns
Functional requirements
- Answer a question with references to the relevant course sections.
- Ask for clarification or abstain when the available evidence is insufficient.
- Suggest exercises based on the learner's permitted progress data.
- Apply a study-plan change only after validating the requested action and authorization.
Nonfunctional requirements
- Never expose another learner's private notes or progress.
- Target a measured p95 response time of five seconds for ordinary questions.
- Track total cost per successfully completed task, including retries and evaluation.
- Bound each request's deadline, model calls, tool calls and concurrent work.
- Version the curriculum, prompts, models and evaluation cases so regressions can be reproduced.
Start with an authenticated request, one retrieval operation and one model response. An ordinary question does not need a planner, a manager agent and a critic by default. Add a component when an observed failure justifies it.
2. Retrieval patterns: improve the missing stage
RAG retrieves external information and makes it available to generation. The useful distinction is which part of that process changes.
| Pattern | Mechanism | Good reason to use it | Cost or failure to test |
|---|---|---|---|
| Basic RAG | Retrieve relevant chunks and answer from them | Small, well-structured corpus; initial baseline | Missing evidence, stale versions, unsupported citations |
| Advanced retrieval pipeline | Rewrite a query, combine lexical/vector candidates, rerank and pack evidence | Baseline search misses identifiers or ranks weak matches first | More calls and latency; rewriting can change intent |
| Parent–child retrieval | Search small child chunks, supply their larger parent sections | A matching sentence needs its definition, table header or exceptions | Duplicate parents, excess context, permissions on the expanded text |
| Adaptive retrieval | Decide when another retrieval step is useful | Some tasks need external facts while others already have enough context | A mistaken skip removes necessary evidence |
| Corrective retrieval | Assess retrieved evidence, then repair retrieval or abstain | Search sometimes returns unrelated or incomplete material | The evaluator can be wrong; external search changes the trust boundary |
“Advanced RAG” is an umbrella description, not a standardized list of mandatory stages. Hybrid search and reranking solve different problems: the first broadens candidate discovery; the second reorders candidates. A reranker cannot recover a document that never entered the candidate set.
Parent–child retrieval without losing rank or authority
Suppose child search returns A2, A1, B3, C1. The parent order should be A, B, C, preserving the strongest child match for each parent. Converting parent IDs to an unordered set and truncating can discard the best result.
def ranked_parent_ids(hits, allowed_parent_ids, limit):
"""hits are already ranked; permissions must come from trusted code."""
if limit <= 0:
return []
result, seen = [], set()
for hit in hits:
parent_id = hit["parent_id"]
if parent_id in seen or parent_id not in allowed_parent_ids:
continue
seen.add(parent_id)
result.append(parent_id)
if len(result) == limit:
break
return result
hits = [{"parent_id": p} for p in ["A", "A", "B", "C"]]
assert ranked_parent_ids(hits, {"A", "C"}, 2) == ["A", "C"]
assert ranked_parent_ids(hits, {"A", "B", "C"}, 1) == ["A"]
assert ranked_parent_ids(hits, {"A"}, 0) == []
This example only handles ordered selection. Production code must load the correct parent revision, recheck access, and pack within a token budget. A public child sentence must not unlock private siblings. If a parent is too large, select the necessary authorized span rather than silently dropping the answer-bearing exception.
Self-RAG and CRAG are specific research methods
Self-RAG trains a model to use reflection tokens for retrieval decisions and assessment of passages and generated content. A generic if needs_retrieval branch followed by a critique prompt is an adaptive workflow, but is not a faithful implementation of that training method. See the Self-RAG paper.
Corrective RAG (CRAG) uses a retrieval evaluator to assess evidence quality and trigger different retrieval actions, with knowledge refinement and web search in the proposed method. It does not define universal thresholds such as “three relevant documents means good evidence.” See the CRAG paper.
For the study assistant, public web search could supplement a dated technology lesson only under an explicit policy. It must not substitute a public answer for a question about private account data. Keep source dates and provenance visible; conflicting sources may require clarification rather than merging them into a confident answer.
3. Execution patterns: who decides the next step?
| Pattern | Control structure | Example | Main tradeoff |
|---|---|---|---|
| Prompt chaining | Application defines a fixed sequence with checks between stages | Extract requirements, then draft a study plan | Predictable but inflexible; early errors propagate |
| Routing | Select a path before executing it | Send account questions to a deterministic account API | Classification mistakes; every path needs evaluation |
| ReAct-style tool loop | Alternate model-selected actions with observations | Search a concept, inspect the result, request a missing prerequisite | Flexible but can loop, misuse tools or act on hostile content |
| Plan and execute | Maintain a plan and execute bounded steps, revising when observations require | Build a multi-week learning plan with prerequisites | Plans become stale; replanning adds cost |
| Evaluator–optimizer | Generate, evaluate against criteria, revise within limits | Repair an exercise that violates its answer rubric | Correlated generator/judge errors; repeated polishing without improvement |
| Orchestrator–workers | A coordinator creates and delegates subtasks dynamically | Review independent parts of an architecture proposal | Coordination, merge conflicts, duplicated work and aggregate spend |
These are execution choices, not measures of sophistication. Anthropic's workflow guidance distinguishes fixed workflows from systems in which the model determines its next actions.
Make the loop state explicit
Read diagram source
flowchart TD
A[Validated request and budget] --> B[Choose next bounded step]
B --> C{Action permitted?}
C -->|No| H[Explain limit or request authorized clarification]
C -->|Yes| D[Execute with operation identity]
D --> E[Record observed result]
E --> F{Goal met and checks passed?}
F -->|Yes| G[Return verified result]
F -->|No| I{Time and step budget remain?}
I -->|Yes| B
I -->|No| J[Return partial result or handoff]
The application's permission gate is separate from model reasoning. A model's explanation does not prove an action was authorized or that its reasoning caused the outcome. Evaluate tool arguments, observed results and resulting business state.
For plan-and-execute, use a versioned queue of remaining steps. Replanning replaces that queue; merely assigning a new list inside a for loop does not change the iterator already being traversed. Record completed actions separately so replanning does not repeat them.
For an evaluator–optimizer loop, specify the rubric, maximum revisions and terminal states. If revision three still fails, return a failed/partial status or route to review. Do not return the final unverified draft as if it passed. A critic can detect missing citations but should not become the authorization service.
For orchestrator–workers, give each worker a scoped input, allowed tools, output schema, deadline and budget. Delegate independent investigations, not competing writes to the same record. The coordinator validates outputs and resolves contradictions. Three workers reading the same misleading evidence do not provide three independent confirmations.
4. Routing, cascading, fallback and speculation are different
| Technique | When another model runs | What it optimizes | Necessary safeguard |
|---|---|---|---|
| Routing | Before the answer, using task features | Match request to an appropriate configuration | Measure routing errors and out-of-distribution cases |
| Cascade | After a first attempt fails an acceptance rule | Use inexpensive attempts where they suffice | A calibrated gate; count both attempts on escalation |
| Provider fallback | After a classified service failure | Availability | Compatible policy, region, tools and remaining deadline |
| Hedged request | A duplicate starts after a delay while the first is still pending | Tail latency | Extra-capacity budget; safe read-only or idempotent operation |
| Draft-and-check workflow | After a draft, a checker assesses the answer | Application-level quality or cost | Judge errors and repair cost |
| Speculative decoding | During token generation, a target model verifies draft tokens | Inference latency/throughput | Correct acceptance/correction algorithm and serving compatibility |
Speculative decoding can preserve the target model's sampling distribution under the algorithm's assumptions. Asking a larger model whether a smaller model's full answer “looks correct” does not provide that property. Its speedup also depends on draft acceptance and hardware workload. See speculative decoding and the original algorithm.
Worked cascade: count the first attempt even on escalation
Assume 10,000 questions. A first model costs $0.002 per attempt, a gate costs $0.001, and 25% escalate to a second model costing $0.012. Ignore other costs initially to isolate the mechanism.
| Component | Calculation | Cost |
|---|---|---|
| First attempt | 10,000 × $0.002 | $20 |
| Acceptance gate | 10,000 × $0.001 | $10 |
| Escalated attempt | 2,500 × $0.012 | $30 |
| Total | $20 + $10 + $30 | $60 |
| Second model for every question | 10,000 × $0.012 | $120 |
This is a hypothetical 50% reduction in these model/gate charges, not proof of equal quality. If 7,500 responses are accepted early and 1% of those are materially wrong, 75 bad answers bypass the second model. Evaluate that error slice, false escalations, second-model failures and end-to-end latency.
The break-even escalation fraction is (0.012 − 0.002 − 0.001) / 0.012 = 75%. Above it, this cascade costs more than the second model alone under these assumptions. The formula changes when the second call uses additional context or when human review, tools and caching matter.
5. Reuse and resilience patterns
Exact and semantic caching
An exact response cache reuses a response for the same effective request. The key must include relevant identity/authorization scope, model and prompt versions, parameters, data version and locale. A matching question string alone is insufficient.
A semantic cache reuses responses for sufficiently equivalent requests. Embedding similarity does not establish equivalent permission, intent or freshness. “Cancel my plan” and “Do not cancel my plan” illustrate why lexical similarity can be dangerous. Learn and test thresholds on the actual workload; 0.95 has no universal meaning across embedding models and distance metrics.
Semantic caching is appropriate only when reuse is demonstrably safe. Prefer live authoritative queries for balances and permissions. Cache expensive reference material separately when the personalized answer cannot safely be reused.
Retry, circuit breaker and bulkhead
| Pattern | Mechanism | Common implementation error |
|---|---|---|
| Bounded retry | Retry selected transient failures with jitter and an overall deadline | Retry every error, including denied or malformed requests |
| Circuit breaker | Stop ordinary calls to a failing dependency; allow controlled recovery probes | Count every user error as an outage; allow unlimited half-open probes |
| Bulkhead | Isolate capacity by dependency or workload | Limit active calls but leave an unlimited waiting queue |
| Degraded mode | Return a useful reduced result | Quietly return stale or incomplete material as current and complete |
A semaphore limits active work, not queued work. Add bounded admission, queue deadlines and per-tenant fairness. Circuit state and probe leases must be concurrency-safe. A local breaker protects one process unless coordinated explicitly; a shared breaker introduces its own dependency.
A timed-out write may already have committed. Reconcile its operation ID before issuing a new write. Changing providers does not bypass this problem, an authorization denial, or a safety policy. See the reliability patterns chapter.
6. Context and cost are runtime budgets
For context packing, reserve output space and preserve the instruction hierarchy, current request, required evidence and valid tool-call/result pairs. Summarize or remove lower-priority history deliberately. Keeping only the most recent messages can discard a critical constraint or break the conversation protocol. Context engineering explains the selection process.
For cost control, reserve a maximum allowance before parallel work starts, then reconcile each call's reported usage. Include retries, failed billable attempts, tool charges and review. Subtracting a global token counter “before” and “after” a request misattributes usage when requests overlap.
Read diagram source
flowchart LR
A[Request budget] --> B[Atomic reservation]
B --> C[Bounded model and tool calls]
C --> D[Per-call usage records]
D --> E[Reconcile reservation]
E --> F[Task cost and outcome]
B -->|Insufficient budget| G[Defer or reduce scope]
A budget alert explains that spending happened; an enforced reservation limits what can start. Reserve conservatively and release unused capacity. A hard cap may interrupt useful work, so define the partial-result experience.
7. Interview practice
1. Search finds the right sentence, but the answer omits its exception. Which pattern helps?
First inspect whether the exception was indexed, retrieved, packed and used. Parent–child retrieval or structure-aware packing helps when the child lacks its necessary context. It will not fix a missing source revision. Evaluate answer completeness and extra token cost.
2. How would you distinguish routing from a cascade on a whiteboard?
Draw routing as a decision before either model executes. Draw a cascade as first attempt → acceptance gate → optional next attempt. Label the latter's accumulated cost and added latency. Neither requires a universal ranking of models by size.
3. The critic approves an answer that cites a nonexistent section. What failed?
The judge's rubric or evidence access may be inadequate. Resolve citation IDs against the versioned corpus with deterministic checks, then assess whether the cited text supports the claim. A second opinion without independent evidence is insufficient.
4. Does a five-call step limit bound cost?
Only partially. Calls can have different token lengths, tool charges and parallel fan-out. Bound tokens, elapsed time, concurrency and spend as well, using shared reservations across branches.
5. When would you avoid a manager agent?
When the sequence is known, one model call suffices, or subtasks share mutable state so coordination dominates useful work. A fixed workflow can be easier to test and operate. Add dynamic delegation when task-specific decomposition provides measured value.
6. Why might the cheapest cascade be the wrong production choice?
Its acceptance gate may silently pass costly errors or its escalation path may violate the latency target. Compare success rate by risk slice, total cost per success and end-to-end latency against the simple baseline.
Final notes and recall card
| Question to remember | What a strong answer includes |
|---|---|
| What failed? | A concrete request and observed failure stage |
| Why this pattern? | The mechanism that addresses that stage |
| What does it cost? | Calls, tokens, latency, operations and maintenance |
| What new failure appears? | Misrouting, stale reuse, unsafe authority or coordination |
| How will we know? | A baseline, representative evaluation and production outcome |
| When does it stop? | Acceptance criteria, deadlines, budgets and explicit failure states |
Closing answer: Start with a bounded workflow. Add retrieval, routing, iteration or delegation to address a measured limitation. Keep access control and side-effect safety in application code, and evaluate the complete path—including the gate that decides a cheap answer is good enough.