Learnastra AI SYSTEM DESIGNAnup Rai

Concept · Understand the mechanism

Production RAG at Scale

By Anup Rai15 min readReviewed September 2026

A production RAG system combines maintained knowledge sources, authorized retrieval and answer generation under explicit quality and operating requirements. Scaling it means sustaining useful, supported answers as traffic, content, permissions and failures change. Query throughput alone is an incomplete success metric.

This chapter develops an interview design from a simple baseline, identifies its limits, and adds complexity only where requirements justify it. All workload numbers are illustrative assumptions, not vendor benchmarks or reports of a real deployment.

1. Clarify scope and requirements

Design a knowledge assistant for 500 organizations with ten million documents in total. Assume 2,000 requests/s on average and bursts up to 10,000 requests/s. Documents include manuals, policies and support articles. The initial product answers questions; it does not execute account changes or invent missing private facts.

Functional requirements:

  1. Ingest supported sources and track updates, deletion and permissions.
  2. Retrieve applicable evidence within the caller's current access scope.
  3. Answer with inspectable source references or an explicit insufficient-evidence outcome.
  4. Support questions spanning several sources when the evidence permits it.
  5. Let operators diagnose failures and release improved retrieval/model configurations.

Non-functional requirements to agree with the interviewer:

  1. Latency: distinguish time to first answer token from time to complete the answer. A p99 two-second complete-answer target is much harder than two-second first-token latency.
  2. Quality: define supported task completion, critical errors and acceptable abstention by query class.
  3. Freshness: specify how soon source changes become searchable; permission revocation has its own enforcement deadline.
  4. Isolation: no cross-tenant or unauthorized within-tenant disclosure, including caches and citations.
  5. Availability: define the degraded outcomes allowed when an index, model or source is unavailable.
  6. Cost: a workload budget including ingestion, serving, evaluation and operational ownership.

For an output of 250 tokens, an illustrative decode speed of 100 tokens/s takes 2.5 seconds just to emit those tokens. That cannot meet a two-second complete-answer target even with instant retrieval. Negotiate response length, model performance, a first-token target or an asynchronous path for long reports; do not promise that adding replicas solves an impossible per-request budget.

2. Size the workload before choosing products

Assume an average of five chunks per document and 1,536-dimensional float16 vectors:

Quantity Calculation Result
Searchable chunks 10M documents × 5 50M chunks
Raw vector payload 50M × 1,536 × 2 bytes 153.6 GB
Three full copies of that payload 153.6 GB × 3 460.8 GB
Requests per average day 2,000 × 86,400 172.8M
Full-path peak requests if a validated answer cache hits 20% 10,000 × 0.8 8,000/s

These exclude index structures, lexical data, source text, metadata, replicas beyond those assumed, working memory and rebuild headroom. Five chunks per document is an assumption to validate against actual lengths and overlap. Cache hit rate must be measured; capacity planning also needs a cold-cache or cache-outage scenario.

At 8,000 uncached requests/s and 250 output tokens/request, generation demand reaches two million output tokens/s during the assumed peak. This is a substantial serving requirement. Confirm provider quotas or measured self-hosted capacity, regional availability and budget before drawing a single model box as though it had unlimited throughput.

3. Start with a minimal useful design

Architecture / visual model
flowchart LR S[Source connector] --> P[Parse and chunk with source versions] P --> I[Search index and metadata] U[Authenticated question] --> R[Retrieve permitted evidence] I --> R R --> G[Generate a supported answer] G --> C[Validate and render citations]
Read diagram source
flowchart LR
    S[Source connector] --> P[Parse and chunk with source versions]
    P --> I[Search index and metadata]
    U[Authenticated question] --> R[Retrieve permitted evidence]
    I --> R
    R --> G[Generate a supported answer]
    G --> C[Validate and render citations]

For an initial corpus, one service and a suitable database/search engine may satisfy the requirements. A shared logical index can be internally distributed; it is incorrect to say that a single logical index cannot scale.

The first design exposes specific problems: ingestion can compete with queries; failures can leave incomplete document versions; similar chunks can crowd out required evidence; a model timeout can consume the request deadline; and a cache added carelessly can expose another user's answer. These are reasons for the next changes.

4. Separate ingestion, query execution and release state

Architecture / visual model
flowchart TD SRC[Sources: files, records, knowledge tools] --> IQ[Durable change queue and checkpoints] IQ --> W[Bounded parse, chunk and embedding workers] W --> RAW[Versioned source and evidence store] W --> IDX[Lexical and vector indexes] W --> MAN[Publication manifest and lineage] AUTH[Identity and authorization policy] --> API[Query API and admission control] USER[Question] --> API API --> CACHE[Scoped cache lookup and validation] CACHE -->|Valid hit| RESP[Answer or explicit incomplete outcome] CACHE -->|Miss| ROUTE[Choose permitted evidence path] ROUTE --> RET[Parallel retrieval with deadlines] IDX --> RET MAN --> RET RET --> FETCH[Resolve current permitted evidence] RAW --> FETCH AUTH --> FETCH FETCH --> RANK[Rerank and pack evidence] RANK --> GEN[Model gateway: quotas, deadline, token budget] GEN --> VALIDATE[Claim and citation checks under product policy] VALIDATE --> RESP API --> TRACE[Restricted traces and operational metrics] RET --> TRACE GEN --> TRACE
Read diagram source
flowchart TD
    SRC[Sources: files, records, knowledge tools] --> IQ[Durable change queue and checkpoints]
    IQ --> W[Bounded parse, chunk and embedding workers]
    W --> RAW[Versioned source and evidence store]
    W --> IDX[Lexical and vector indexes]
    W --> MAN[Publication manifest and lineage]
    AUTH[Identity and authorization policy] --> API[Query API and admission control]
    USER[Question] --> API
    API --> CACHE[Scoped cache lookup and validation]
    CACHE -->|Valid hit| RESP[Answer or explicit incomplete outcome]
    CACHE -->|Miss| ROUTE[Choose permitted evidence path]
    ROUTE --> RET[Parallel retrieval with deadlines]
    IDX --> RET
    MAN --> RET
    RET --> FETCH[Resolve current permitted evidence]
    RAW --> FETCH
    AUTH --> FETCH
    FETCH --> RANK[Rerank and pack evidence]
    RANK --> GEN[Model gateway: quotas, deadline, token budget]
    GEN --> VALIDATE[Claim and citation checks under product policy]
    VALIDATE --> RESP
    API --> TRACE[Restricted traces and operational metrics]
    RET --> TRACE
    GEN --> TRACE

The manifest records which source version and representations are ready. Stage new versions, verify completeness, then publish them under a defined read policy. Separate queues and resource limits reduce contention; separate read replicas may help, but replication still consumes resources and introduces lag.

Use source/version identities for idempotent processing. A newer update must not be overwritten by a delayed older event. Failed records enter a visible quarantine with ownership and replay. Reconcile source inventory with published records to detect lost events, orphaned vectors and incomplete deletes. See data engineering.

5. Choose retrieval depth and route by evidence need

Request Appropriate starting path What to avoid
A greeting Simple response without knowledge retrieval Calling every search backend
Current private policy Authorized retrieval of the applicable source Guessing from pretrained knowledge
Current order status Scoped operational record lookup Treating a stale document index as the order database
Comparison across periods Retrieve each required fact, then compare Assuming every comparison requires an agent
Variable evidence-dependent investigation Bounded iterative retrieval Unbounded searches and model-controlled permissions
Exhaustive corpus summary Enumerate the eligible set and track coverage Claiming that top-k similarity search visited everything

A query beginning “what is” can require private or current facts. Prefix regexes do not establish that retrieval is unnecessary. A router's errors can remove needed evidence before the retriever runs; evaluate route confusion and downstream outcomes.

Domain routing can reduce fan-out, but uncertain or multi-domain questions need an appropriate broader path. A summary index can route to detailed chunks, yet omitted summary concepts create a recall ceiling. Keep a fallback for queries that the coarse representation cannot cover.

Choose exact, lexical, dense, graph or structured access according to the required evidence. Hybrid retrieval and reranking are useful options, not mandatory stages for every query.

6. Compare RAG with long-context input

A small permitted corpus can sometimes be supplied directly to a model. This avoids retrieval misses but may increase token processing, distract from relevant evidence and complicate repeated updates. Retrieval can select an evidence packet from a larger corpus, and a long-context model can then synthesize that packet.

Compare actual model limits, answer quality, prompt-cache conditions, cost, freshness and access scope. Not every current model supports a million tokens. Fitting within the limit does not prove that all relevant details are used accurately.

“Lost in the middle” describes measured sensitivity to evidence position in studied models/tasks. It is not a universal fixed percentage penalty, and RAG does not eliminate it: retrieved context can still be lengthy or poorly ordered. Test evidence ordering and context size for the selected model. Research paper.

Reserve space for instructions, conversation state, tool data and output before selecting evidence. Preserve necessary qualifiers and contradictions; simply dropping the oldest chunk can delete the fact that makes the answer valid. See context engineering.

7. Make caching safe before making it effective

Cache Reused work Required identity and validation
Exact answer Prior complete answer Query, relevant conversation state, scope, policy, source/configuration versions
Semantic answer Answer to an equivalent information need All exact-cache checks plus validated semantic equivalence
Retrieval results Candidate IDs or evidence packet Query representation, filters, source/index versions, current permissions
Document/representation Source bytes or computed derivatives Source version, parser/embedding configuration and permitted use
Model prefix/KV cache Repeated prompt processing Serving/provider cache contract and scope isolation

“Can a refurbished item be returned after 14 days?” and “after 40 days?” may be close in embedding space and require different answers. A cosine threshold such as 0.95 does not prove equivalence. Numeric bounds, dates, entities, negation and user-specific records deserve explicit checks; unsuitable query classes can bypass semantic answer caching entirely.

Store dependency IDs and versions, but recognize their limit. A new policy that was absent from the previous answer can invalidate it even though none of its cited documents changed. Use an appropriate corpus/policy epoch, source-change invalidation or conservative expiration. Negative/no-answer entries also need invalidation when new evidence arrives.

Authenticate before lookup, enforce scope in cache selection, and revalidate permissions/freshness before exposure. A raw query hash is not an adequate cross-user answer key. Use structured serialization to avoid ambiguous concatenations. Time-to-live is a staleness bound only under the stated policy; it is not immediate revocation.

Prevent cache stampedes with bounded request coalescing within the same valid scope. Do not coalesce different tenants merely because the query string matches. Record invalid-hit rejection, not just hit rate. A lower hit rate can be the correct result of tighter correctness requirements.

8. Work the latency budget and failure paths

Consider these assumed stage durations for one illustrative request:

Authentication/admission                       20 ms
Dense branch: embedding 70 + search 60        130 ms
Lexical branch, run concurrently               90 ms
Fusion                                        10 ms
Reranking                                    100 ms
Evidence fetch/packing                        40 ms
Complete generation                        1,200 ms
Queue/network allowance                      150 ms
Total: 20 + max(130, 90) + 10 + 100 + 40 + 1,200 + 150
     = 1,650 ms

Parallel retrieval saves the shorter branch's serial contribution; it does not remove query embedding from the dense branch. These durations are a worked budget, not evidence of a p99 guarantee. Adding per-stage p99 values also does not yield an exact end-to-end p99; measure the complete latency distribution under load.

Carry an absolute deadline through downstream calls. Bound queues and retries, cancel unnecessary work, and reserve time for the response. Batch compatible embedding/reranking requests while limiting added queue delay. Every queued request must resolve successfully, fail, or be cancelled; exceptions must not leave futures waiting forever. Validate batch result counts and mappings.

Failure Allowed response, if the product contract permits Essential restriction
Dense search unavailable Lexical-only result with measured degraded quality Do not label it the full retrieval path
Reranker times out Use first-stage order Only if minimum evidence requirements still hold
Generation unavailable Return permitted source links or an explicit outage Do not invent an answer
Authorization unavailable Deny protected evidence access Never bypass scope checks to improve uptime
Required source missing Clarify or report incomplete evidence A disclaimer does not validate unsupported claims

Silently dropping every failed branch can make a partial result look complete. A question requiring two domains must not receive a complete-sounding comparison when one domain failed.

Speculative retrieval while the user types may reduce latency but processes unsubmitted text and wastes work on changing queries. It needs an explicit product/privacy decision, cancellation and careful final-query matching. It is not a default requirement for RAG.

9. Bound corrective and agentic retrieval

Corrective retrieval can assess weak evidence and choose another permitted search. CRAG is a particular research design involving retrieval evaluation and corrective processing; Self-RAG uses trained retrieval/reflection mechanisms. Printing a confidence label from an arbitrary model does not reproduce either method or establish correctness. CRAG, Self-RAG.

  1. Preserve the original question and hard constraints through reformulation.
  2. Track evidence found, missing facts, repeated queries and remaining budget.
  3. Restrict source selection and tool access in application code.
  4. Stop on success, missing user input, exhausted budget or no progress.
  5. Return an explicit incomplete outcome when required evidence remains absent.

Do not automatically send a private query to public web search after a low relevance score. A public fallback must be permitted for that data and information need. Model self-reported confidence is not a calibrated probability; route using validated task checks and measured risk. See agentic RAG.

10. Partition for capacity and isolation

Arrangement Benefit Responsibility
Shared index with enforced tenant/document filtering Efficient shared capacity Correct filter enforcement, quotas and scoped caches
Dedicated tenant index Easier placement and some operational separation More indexes and potentially idle capacity
Dedicated infrastructure Stronger resource and failure isolation Higher provisioning and maintenance cost
Hybrid placement Match exceptional tenants to special requirements Routing, migration and policy consistency

These are degrees of deployment separation, not automatic security guarantees. A dedicated index can still leak through an application bug, shared cache or asset URL. A pooled index can enforce strong authorization when correctly designed. Tenant membership alone is insufficient when documents have different permissions within the tenant. Microsoft's multitenant RAG guidance.

Derive tenant identity from authenticated state. Do not trust a request body's tenant ID or use a language model to decide access. Apply mandatory constraints during retrieval where supported, and verify current authorization before evidence leaves the trusted data layer. Runtime security checks must not rely on language assertions that can be disabled.

Hash partitioning spreads data but can require broad query fan-out. Time/domain partitions can reduce selected queries' work but create skew or cross-partition requests. Tenant partitioning can produce hot large tenants and many small partitions. Replicas add read capacity and failure tolerance; shards distribute data. Neither removes the need to test selective filtering and p99 behavior.

Use per-tenant admission, concurrent-work and token budgets as well as request counts. A single giant prompt can cost more than many short queries. The budget decision and reservation must be atomic; “read current spend, check, then increment” races across workers. Settle reservations against actual usage and recover abandoned reservations conservatively.

Likewise, a fixed-window rate limiter needs atomic counter/expiry behavior and clear window boundaries. Resetting a one-second TTL on every request can keep a busy counter alive indefinitely. Redis documents atomic scripting patterns for this concern; choose the algorithm that matches the desired burst behavior. Redis INCR documentation.

11. Calculate cost across the whole day

For sustained rate q, full-answer cache hit fraction h and average uncached variable cost c:

daily uncached variable cost = q × 86,400 × (1 − h) × c

At the assumed average 2,000 requests/s, 20% cache hits and an illustrative $0.0015 per uncached request, the result is $207,360/day. Cache service, ingestion, index infrastructure, evaluation and operations are additional. This is an arithmetic scenario, not quoted provider pricing.

Sustained rate Cost per request, no answer cache Daily variable cost
1,000/s $0.0001 $8,640
1,000/s $0.001 $86,400
10,000/s $0.001 $864,000

A cost proposal that ignores the 86,400 seconds in a day can be wrong by orders of magnitude. Distinguish peak capacity from average billable traffic, and include tokens for planning, retries, reranking and evaluation.

When cost rises at constant request volume, inspect input/output lengths, model routing, cache validity/hit rate, retrieval rounds, retries, provider prices and duplicate ingestion. Route proven simple tasks to a cheaper model when evaluation supports it. Do not save money by answering before the evidence is sufficient or by weakening access checks.

12. Keep source and model migrations reversible

Do not re-embed a vector simply because it is ninety days old. Reprocess when its source, embedding contract, parser or relevant representation changes, or when measured quality requires a new approach. Model drift and source freshness are different problems.

For an embedding migration, build a new index with a compatible query encoder, compare it on held-out queries, shadow suitable traffic and switch a versioned release pointer. Maintain rollback within data-retention policy. A version switch cannot restore content whose permission was revoked or whose deletion is required.

Incremental ingestion requires checkpoints, backpressure and reconciliation. Deletion must reach source copies, derivatives, indexes, caches and citation assets. Tombstones/current authorization can stop exposure before asynchronous physical cleanup finishes, according to the required revocation contract.

13. Observe quality and operate incidents

Signal What it tells you What it does not prove
Complete latency and first-token latency User-visible speed Answer correctness
Queue depth, throttling and branch failures Capacity/dependency pressure Which evidence was semantically needed
Source freshness and publication gaps Data availability and lag Parsing preserved every qualifier
Labeled recall and claim support Measured retrieval/answer properties Truth beyond available labels/evidence
Scope violations and revocation tests Access-control behavior Safety of every untested derivative
Cost per supported task completion Efficiency of useful outcomes Unbiased quality if success labels are poor

Keep restricted traces with request IDs, selected routes, source/index/model versions, candidate and evidence IDs, timing, fallback state, token usage and outcomes. Raw prompts or source text may be sensitive; apply minimization, access controls and retention rather than logging everything by default.

Recall and semantic correctness usually require labels or calibrated review; they are not free real-time counters. Use representative samples plus targeted severe-error suites. Keep targeted investigations separate from population estimates. User thumbs-up and silence are not verified correctness. See RAG evaluation.

Set alerts from the actual SLO, change sensitivity and incident cost. A falling cache hit rate does not automatically justify lowering the similarity threshold. A latency regression may be model queueing, not a search-capacity problem. Correlate signals, inspect examples and assign an owner to the failing stage.

If three required stages each succeed with independent probability 0.95, their joint probability is 0.95³ ≈ 0.857. Independence is an assumption; shared outages and correlated content errors can invalidate that product. Measure complete outcomes instead of presenting multiplied stage averages as observed reliability.

14. Adapt the design to the task

For customer support, separate current account/order tools from document policy retrieval and scope both to the caller. Human escalation follows validated risk and evidence criteria, not an arbitrary model-generated confidence number.

For enterprise knowledge search, prioritize connectors, current document permissions, source versions and cross-domain evidence. Dedicated infrastructure may be appropriate for contractual requirements or large noisy tenants, not merely because a tenant buys a higher plan.

For a request to compare every contract in an eligible set, enumerate that set and report coverage. If 47 contracts are eligible and clauses are found in only 43, classify the remaining four: absent clause, parse failure, inaccessible source or unresolved search. Do not silently summarize 43 as though all 47 were processed. A coverage-tracked workflow may be more reliable than an open-ended agent.

Interview practice

Q1: How would you meet 10,000 QPS and a two-second p99?

First clarify whether two seconds means first token or complete answer, output length and traffic distribution. Quantify uncached requests and token throughput, validate service capacity and quotas, and allocate a measured deadline across stages. I would use admission control, scoped caching where safe, parallel independent retrieval and bounded queues. I would not assume cache hits or replicas make an infeasible complete-answer target possible.

Q2: Why can semantic caching return a dangerous answer?

Similar phrasing can hide different numbers, dates, permissions or entities. The cached answer may also become wrong after a new source appears. I restrict suitable query classes, validate equivalence and current scope, track versions/dependencies and treat freshness rejection as correct behavior.

Q3: What if retrieved passages are plausible but insufficient?

Check required fact coverage and applicability. Use a bounded permitted follow-up search, clarification or an insufficient-evidence outcome. A fluent answer with a disclaimer or high self-reported confidence does not repair missing evidence.

Q4: How do you choose pooled versus dedicated indexes?

Compare isolation requirements, tenant size/skew, filtering capabilities, failure boundaries and operational cost. Both require application authorization and scoped derivatives. Dedicated storage alone does not secure caches, model inputs or citations.

Q5: What causes a cost spike without more requests?

More tokens, expensive route selection, reduced valid cache reuse, repeated searches, retries, price changes or duplicate ingestion. Attribute usage by stage and version, then fix the measured cause while preserving quality and access requirements.

Q6: What does a partial retrieval outage mean for the answer?

It depends on the required evidence. A tested lexical-only fallback may satisfy one query, while a cross-domain comparison becomes incomplete. Propagate branch status, enforce the evidence contract and avoid presenting partial coverage as complete.

Q7: How do you ensure rate limits work across workers?

Make the admission decision and counter/reservation update atomic under the chosen consistency boundary. Define fixed/sliding/token-bucket semantics, expiration, retries and dependency-outage policy. Separate request limits from token, concurrency and spend limits.

Q8: What is your release process?

Version sources, indexes, encoders, prompts and routing policy. Run access/freshness checks and matched quality/latency/cost evaluations, inspect severe regressions, then use controlled exposure with rollback. Continue sampling live outcomes and reconcile ingestion/deletion state.

Closing recommendation

Start with an authorized, observable retrieval path and a source lifecycle that can be trusted. Add caching, routing, reranking, partitioning and agentic steps when their measured benefits exceed their cost and failure burden. Close the interview by naming the limiting resource, the remaining quality uncertainty, and the first experiment that would resolve it.

Recall card: Scope → capacity → baseline → evidence quality → isolation → deadlines → cost → lifecycle → measured release.

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← RAG Evaluation Patterns
NEXT LESSONData Engineering for AI →

Explore the diagram