Learnastra AI SYSTEM DESIGNAnup Rai

Complete design interview

Design an Enterprise Knowledge Assistant

By Anup Rai12 min readReviewed September 2026

Interview problem: design a read-only assistant that answers employee questions from internal policies, procedures and research, with inspectable citations and current access control.

This is a hypothetical Learnastra interview exercise. Scale, targets and costs are assumptions to challenge with the interviewer. The design is not a claim about a deployed financial-services system.

RAG, or retrieval-augmented generation, supplies external evidence to a model before it answers. In this case, success means finding the applicable, authorized version of a document and using it correctly. A fluent answer to the wrong policy version is a failure.

1. Clarify scope and requirements

Ask which sources are authoritative, whether permissions include individual exceptions, whether historical “as of” questions matter, and whether external model APIs are allowed. For this exercise, assume all document content, embeddings, inference and telemetry stay inside the organization's approved private environment.

Functional requirements

  1. Answer natural-language questions using internal documents and link each material factual claim to supporting passages.
  2. Support follow-up questions while preserving the employee's intended scope and date.
  3. Ingest new documents, updates, permission changes and deletions from approved sources.
  4. Restrict every search, passage, document view and saved answer to the user's current permitted scope.
  5. Clarify ambiguous requests and abstain when evidence is insufficient or conflicting.
  6. Allow feedback and operator investigation without exposing unnecessary private text.

Deferred: executing business actions, unrestricted web browsing, arbitrary SQL generation, and a universal summary of every document. These introduce different requirements and authority boundaries.

Nonfunctional requirements

  1. Support 500,000 documents and 5,000 employees; distinguish connected users from active requests.
  2. Target p95 end-to-end completion below five seconds for the agreed ordinary-question workload. Measure time to first useful output separately.
  3. Target 99.9% request availability over a defined monthly window, with a labeled search-only fallback when generation is unavailable.
  4. Make 99% of ordinary content updates searchable within one hour of source publication; expose backlog and source outages.
  5. Enforce permission revocations and deletions through a higher-priority deny path, with a separately agreed propagation bound. Never use the one-hour content target as permission to serve revoked data.
  6. On a reviewed, representative answerable set, target at least 90% fully correct, evidence-supported answers. Report abstention, coverage and uncertainty separately, plus severe-risk tests.
  7. Enforce request, tenant, token, queue and cost limits; retain auditable version IDs and decisions with a defined retention policy.

Interview tip: “90% accuracy” is incomplete until you state the cases, rubric and denominator. Answering only the easiest questions can inflate accuracy while making the product unhelpful.

2. Estimate the workload before naming databases

Assume ten searchable passages per document on average and 1,024 float32 values per dense vector.

Quantity Calculation Implication
Searchable passages 500,000 × 10 = 5 million Passage count drives the index
Raw vector bytes 5M × 1,024 × 4 = 20.48 GB Excludes text, graph index, metadata and allocator overhead
Two vector copies 40.96 GB of raw vectors Still not the total RAM/storage requirement
Heavy users 500 × 100 queries/day = 50,000/day The remaining employees may use it occasionally
Illustrative monthly volume 50,000 × 30 = 1.5M queries Explicit 30-day assumption
Average over 24 hours 50,000 / 86,400 ≈ 0.58 requests/s Misleading for daytime capacity if used alone
Average over an eight-hour workday 50,000 / 28,800 ≈ 1.74 requests/s Better starting assumption for this usage pattern
Illustrative 10× workday peak About 17.4 requests/s Validate actual burst distribution
In-flight work at three-second mean 17.4 × 3 ≈ 52 requests Little's-law sizing clue, not a tail-latency guarantee

Also measure source byte size, table/image prevalence, update rate, chunk-length distribution and multilingual mix. Quantization may reduce vector storage at a recall cost. A generator's weights, KV cache, batching and token mix determine GPU capacity; document count cannot size the generator.

If 1% of documents change daily, 5,000 documents × ten passages means roughly 50,000 passages to process per day before reuse/deduplication. Bursty source migrations can dominate that average.

3. Start with the smallest complete design

Architecture / visual model
flowchart LR U[Employee] --> A[Authenticated query API] A --> P[Current permission check] P --> S[Keyword search over approved corpus] S --> E[Authorized evidence and source IDs] E --> M[Private model endpoint] M --> V[Validate citations and current access] V --> R[Answer or abstention] C[Versioned source connector] --> I[Parse and index] I --> S
Read diagram source
flowchart LR
    U[Employee] --> A[Authenticated query API]
    A --> P[Current permission check]
    P --> S[Keyword search over approved corpus]
    S --> E[Authorized evidence and source IDs]
    E --> M[Private model endpoint]
    M --> V[Validate citations and current access]
    V --> R[Answer or abstention]
    C[Versioned source connector] --> I[Parse and index]
    I --> S

Pilot one well-owned corpus. Keyword search supplies a measurable baseline, particularly for policy IDs and exact terminology. Add a model answer only if it improves time to a correct result relative to source links alone.

Find the baseline's failures

Observed failure Change to test Benefit Added cost or new risk
Paraphrases miss the right policy Dense retrieval plus keyword search Broader candidate recall Embedding/index lifecycle and extra latency
Relevant passage ranks below weak matches Rerank a shortlist Better evidence ordering Compute and possible domain-specific ranking errors
Table exception is separated from its value Structural parsing and parent expansion Preserves interpretation Larger context; parent access must be checked
Answer mixes old/new policy Versioned publication and query validation Coherent evidence Catalog and reconciliation complexity
Cache survives access removal Policy epoch plus current authorization check Prevents reuse outside current scope More validation and lower hit rate
Model invents a citation Resolve cited IDs and assess claim support Detects malformed/unsupported answers False abstentions; support checking is imperfect

A large model context window does not remove the need for permissions, source selection, effective dates or a latency budget.

4. Detailed architecture: separate ingestion, query and control

Architecture / visual model
flowchart TD subgraph ING[Ingestion path in private environment] SRC[Approved source connectors] --> Q[Change queue and durable checkpoints] Q --> PAR[Parse structure and preserve source locations] PAR --> CH[Versioned passages and embeddings] CH --> VS[(Vector index)] CH --> KS[(Keyword index)] PAR --> DS[(Immutable document versions)] VS --> PUB[Readiness checks and publication CAS] KS --> PUB DS --> PUB end subgraph CTRL[Authoritative control state] CAT[(Published revisions and job states)] ACL[(Identity groups, ACLs and deny tombstones)] REC[Reconciliation and freshness monitor] end PUB --> CAT SRC -->|Revoke or delete| ACL REC -. checks .-> CAT REC -. repairs .-> Q subgraph QUERY[Interactive query path] U[Employee] --> API[Auth, admission and deadline] API --> AUTH[Resolve current access scope] AUTH --> RET[Parallel lexical and dense retrieval] RET --> VAL[Validate versions and authorization] VAL --> RR[Rerank and pack evidence] RR --> GEN[Private model pool] GEN --> OUT[Check claims, citations and access before release] OUT --> UI[Answer with versioned source links] end ACL --> AUTH ACL --> VAL CAT --> VAL RET --> VS RET --> KS VAL --> DS ACL --> OUT QUERY -. metadata and timings .-> OBS[Restricted evaluation and operations telemetry]
Read diagram source
flowchart TD
    subgraph ING[Ingestion path in private environment]
        SRC[Approved source connectors] --> Q[Change queue and durable checkpoints]
        Q --> PAR[Parse structure and preserve source locations]
        PAR --> CH[Versioned passages and embeddings]
        CH --> VS[(Vector index)]
        CH --> KS[(Keyword index)]
        PAR --> DS[(Immutable document versions)]
        VS --> PUB[Readiness checks and publication CAS]
        KS --> PUB
        DS --> PUB
    end
    subgraph CTRL[Authoritative control state]
        CAT[(Published revisions and job states)]
        ACL[(Identity groups, ACLs and deny tombstones)]
        REC[Reconciliation and freshness monitor]
    end
    PUB --> CAT
    SRC -->|Revoke or delete| ACL
    REC -. checks .-> CAT
    REC -. repairs .-> Q
    subgraph QUERY[Interactive query path]
        U[Employee] --> API[Auth, admission and deadline]
        API --> AUTH[Resolve current access scope]
        AUTH --> RET[Parallel lexical and dense retrieval]
        RET --> VAL[Validate versions and authorization]
        VAL --> RR[Rerank and pack evidence]
        RR --> GEN[Private model pool]
        GEN --> OUT[Check claims, citations and access before release]
        OUT --> UI[Answer with versioned source links]
    end
    ACL --> AUTH
    ACL --> VAL
    CAT --> VAL
    RET --> VS
    RET --> KS
    VAL --> DS
    ACL --> OUT
    QUERY -. metadata and timings .-> OBS[Restricted evaluation and operations telemetry]

Possible components include PostgreSQL for authoritative metadata, a private object store for source versions, Qdrant for dense retrieval, Elasticsearch/OpenSearch for lexical search, and a measured local model-serving pool. A single search engine supporting both retrieval modes can reduce operations; two engines permit independent tuning but create more publication and recovery work.

BGE-M3 provides a concrete 1,024-dimensional multilingual embedding baseline. Its availability does not make it the best September 2026 choice for every corpus. Compare current licensed candidates on the organization's own recall, latency and hardware constraints. Model names are configuration, not the architecture. An approved external API would be a separate data-boundary decision and needs an equivalent evaluation.

5. Define the API and records

Query contract

POST /v1/answers
Authenticated principal comes from the server session.
Body: {question, conversation_id?, as_of_date?}
Response: {request_id, status, answer, citations[], evidence_revision}
Citation: {document_id, version, passage_id, source_location}
Status: answered | needs_clarification | insufficient_evidence | unavailable

Do not accept a caller-supplied list of groups as proof of membership. Source links resolve through an authenticated document endpoint, not an unrestricted storage URL. An evidence_revision identifies the selected evidence set; it is not proof that every external source was globally synchronized at that instant.

Record Essential fields Purpose
Document source ID, current published version, effective dates, deletion state Authoritative lifecycle
Document version immutable content hash, parser version, source location Reproducible evidence
Passage document/version, passage ID, text/span, embedding version Retrieval and precise citation
Access policy resource, allow/deny rules, policy epoch Current authorization
Ingestion job source/version, stage states, retry count, error, timestamps Repair and freshness measurement
Answer metadata request ID, principal scope, evidence IDs, model/prompt versions, timings Investigation without mandatory full-text logging

Store source time and ingestion time separately. A future-effective policy can be published and searchable without being the policy applicable today.

6. Publish coherent document versions

Idempotency means repeating an operation has the same intended effect as performing it once. Use immutable (document_id, version, passage_id) identities. Idempotent writes replace the same versioned object on retry; they do not mutate the currently published version in place.

Architecture / visual model
sequenceDiagram participant C as Connector participant W as Worker participant I as Search indexes participant S as Document store participant P as Catalog C->>P: Register desired version 9 C->>W: Ingest version 9 W->>W: Parse, validate and embed par Prepare indexes W->>I: Write version 9 with deterministic IDs I-->>W: Confirm searchable readiness and Prepare source W->>S: Store immutable version 9 S-->>W: Confirm durable object end W->>P: Compare-and-swap publish if desired version is still 9 alt Version 10 already supersedes it P-->>W: Do not publish stale job else Still current and every stage ready P-->>W: Publish version 9 end
Read diagram source
sequenceDiagram
    participant C as Connector
    participant W as Worker
    participant I as Search indexes
    participant S as Document store
    participant P as Catalog
    C->>P: Register desired version 9
    C->>W: Ingest version 9
    W->>W: Parse, validate and embed
    par Prepare indexes
        W->>I: Write version 9 with deterministic IDs
        I-->>W: Confirm searchable readiness
    and Prepare source
        W->>S: Store immutable version 9
        S-->>W: Confirm durable object
    end
    W->>P: Compare-and-swap publish if desired version is still 9
    alt Version 10 already supersedes it
        P-->>W: Do not publish stale job
    else Still current and every stage ready
        P-->>W: Publish version 9
    end

“Write acknowledged” may precede search visibility. The adapters must meet the readiness contract, including the engines' refresh/consistency settings, before publication. Search results must match the catalog's published revision; a ready=true flag alone can leave several versions visible.

For ordinary content updates, keep the prior published version until the replacement is ready and label freshness where relevant. For revocation or deletion, deny immediately through authoritative control state even if physical index removal is delayed. Keeping the old readable version is not an acceptable fallback for a revoked document.

Avoid shipping a list of all 500,000 document/version pairs with every query. Use engine-supported authorization filters and partitioning, then validate candidate versions and ACLs against the authoritative catalog before reranking or generation. Bound refilling when stale candidates reduce recall. A large stale-index backlog can produce abstentions even when source evidence exists, so monitor it and reconcile.

7. Retrieve, rank and build the answer

  1. Authenticate, apply admission limits and resolve the request's date and access scope.
  2. Resolve follow-up references without allowing conversation text to change authority.
  3. Search permitted lexical and dense candidates in parallel.
  4. Deduplicate by versioned passage ID and discard stale, deleted or unauthorized candidates.
  5. Combine ranks, optionally rerank, then select a bounded evidence set.
  6. Expand necessary parent context while rechecking its access and version.
  7. Pack evidence with stable source IDs; reserve output tokens and preserve exceptions.
  8. Generate a structured answer with cited source IDs and an explicit insufficiency path.
  9. Validate citation IDs, source versions and current access; assess whether material claims are supported.
  10. Return the answer or a precise clarification/abstention and record outcome metadata.

Rank fusion example

Reciprocal rank fusion combines rank positions without requiring comparable engine scores. With rank starting at one:

score(document) = Σ weight_i / (k + rank_i(document))

For equal weights and illustrative k=60, passage A ranked 1st and 10th scores 1/61 + 1/70 ≈ 0.03068; passage B ranked 3rd in both scores 2/63 ≈ 0.03175, so B ranks ahead. A missing document contributes zero for that list. Use each document once per list. Tune the candidate counts and fusion constant on development data, not a protected release set.

The executable fusion below uses one contribution per passage per list. It returns the supplied passage objects; filtering and version validation happen before this function.

def reciprocal_rank_fusion(result_lists, weights=(0.5, 0.5), k=60):
    import math
    if len(result_lists) != len(weights) or not math.isfinite(k) or k <= 0:
        raise ValueError("Use a positive finite k and one weight per list")
    if any(not math.isfinite(w) or w < 0 for w in weights) or not any(weights):
        raise ValueError("Weights must be finite, nonnegative and not all zero")
    scores, records = {}, {}
    for results, weight in zip(result_lists, weights):
        seen = set()
        rank = 0
        for passage in results:
            if passage.id in seen:
                continue
            seen.add(passage.id)
            rank += 1
            scores[passage.id] = scores.get(passage.id, 0) + weight / (k + rank)
            records[passage.id] = passage
    ordered = sorted(scores, key=lambda key: (-scores[key], key))
    return [records[key] for key in ordered]

An initial experiment might retrieve up to 100 per engine, fuse to 50, rerank to ten and pack within 12,000 evidence tokens. These are experiment settings, not promises of good retrieval. The reranking chapter explains why a shortlist cannot recover omitted evidence.

Generation with bounded evidence

A citation to a real paragraph may still fail to support the claim. Check IDs deterministically, then evaluate support using reviewed labels and appropriate verifiers. Treat retrieved text as evidence, not instructions to override the application. Confidence should be calibrated against outcomes; a similarity score is not the probability the final answer is correct.

For an illustrative five-second completion budget, allocate 0.2 seconds to admission/auth, 0.5 to retrieval and validation, 0.5 to reranking/packing, 3.3 to generation and 0.5 to final checks. These sum to five seconds as a planning allocation. Adding component p95s does not establish an end-to-end p95; test the complete path under load.

Do not stream unchecked sensitive content and then claim a final access check protects it. Buffer until required checks pass, or design an explicitly bounded streaming validation/revocation policy. Access may change during a request: define when a decision is authoritative, use current policy epochs at release, and cancel invalidated work. Already delivered bytes cannot be recalled.

8. Cache, scale and recover

Concern Mechanism Failure test
Repeated safe questions Exact cache keyed by principal/scope, question/history, policy epoch, evidence and model/prompt versions Same wording after revocation or policy change
Semantic cache Add only after equivalence and freshness evaluation Negation, dates, departments and near-matching questions
Burst traffic Bounded queues, per-tenant fairness and token-aware admission Slow generations monopolize the pool
Search growth Shards plus replicas sized from measured passage/index workload Node loss during ingestion and query load
Embedding migration Separate index version, dual evaluation, controlled cutover Old/new vectors mixed in one incompatible space
Model outage Circuit breaker and labeled authorized-search fallback Fallback silently presented as a complete answer
Source outage Lag alerts, current published content where still permitted Silent hours-long freshness breach

Qdrant's optimizer threshold is based on vector-data size, not a count of documents. Avoid memorizing “four shards and three nodes” as a capacity solution. Verify filtered recall, index build time, replica recovery and full-load latency on the actual distribution.

TTL is a cleanup/freshness mechanism, not authorization. Check source access on cache hits and document views. Saved answers and conversation memory need the same permission and retention policy. If authoritative access state is unavailable, fail closed for protected content rather than trusting an old allow decision.

9. Cost worksheet and decision tradeoffs

For a strict-local design, use measured GPU capacity and total operating costs. The following is an illustrative monthly budget, not a hardware quote:

Line Monthly allocation
Active generation fleet including peak headroom $9,000
Failure-reserve generation capacity $4,500
Embedding and reranking $1,500
Search and metadata services $2,000
Storage, backups and networking $1,000
Observability and evaluation infrastructure $1,000
Allocated engineering, on-call and source maintenance $12,000
Total $31,000

At 1.5M requests this is approximately $0.02067/request. At 90% verified success, $31,000 / 1.35M ≈ $0.02296/success. At half the volume on the same fleet, cost per request doubles. Do not add a second idle-capacity charge when that capacity is already included.

For comparison only, assume an external service were permitted and its hypothetical rates were $2/M input tokens and $12/M output tokens. A 2,000-input/500-output call costs $0.004 + $0.006 = $0.010, or $15,000 for 1.5M calls. Add embedding, retrieval, retries, evaluation, staff and review before comparing it with the full local budget. These are arithmetic assumptions, not current vendor prices.

No cost saving overrides the private-processing requirement. Once two options are feasible, compare quality, latency, capacity utilization, operator effort and migration cost. Batch offline embeddings where useful; do not force interactive questions to wait for an offline batching window.

10. Evaluation, rollout and ownership

  1. Collect representative questions with reviewed source passages and effective dates.
  2. Evaluate retrieval coverage, answer correctness/support, abstention and citation quality separately.
  3. Include tables, acronyms, languages, conflicting versions and multi-document answers.
  4. Test permission revocation during retrieval, generation, cache lookup and saved-answer access.
  5. Test partial ingestion, duplicate events, out-of-order versions and deletion while a worker is retrying.
  6. Load-test the expected token distribution, burst rate and node-loss condition.
  7. Pilot one corpus with source owners, compare against ordinary search, then expand.

Assign owners for source freshness, authorization, retrieval quality, answer evaluation and user recovery. Record request IDs and stage timings by default; full query/document logging requires a specific purpose, access control and retention. See RAG evaluation.

A source ingestion success rate can look healthy while one important policy remains stale. Monitor oldest unprocessed change, publication lag by source, tombstone propagation and reconciliation discrepancies. Roll back model/prompt/index changes independently where possible; a rollback must not restore revoked access or deleted content.

Interview follow-ups

1. Why not place the whole archive in a large context window? It exceeds practical per-request evidence, cost and latency budgets, mixes applicability and permissions, and makes citations difficult. A bounded authorized document packet can be a useful long-context alternative; evaluate that specific workload.

2. The vector write succeeds and keyword indexing fails. What is visible? The old permitted published revision remains available for ordinary updates. The new revision is not published until its required stages are searchable. A delete/revoke event separately denies access regardless of ingestion success.

3. How does a policy change affect a cached answer? Its evidence/version key becomes stale and the current publication/access checks reject reuse. Propagate invalidation, but do not depend on invalidation alone to enforce permissions.

4. Why can a faithful answer be wrong? It may faithfully repeat an obsolete policy, one applying to a different department/date, or an erroneous source. Evaluate applicability and source authority as well as textual support.

5. What would you investigate if users like the answers but cannot complete tasks? Satisfaction may reward fluency. Compare supported correctness, time to the authoritative source, repeated searches and real task outcomes. Review the unsuccessful cohort and source coverage.

6. What changes at ten times the load? Measure the new peak token throughput, index filtering cost and ingestion backlog. Add admission, batching and serving capacity according to bottlenecks; confirm failover capacity. Multiplying replicas without testing the model pool and metadata service is insufficient.

Closing remarks and recall table

Remember Explain it concretely
Applicable Correct source, effective date and document version
Authorized Current access before evidence processing and release
Supported Claims trace to inspectable passages
Recoverable Versioned jobs, idempotent writes and reconciliation
Measured Complete-task quality, freshness, latency and cost

60-second interview answer

I would build a read-only assistant over one owned corpus, starting with authenticated search and a model that answers from selected evidence. Versioned ingestion keeps ordinary updates coherent, while a separate deny path handles deletions and revocations. I would test hybrid retrieval and reranking against the baseline, verify citations and applicability, and measure supported answers, freshness, latency and full operating cost. Expansion follows demonstrated recovery and authorization behavior under load.

Remember: Applicable evidence → Current access → Supported answer → Measured outcome.

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← AI Anti-Patterns: Diagnose the Assumption Before Replacing the Design
NEXT LESSONDesign a Conversational Customer-Support Agent →

Explore the diagram