Interview problem: design a read-only assistant that answers employee questions from internal policies, procedures and research, with inspectable citations and current access control.
This is a hypothetical Learnastra interview exercise. Scale, targets and costs are assumptions to challenge with the interviewer. The design is not a claim about a deployed financial-services system.
RAG, or retrieval-augmented generation, supplies external evidence to a model before it answers. In this case, success means finding the applicable, authorized version of a document and using it correctly. A fluent answer to the wrong policy version is a failure.
1. Clarify scope and requirements
Ask which sources are authoritative, whether permissions include individual exceptions, whether historical “as of” questions matter, and whether external model APIs are allowed. For this exercise, assume all document content, embeddings, inference and telemetry stay inside the organization's approved private environment.
Functional requirements
- Answer natural-language questions using internal documents and link each material factual claim to supporting passages.
- Support follow-up questions while preserving the employee's intended scope and date.
- Ingest new documents, updates, permission changes and deletions from approved sources.
- Restrict every search, passage, document view and saved answer to the user's current permitted scope.
- Clarify ambiguous requests and abstain when evidence is insufficient or conflicting.
- Allow feedback and operator investigation without exposing unnecessary private text.
Deferred: executing business actions, unrestricted web browsing, arbitrary SQL generation, and a universal summary of every document. These introduce different requirements and authority boundaries.
Nonfunctional requirements
- Support 500,000 documents and 5,000 employees; distinguish connected users from active requests.
- Target p95 end-to-end completion below five seconds for the agreed ordinary-question workload. Measure time to first useful output separately.
- Target 99.9% request availability over a defined monthly window, with a labeled search-only fallback when generation is unavailable.
- Make 99% of ordinary content updates searchable within one hour of source publication; expose backlog and source outages.
- Enforce permission revocations and deletions through a higher-priority deny path, with a separately agreed propagation bound. Never use the one-hour content target as permission to serve revoked data.
- On a reviewed, representative answerable set, target at least 90% fully correct, evidence-supported answers. Report abstention, coverage and uncertainty separately, plus severe-risk tests.
- Enforce request, tenant, token, queue and cost limits; retain auditable version IDs and decisions with a defined retention policy.
Interview tip: “90% accuracy” is incomplete until you state the cases, rubric and denominator. Answering only the easiest questions can inflate accuracy while making the product unhelpful.
2. Estimate the workload before naming databases
Assume ten searchable passages per document on average and 1,024 float32 values per dense vector.
| Quantity | Calculation | Implication |
|---|---|---|
| Searchable passages | 500,000 × 10 = 5 million | Passage count drives the index |
| Raw vector bytes | 5M × 1,024 × 4 = 20.48 GB | Excludes text, graph index, metadata and allocator overhead |
| Two vector copies | 40.96 GB of raw vectors | Still not the total RAM/storage requirement |
| Heavy users | 500 × 100 queries/day = 50,000/day | The remaining employees may use it occasionally |
| Illustrative monthly volume | 50,000 × 30 = 1.5M queries | Explicit 30-day assumption |
| Average over 24 hours | 50,000 / 86,400 ≈ 0.58 requests/s | Misleading for daytime capacity if used alone |
| Average over an eight-hour workday | 50,000 / 28,800 ≈ 1.74 requests/s | Better starting assumption for this usage pattern |
| Illustrative 10× workday peak | About 17.4 requests/s | Validate actual burst distribution |
| In-flight work at three-second mean | 17.4 × 3 ≈ 52 requests | Little's-law sizing clue, not a tail-latency guarantee |
Also measure source byte size, table/image prevalence, update rate, chunk-length distribution and multilingual mix. Quantization may reduce vector storage at a recall cost. A generator's weights, KV cache, batching and token mix determine GPU capacity; document count cannot size the generator.
If 1% of documents change daily, 5,000 documents × ten passages means roughly 50,000 passages to process per day before reuse/deduplication. Bursty source migrations can dominate that average.
3. Start with the smallest complete design
Read diagram source
flowchart LR
U[Employee] --> A[Authenticated query API]
A --> P[Current permission check]
P --> S[Keyword search over approved corpus]
S --> E[Authorized evidence and source IDs]
E --> M[Private model endpoint]
M --> V[Validate citations and current access]
V --> R[Answer or abstention]
C[Versioned source connector] --> I[Parse and index]
I --> S
Pilot one well-owned corpus. Keyword search supplies a measurable baseline, particularly for policy IDs and exact terminology. Add a model answer only if it improves time to a correct result relative to source links alone.
Find the baseline's failures
| Observed failure | Change to test | Benefit | Added cost or new risk |
|---|---|---|---|
| Paraphrases miss the right policy | Dense retrieval plus keyword search | Broader candidate recall | Embedding/index lifecycle and extra latency |
| Relevant passage ranks below weak matches | Rerank a shortlist | Better evidence ordering | Compute and possible domain-specific ranking errors |
| Table exception is separated from its value | Structural parsing and parent expansion | Preserves interpretation | Larger context; parent access must be checked |
| Answer mixes old/new policy | Versioned publication and query validation | Coherent evidence | Catalog and reconciliation complexity |
| Cache survives access removal | Policy epoch plus current authorization check | Prevents reuse outside current scope | More validation and lower hit rate |
| Model invents a citation | Resolve cited IDs and assess claim support | Detects malformed/unsupported answers | False abstentions; support checking is imperfect |
A large model context window does not remove the need for permissions, source selection, effective dates or a latency budget.
4. Detailed architecture: separate ingestion, query and control
Read diagram source
flowchart TD
subgraph ING[Ingestion path in private environment]
SRC[Approved source connectors] --> Q[Change queue and durable checkpoints]
Q --> PAR[Parse structure and preserve source locations]
PAR --> CH[Versioned passages and embeddings]
CH --> VS[(Vector index)]
CH --> KS[(Keyword index)]
PAR --> DS[(Immutable document versions)]
VS --> PUB[Readiness checks and publication CAS]
KS --> PUB
DS --> PUB
end
subgraph CTRL[Authoritative control state]
CAT[(Published revisions and job states)]
ACL[(Identity groups, ACLs and deny tombstones)]
REC[Reconciliation and freshness monitor]
end
PUB --> CAT
SRC -->|Revoke or delete| ACL
REC -. checks .-> CAT
REC -. repairs .-> Q
subgraph QUERY[Interactive query path]
U[Employee] --> API[Auth, admission and deadline]
API --> AUTH[Resolve current access scope]
AUTH --> RET[Parallel lexical and dense retrieval]
RET --> VAL[Validate versions and authorization]
VAL --> RR[Rerank and pack evidence]
RR --> GEN[Private model pool]
GEN --> OUT[Check claims, citations and access before release]
OUT --> UI[Answer with versioned source links]
end
ACL --> AUTH
ACL --> VAL
CAT --> VAL
RET --> VS
RET --> KS
VAL --> DS
ACL --> OUT
QUERY -. metadata and timings .-> OBS[Restricted evaluation and operations telemetry]
Possible components include PostgreSQL for authoritative metadata, a private object store for source versions, Qdrant for dense retrieval, Elasticsearch/OpenSearch for lexical search, and a measured local model-serving pool. A single search engine supporting both retrieval modes can reduce operations; two engines permit independent tuning but create more publication and recovery work.
BGE-M3 provides a concrete 1,024-dimensional multilingual embedding baseline. Its availability does not make it the best September 2026 choice for every corpus. Compare current licensed candidates on the organization's own recall, latency and hardware constraints. Model names are configuration, not the architecture. An approved external API would be a separate data-boundary decision and needs an equivalent evaluation.
5. Define the API and records
Query contract
POST /v1/answers
Authenticated principal comes from the server session.
Body: {question, conversation_id?, as_of_date?}
Response: {request_id, status, answer, citations[], evidence_revision}
Citation: {document_id, version, passage_id, source_location}
Status: answered | needs_clarification | insufficient_evidence | unavailable
Do not accept a caller-supplied list of groups as proof of membership. Source links resolve through an authenticated document endpoint, not an unrestricted storage URL. An evidence_revision identifies the selected evidence set; it is not proof that every external source was globally synchronized at that instant.
| Record | Essential fields | Purpose |
|---|---|---|
| Document | source ID, current published version, effective dates, deletion state | Authoritative lifecycle |
| Document version | immutable content hash, parser version, source location | Reproducible evidence |
| Passage | document/version, passage ID, text/span, embedding version | Retrieval and precise citation |
| Access policy | resource, allow/deny rules, policy epoch | Current authorization |
| Ingestion job | source/version, stage states, retry count, error, timestamps | Repair and freshness measurement |
| Answer metadata | request ID, principal scope, evidence IDs, model/prompt versions, timings | Investigation without mandatory full-text logging |
Store source time and ingestion time separately. A future-effective policy can be published and searchable without being the policy applicable today.
6. Publish coherent document versions
Idempotency means repeating an operation has the same intended effect as performing it once. Use immutable (document_id, version, passage_id) identities. Idempotent writes replace the same versioned object on retry; they do not mutate the currently published version in place.
Read diagram source
sequenceDiagram
participant C as Connector
participant W as Worker
participant I as Search indexes
participant S as Document store
participant P as Catalog
C->>P: Register desired version 9
C->>W: Ingest version 9
W->>W: Parse, validate and embed
par Prepare indexes
W->>I: Write version 9 with deterministic IDs
I-->>W: Confirm searchable readiness
and Prepare source
W->>S: Store immutable version 9
S-->>W: Confirm durable object
end
W->>P: Compare-and-swap publish if desired version is still 9
alt Version 10 already supersedes it
P-->>W: Do not publish stale job
else Still current and every stage ready
P-->>W: Publish version 9
end
“Write acknowledged” may precede search visibility. The adapters must meet the readiness contract, including the engines' refresh/consistency settings, before publication. Search results must match the catalog's published revision; a ready=true flag alone can leave several versions visible.
For ordinary content updates, keep the prior published version until the replacement is ready and label freshness where relevant. For revocation or deletion, deny immediately through authoritative control state even if physical index removal is delayed. Keeping the old readable version is not an acceptable fallback for a revoked document.
Avoid shipping a list of all 500,000 document/version pairs with every query. Use engine-supported authorization filters and partitioning, then validate candidate versions and ACLs against the authoritative catalog before reranking or generation. Bound refilling when stale candidates reduce recall. A large stale-index backlog can produce abstentions even when source evidence exists, so monitor it and reconcile.
7. Retrieve, rank and build the answer
- Authenticate, apply admission limits and resolve the request's date and access scope.
- Resolve follow-up references without allowing conversation text to change authority.
- Search permitted lexical and dense candidates in parallel.
- Deduplicate by versioned passage ID and discard stale, deleted or unauthorized candidates.
- Combine ranks, optionally rerank, then select a bounded evidence set.
- Expand necessary parent context while rechecking its access and version.
- Pack evidence with stable source IDs; reserve output tokens and preserve exceptions.
- Generate a structured answer with cited source IDs and an explicit insufficiency path.
- Validate citation IDs, source versions and current access; assess whether material claims are supported.
- Return the answer or a precise clarification/abstention and record outcome metadata.
Rank fusion example
Reciprocal rank fusion combines rank positions without requiring comparable engine scores. With rank starting at one:
score(document) = Σ weight_i / (k + rank_i(document))
For equal weights and illustrative k=60, passage A ranked 1st and 10th scores 1/61 + 1/70 ≈ 0.03068; passage B ranked 3rd in both scores 2/63 ≈ 0.03175, so B ranks ahead. A missing document contributes zero for that list. Use each document once per list. Tune the candidate counts and fusion constant on development data, not a protected release set.
The executable fusion below uses one contribution per passage per list. It returns the supplied passage objects; filtering and version validation happen before this function.
def reciprocal_rank_fusion(result_lists, weights=(0.5, 0.5), k=60):
import math
if len(result_lists) != len(weights) or not math.isfinite(k) or k <= 0:
raise ValueError("Use a positive finite k and one weight per list")
if any(not math.isfinite(w) or w < 0 for w in weights) or not any(weights):
raise ValueError("Weights must be finite, nonnegative and not all zero")
scores, records = {}, {}
for results, weight in zip(result_lists, weights):
seen = set()
rank = 0
for passage in results:
if passage.id in seen:
continue
seen.add(passage.id)
rank += 1
scores[passage.id] = scores.get(passage.id, 0) + weight / (k + rank)
records[passage.id] = passage
ordered = sorted(scores, key=lambda key: (-scores[key], key))
return [records[key] for key in ordered]
An initial experiment might retrieve up to 100 per engine, fuse to 50, rerank to ten and pack within 12,000 evidence tokens. These are experiment settings, not promises of good retrieval. The reranking chapter explains why a shortlist cannot recover omitted evidence.
Generation with bounded evidence
A citation to a real paragraph may still fail to support the claim. Check IDs deterministically, then evaluate support using reviewed labels and appropriate verifiers. Treat retrieved text as evidence, not instructions to override the application. Confidence should be calibrated against outcomes; a similarity score is not the probability the final answer is correct.
For an illustrative five-second completion budget, allocate 0.2 seconds to admission/auth, 0.5 to retrieval and validation, 0.5 to reranking/packing, 3.3 to generation and 0.5 to final checks. These sum to five seconds as a planning allocation. Adding component p95s does not establish an end-to-end p95; test the complete path under load.
Do not stream unchecked sensitive content and then claim a final access check protects it. Buffer until required checks pass, or design an explicitly bounded streaming validation/revocation policy. Access may change during a request: define when a decision is authoritative, use current policy epochs at release, and cancel invalidated work. Already delivered bytes cannot be recalled.
8. Cache, scale and recover
| Concern | Mechanism | Failure test |
|---|---|---|
| Repeated safe questions | Exact cache keyed by principal/scope, question/history, policy epoch, evidence and model/prompt versions | Same wording after revocation or policy change |
| Semantic cache | Add only after equivalence and freshness evaluation | Negation, dates, departments and near-matching questions |
| Burst traffic | Bounded queues, per-tenant fairness and token-aware admission | Slow generations monopolize the pool |
| Search growth | Shards plus replicas sized from measured passage/index workload | Node loss during ingestion and query load |
| Embedding migration | Separate index version, dual evaluation, controlled cutover | Old/new vectors mixed in one incompatible space |
| Model outage | Circuit breaker and labeled authorized-search fallback | Fallback silently presented as a complete answer |
| Source outage | Lag alerts, current published content where still permitted | Silent hours-long freshness breach |
Qdrant's optimizer threshold is based on vector-data size, not a count of documents. Avoid memorizing “four shards and three nodes” as a capacity solution. Verify filtered recall, index build time, replica recovery and full-load latency on the actual distribution.
TTL is a cleanup/freshness mechanism, not authorization. Check source access on cache hits and document views. Saved answers and conversation memory need the same permission and retention policy. If authoritative access state is unavailable, fail closed for protected content rather than trusting an old allow decision.
9. Cost worksheet and decision tradeoffs
For a strict-local design, use measured GPU capacity and total operating costs. The following is an illustrative monthly budget, not a hardware quote:
| Line | Monthly allocation |
|---|---|
| Active generation fleet including peak headroom | $9,000 |
| Failure-reserve generation capacity | $4,500 |
| Embedding and reranking | $1,500 |
| Search and metadata services | $2,000 |
| Storage, backups and networking | $1,000 |
| Observability and evaluation infrastructure | $1,000 |
| Allocated engineering, on-call and source maintenance | $12,000 |
| Total | $31,000 |
At 1.5M requests this is approximately $0.02067/request. At 90% verified success, $31,000 / 1.35M ≈ $0.02296/success. At half the volume on the same fleet, cost per request doubles. Do not add a second idle-capacity charge when that capacity is already included.
For comparison only, assume an external service were permitted and its hypothetical rates were $2/M input tokens and $12/M output tokens. A 2,000-input/500-output call costs $0.004 + $0.006 = $0.010, or $15,000 for 1.5M calls. Add embedding, retrieval, retries, evaluation, staff and review before comparing it with the full local budget. These are arithmetic assumptions, not current vendor prices.
No cost saving overrides the private-processing requirement. Once two options are feasible, compare quality, latency, capacity utilization, operator effort and migration cost. Batch offline embeddings where useful; do not force interactive questions to wait for an offline batching window.
10. Evaluation, rollout and ownership
- Collect representative questions with reviewed source passages and effective dates.
- Evaluate retrieval coverage, answer correctness/support, abstention and citation quality separately.
- Include tables, acronyms, languages, conflicting versions and multi-document answers.
- Test permission revocation during retrieval, generation, cache lookup and saved-answer access.
- Test partial ingestion, duplicate events, out-of-order versions and deletion while a worker is retrying.
- Load-test the expected token distribution, burst rate and node-loss condition.
- Pilot one corpus with source owners, compare against ordinary search, then expand.
Assign owners for source freshness, authorization, retrieval quality, answer evaluation and user recovery. Record request IDs and stage timings by default; full query/document logging requires a specific purpose, access control and retention. See RAG evaluation.
A source ingestion success rate can look healthy while one important policy remains stale. Monitor oldest unprocessed change, publication lag by source, tombstone propagation and reconciliation discrepancies. Roll back model/prompt/index changes independently where possible; a rollback must not restore revoked access or deleted content.
Interview follow-ups
1. Why not place the whole archive in a large context window? It exceeds practical per-request evidence, cost and latency budgets, mixes applicability and permissions, and makes citations difficult. A bounded authorized document packet can be a useful long-context alternative; evaluate that specific workload.
2. The vector write succeeds and keyword indexing fails. What is visible? The old permitted published revision remains available for ordinary updates. The new revision is not published until its required stages are searchable. A delete/revoke event separately denies access regardless of ingestion success.
3. How does a policy change affect a cached answer? Its evidence/version key becomes stale and the current publication/access checks reject reuse. Propagate invalidation, but do not depend on invalidation alone to enforce permissions.
4. Why can a faithful answer be wrong? It may faithfully repeat an obsolete policy, one applying to a different department/date, or an erroneous source. Evaluate applicability and source authority as well as textual support.
5. What would you investigate if users like the answers but cannot complete tasks? Satisfaction may reward fluency. Compare supported correctness, time to the authoritative source, repeated searches and real task outcomes. Review the unsuccessful cohort and source coverage.
6. What changes at ten times the load? Measure the new peak token throughput, index filtering cost and ingestion backlog. Add admission, batching and serving capacity according to bottlenecks; confirm failover capacity. Multiplying replicas without testing the model pool and metadata service is insufficient.
Closing remarks and recall table
| Remember | Explain it concretely |
|---|---|
| Applicable | Correct source, effective date and document version |
| Authorized | Current access before evidence processing and release |
| Supported | Claims trace to inspectable passages |
| Recoverable | Versioned jobs, idempotent writes and reconciliation |
| Measured | Complete-task quality, freshness, latency and cost |
60-second interview answer
I would build a read-only assistant over one owned corpus, starting with authenticated search and a model that answers from selected evidence. Versioned ingestion keeps ordinary updates coherent, while a separate deny path handles deletions and revocations. I would test hybrid retrieval and reranking against the baseline, verify citations and applicability, and measure supported answers, freshness, latency and full operating cost. Expansion follows demonstrated recovery and authorization behavior under load.
Remember: Applicable evidence → Current access → Supported answer → Measured outcome.