An answer framework is a way to organize reasoning so another person can follow it. For system design, use the conventional sequence: clarify requirements, estimate scale, propose a baseline, define contracts, investigate bottlenecks, compare changes, and explain evaluation and operations.
Remember: requirements → baseline → evidence → improvements → closing decision.
The practice mnemonics SPIDER and ETA on this page are optional memory aids used in this guide, not globally standardized interview methods. STAR is an established structure for behavioral answers; “L” makes reflection explicit. A framework should support the conversation rather than dictate every minute.
Choose the structure that fits the question
| Question | Useful structure | What the answer must establish |
|---|---|---|
| Design a system | Requirements → scale → baseline → deep dive → evaluation/operations | The components satisfy a coherent product contract |
| Explain a concept | Definition → mechanism/example → limits | You understand what the term means and when it applies |
| Compare options | Constraints → alternatives → evidence → decision → reversal trigger | The choice follows from the requirements |
| Diagnose a failure | Impact → containment → hypotheses → discriminating evidence → repair | The fix addresses the observed cause |
| Describe past work | Situation → task → action → result → reflection | Your actual contribution and what you learned |
If a term is unfamiliar, say so. Explain the capability you need and reason from the known contract. Do not guess SDK behavior from a product name or present a hypothetical deployment as experience.
System design framework: requirements first
If you use SPIDER, read it as Scope, Prioritize, Initial design, Deep dive, Evaluation, Reliability. Reliability, access control and evaluation also influence earlier decisions; they are not deferred until the final letter.
1. Scope the outcome
For “design an employee assistant,” clarify:
- Who uses it, and what counts as successful completion?
- Does it answer questions, create drafts or change records?
- Which sources are authoritative, and who may access them?
- Which mistakes are tolerable, and which must block a release?
- What workload and operating constraints should the design support?
A public handbook search box and a payroll-changing agent need different permissions and recovery. Summarize the scope before drawing, and do not spend the whole session gathering every possible requirement.
2. Write functional and nonfunctional requirements separately
Functional requirements describe required behavior. Nonfunctional requirements describe quality attributes or operating constraints on that behavior. Teams sometimes classify a security rule differently; make the requirement explicit rather than arguing about the heading.
| Functional requirement | Corresponding measurable constraint |
|---|---|
| Answer a policy question with evidence | Define supported-answer quality and final-answer latency |
| Ingest document changes | Define update lag and behavior when freshness is unknown |
| Enforce source permissions | No unauthorized source text in retrieval, model input or output |
| Let the user open a citation | Resolve the exact edition and recheck current access |
| Escalate an unsupported question | Define ownership, response expectation and visible status |
Distinguish a hard constraint, which rules out an option if violated, from a preference, which can be traded against another benefit. A weighted quality/cost score cannot compensate for failing an access requirement.
3. Estimate only what changes a decision
Write the units and assumptions. Calculate request rate, token throughput, storage or concurrency where they constrain the architecture. Average traffic does not determine peak capacity. Inference sizing and cost models provide detailed examples.
4. Draw a complete baseline
Show the actor, input, services, stores and output. Separate data preparation from live requests, and mark authorization and state ownership. A generic “AI layer” box is not enough to explain where evidence comes from or which operation may be retried.
Then trace one request end to end. Choose one or two deep dives from the dominant risks or the interviewer's direction. Add components when a requirement or measured failure justifies them.
5. Close the design with evidence and operations
| Term | Meaning |
|---|---|
| Baseline | Current solution or simpler alternative used for comparison |
| Evaluation | Measurement of behavior against stated criteria |
| Release gate | Evidence required before a particular deployment decision |
| Rollout | Sequence of users, workloads or capabilities receiving the change |
| Rollback | Tested action that restores a usable earlier deployment or mode |
Good offline results do not by themselves authorize broad exposure. Name the on-call owner, dashboard signals, fallback behavior and unresolved assumptions.
Worked interview: employee policy assistant
This is an illustrative practice design, not a report of an actual deployment. A 45-minute session might allocate 5 minutes to requirements, 5 to scale, 8 to the baseline, 15 to deep dives, 7 to evaluation/cost and 5 to closing discussion. Adapt to the interviewer's direction.
Functional requirements
- Answer employee policy questions using authorized, current source material.
- Attach passage-level citations with source and edition identifiers.
- Return an explicit limitation when evidence is missing or contradictory.
- Ingest updates and deletions from an owned policy corpus.
- Provide an escalation route to the policy owner.
Out of scope: payroll writes, benefit enrollment, legal advice and autonomous HR decisions. This first release answers questions only.
Nonfunctional requirements
- No cross-user disclosure through retrieval, model prompts, citations, logs or caches.
- Updates visible within 15 minutes under normal operation; stale or unknown freshness must be visible to the serving policy.
- Revoked source access must take effect before new disclosure; an unavailable permission check must not allow access.
- p95 time to a verified final answer below 6 seconds at the agreed peak workload; first-token latency is measured separately.
- At least 95% supported correct answers on the predefined answerable evaluation population, with uncertainty and slice results reported separately. Unanswerable cases and critical access failures have separate checks.
- A 99.9% monthly service-availability target for the defined query operation; an error or dependency-unavailable response does not count as a successful answer merely because it uses HTTP 200.
The numerical targets are interview assumptions. Define the final acceptance criteria with the stakeholder rather than suggesting that these are universal defaults. A finite test suite cannot prove the absence of all security failures.
Scale and initial estimates
| Assumption or calculation | Value | Design implication |
|---|---|---|
| Employees | 50,000 | Identity and source permissions already exist |
| Daily active users | 20% | 10,000 active users |
| Questions per active user/day | 5 | 50,000 requests/workday |
| Eight-hour active window | 50,000 / 28,800 | Mean ≈ 1.74 requests/second |
| Provisional burst factor | 5 × mean | ≈ 8.68 requests/second; test at 10 and above |
| Mean in-flight request duration | 3 seconds at 10 requests/second | Little's Law estimates 30 concurrent requests, not a p95 guarantee |
| Indexed corpus | 100,000 documents × 10 chunks | 1 million chunks |
| Dense vector storage | 1M × 1,536 dimensions × 4 bytes | 6.144 GB decimal, before index, metadata, replicas and source bytes |
| Input/output assumption | 4,000 / 500 tokens per request | At 10 requests/second: 2.4M input and 0.3M output tokens/minute |
Confirm model quotas, queueing and realistic output lengths in load tests. A concurrency estimate is not a provider quota or capacity commitment.
Start with a baseline
Read diagram source
flowchart LR
S[Versioned policy sources] --> I[Parse and index]
I --> X[Search index with source IDs]
U[Employee query] --> A[Authenticate and check access]
A --> R[Retrieve permitted passages]
X --> R
R --> P[Pack evidence]
P --> M[Model drafts answer]
M --> V[Validate citations and support]
V --> O[Answer or limitation]
Begin with one corpus, lexical retrieval or a simple retrieval method appropriate to its content, and a model deployment that meets data constraints. Compare to the existing employee search experience. Do not assume a vector database, reranker and agent loop are all necessary.
Find the flaws before adding components
| Failure in the baseline | How to distinguish it | Improvement | Cost or limitation |
|---|---|---|---|
| Current policy never reaches the index | Compare source and indexed edition | Versioned ingestion jobs, replay and freshness status | More state, reconciliation and operator work |
| Correct passage is absent from candidates | Measure candidate recall on labeled cases | Improve parsing, lexical/dense retrieval or query handling | Embedding/index costs and new evaluation |
| Correct passage is present but ranked too low | Inspect candidate list versus packed context | Evaluate a reranker on those cases | Extra latency and call cost; cannot fix missing sources |
| Permission changes after indexing | Check current authoritative policy at exposure boundaries | Revision-aware permission enforcement and dependency invalidation | Extra dependency and fail-closed behavior |
| Citation looks valid but does not support the answer | Human or calibrated evidence-support grading | Require specific evidence and constrain unsupported claims | Validation cost; no automatic proof of truth |
| Dependency stalls consume all workers | Queue age, deadlines and saturation | Bounded admission, backpressure and circuit breaking | Some requests receive a clear unavailable response |
Detailed architecture and contracts
Read diagram source
flowchart TD
S[Source change or deletion] --> Q[Durable ingestion queue]
Q --> W[Versioned parser and index worker]
W --> X[Search index and source revision catalog]
S --> P[Authoritative permissions and revocations]
U[Authenticated employee] --> G[Gateway: identity, admission and deadline]
G --> R[Retrieval with current access enforcement]
X --> R
P --> R
R --> K[Evidence packing and optional reranking]
K --> M[Approved model deployment]
M --> V[Output and evidence checks]
P --> V
V --> A[Answer with citation dependencies]
V --> H[Limitation or policy-owner escalation]
G --> T[Restricted traces, metrics and usage ledger]
W --> T
V --> T
Explain these records before selecting a storage product:
| Record | Essential fields and rules |
|---|---|
| Query | Server-derived employee identity, request ID, question, deadline and locale |
| Source revision | Source ID, edition, content digest, effective time, ingestion status and current permission reference |
| Retrieved passage | Source/edition/span, retrieval score and authorization evidence |
| Answer | Request ID, model/prompt/index versions, status and cited source dependencies |
| Ingestion job | Source ID + revision as logical identity, attempt ID, status and bounded retry state |
A request-supplied employee ID does not establish identity. A cached group list does not prove that access is current. If source APIs cannot provide the promised revocation semantics, narrow the requirement or redesign the access boundary explicitly; do not promise instantaneous source truth from stale tags.
For this first release, avoid shared generated-answer caching. If later added, key by the relevant security scope and versions, record source dependencies, and invalidate/recheck on revocation. Do not disclose an old answer before its current authorization is established. See access control.
Recovery and degradation
- Ingestion worker crash: reclaim an expired lease and retry the same source revision. Publish only a complete, current revision; a late worker cannot overwrite a newer one.
- Source deletion: record a tombstone, remove searchable content and invalidate dependent artifacts. Serving checks the authoritative revision state during propagation.
- Model timeout: classify the outcome and remaining deadline. A read-only inference retry still consumes cost; any future tool write would require operation reconciliation before replay.
- Permission service outage: return unavailable for protected content. An older cached answer is not a safe shortcut around the failed check.
- Insufficient evidence: report the limitation and offer the escalation path. Do not manufacture an answer to improve the completion metric.
- Overload: bound the queue, apply per-user limits and shed work explicitly rather than allowing unbounded latency.
See reliability patterns and durable execution.
Evaluate the decision
Use the same held-out tasks for the search baseline and candidate. Include authoritative answers, unanswerable questions, stale editions, conflicting passages, access changes, injection attempts and relevant language slices. Measure retrieval separately from the final answer.
Report supported correct answers over the predefined eligible population, explicit limitations, incorrect answers and unresolved measurement cases. Do not silently exclude timeouts or bad answers from the denominator. Calibrate any LLM judge against appropriate human labels. See capability assessment.
Load-test end-to-end latency, including queueing and validation, and test degraded dependencies. Adding component p95 values does not calculate the system's p95. A pilot should also measure whether employees finish their task and how much human escalation work remains.
Calculate full cost and benefit
Assume 22 workdays/month: 1.1 million questions. The following figures are illustrative budget assumptions, not provider quotes or promised savings. An escalated question takes two minutes at a loaded rate of $60/hour, so each escalation costs $2.
| Monthly cost | Current search baseline | Proposed assistant |
|---|---|---|
| Common search/service operations | $6,000 | $6,000 |
| Model calls at assumed $0.008/question | $0 | $8,800 |
| Added indexing/evaluation variable costs | $0 | $1,000 |
| Human escalations | 5% × 1.1M × $2 = $110,000 | 2% × 1.1M × $2 = $44,000 |
| Additional assistant operations | $0 | $5,000 |
| Implementation amortization | $0 | $2,000 |
| Total | $116,000 | $66,800 |
The proposal saves $49,200/month only if the assumed escalation reduction occurs without unacceptable quality loss. If both systems produce 1 million independently verified accepted outcomes, the cost per 1,000 accepted outcomes is $116 versus $66.80. Count ordinary employee time separately if it differs; it is not included in this worksheet.
The candidate's non-escalation costs are $22,800. Its break-even escalation rate is (116,000 − 22,800) / (1,100,000 × 2) ≈ 4.24%. At 5%, the candidate costs $132,800 and is more expensive than the baseline. Measure the time per escalation and its tail, not just the rate.
Closing remarks
“I would start with a read-only assistant over one owned policy corpus. The main decisions are versioned evidence, current access checks and measurable answer support. Retrieval improvements are experiments against the baseline. I would pilot the system only after permission and recovery checks pass, then expand if final-answer quality, latency and full operating cost meet the agreed criteria. The largest unproven economic assumption is the reduction in human escalations.”
Explain a concept: definition, example, limits
ETA is a local memory aid for Explain simply, Technical details, Applications and tradeoffs. Start with the standard definition rather than an analogy alone.
For a KV cache:
- Definition: a KV cache retains previously computed attention keys and values for reuse during compatible autoregressive decoding.
- Mechanism: a new token produces its own states; attention can reuse earlier keys and values. Other computations and access to prior cache entries remain necessary.
- Decision: cache memory may limit concurrent requests even when the model weights fit in memory.
For uniform full-attention layers, an uncompressed cache estimate is:
bytes = 2 × batch × cached_tokens × layers × KV_heads × head_width × bytes_per_element
For one sequence, 8,192 cached tokens, 32 layers, 8 KV heads, head width 128 and BF16 (2 bytes), that is 1 GiB. With 32 KV heads, it is 4 GiB. Hybrid, sliding-window or compressed architectures require their own formulas; parameter count alone does not specify cache size.
Paging reduces allocation waste; grouped-query attention changes the KV-head count. Neither makes all decoding work free. Provider prompt-cache pricing and eligibility are separate contracts from the mathematical cache estimate.
Defend a tradeoff with comparable evidence
Use options → hard gates → measurements → recommendation → reversal trigger.
| Step | API versus self-hosting example |
|---|---|
| Options | Approved managed deployment; supported self-hosted deployment |
| Hard gates | Data location, license, feature support and minimum quality |
| Measurements | Task outcomes, peak load, latency and total cost including staff |
| Recommendation | Choose the feasible deployment with the stronger measured fit |
| Reversal trigger | Sustained utilization, a contract change or a demonstrated cost/quality advantage |
An interface abstraction can reduce integration work, but cannot remove an embedding migration, tokenization difference or behavior change. Compare actual current deployments in the model-selection guide. Avoid arbitrary star scores and unexplained weights.
Debug with hypotheses and discriminating evidence
First define the user impact and contain it. Then compare failing and known-good cases with versions and timestamps. A recent change is a hypothesis, not automatic proof of cause.
Read diagram source
flowchart TD
A[Wrong or unsupported answer] --> B{Authoritative source correct?}
B -->|No| C[Repair source quality or coverage]
B -->|Yes| D{Current source parsed and indexed?}
D -->|No| E[Repair ingestion and replay version]
D -->|Yes| F{Required evidence reaches context?}
F -->|No| G[Inspect permissions, retrieval and packing]
F -->|Yes| H[Inspect interpretation, output and grading]
C --> I[Verify repair and adjacent regression cases]
E --> I
G --> I
H --> I
| Hypothesis | Evidence that distinguishes it | Suitable response |
|---|---|---|
| Retrieval regressed | Needed passage exists but is absent from candidates | Inspect parser, filters, index and query behavior |
| Context packing dropped evidence | Passage is retrieved but missing from model input | Fix selection/truncation and retest the budget |
| Generator changed | Same evidence and grading, different output quality | Compare model/prompt settings on controlled cases |
| Judge changed | Same outputs receive different scores | Recalibrate or restore the measurement before judging a release |
| Timeout population grew | More unresolved calls appear in the score denominator | Investigate quotas, queueing and dependency latency |
Avoid changing the model and retrieval pipeline together when trying to attribute the failure. In an incident, containment can precede complete diagnosis; record what changed so you can still investigate.
Behavioral answers: STAR with reflection
STAR means Situation, Task, Action and Result. It helps organize an actual experience. The National Careers Service also includes learning in its guidance; this guide's STAR-L makes that reflection explicit. STAR guidance.
| Part | Include | Avoid |
|---|---|---|
| Situation | Relevant context and constraints | A long company history |
| Task | Your specific responsibility | Taking credit for the whole team |
| Action | Decisions, alternatives, collaboration and implementation | “We used AI” without your contribution |
| Result | Supported quantitative or qualitative outcome | Invented metrics or overstated causality |
| Learning | What you changed in later practice | A generic lesson unrelated to the event |
A concrete qualitative result is valid when a trustworthy metric was not available. Say what was measured, what was observed and what remains uncertain. See behavioral preparation.
Handle interruptions and unknowns
| Interviewer challenge | Useful response |
|---|---|
| “Why not put everything in a long context?” | Compare the same authorized corpus, freshness, quality and cost; direct context may be simpler for a small packet |
| “Just add a reranker.” | First establish that the needed evidence is in the candidates; a reranker cannot recover absent evidence |
| “We need to launch next week.” | Offer a smaller scope with named owners and release evidence; do not silently weaken a hard constraint |
| “You made an authorization mistake.” | State the exposure path, repair it and check related caches/artifacts; preserve unaffected decisions |
| “Explain a tool you have not used.” | State what you know, ask for the relevant contract and reason from it; do not invent production experience |
An interruption changes the next explanation. It does not erase the requirements already agreed.
Interview questions with developed answers
1. How would you open an ambiguous system design question?
Identify the user outcome, action boundary, authoritative data and consequential failures. State bounded assumptions, number the key requirements and draw a baseline. Gather enough information to make decisions without consuming the entire session.
2. Can a weighted score compensate for a failed residency requirement?
No. If residency is a hard constraint, eliminate the deployment before comparing preferences. A high quality or low price score does not make it eligible.
3. Why distinguish time to first token from final-answer latency?
Streaming can produce early text while the complete answer remains unavailable or unverified. Measure the boundary the user requirement actually specifies, including queueing and validation.
4. Does adding component p95 latencies produce the end-to-end p95?
No. Percentiles do not generally add. Dependencies, parallelism and correlation matter. Use component measurements for diagnosis and measure the complete request distribution for the SLO.
5. How can an apparently cheaper model increase total cost?
It can require more retries, fallbacks, review or longer handling time. Compare the same workload and accepted outcomes with all incremental operating and implementation costs.
6. What does a concurrency estimate from Little's Law establish?
For a stable system, mean in-flight work equals mean arrival rate times mean time in the system. It does not prove a tail-latency target, peak capacity or behavior during an unstable backlog.
7. When should you add a reranker?
When relevant evidence is in the candidate set but ordering or selection is inadequate, and testing shows enough improvement to justify cost and latency. First fix missing or incorrectly parsed sources.
8. Why is a cache key containing a permission set insufficient on its own?
The permission set may be stale, and source content or policy may change. Establish current authorization and invalidate or revalidate affected dependencies before disclosure.
9. What should happen when the evaluator changes during an experiment?
Version it and distinguish a measurement change from a product change. Re-score comparable stored outputs where permitted, calibrate the new evaluator, and avoid attributing the whole score movement to the model.
10. Where does a formula belong in a concept answer?
After defining the concept and variables, when it explains a mechanism or decision. The KV-cache formula is useful because it connects context length and KV-head count to serving memory.
11. Can you promise all security requirements are proved by a passing suite?
No. Tests provide evidence for the cases and boundaries exercised. Combine implementation controls, threat analysis, adversarial testing and monitoring, and state the limits of the evidence.
12. How do you recover from an incorrect design decision during the interview?
Acknowledge the error and consequence, update the affected decision, and explain how the repair closes the failure path. Do not defend the mistake or restart unrelated parts of the design.
13. What if your behavioral result has no reliable number?
Give concrete qualitative evidence and its limits. Explain your contribution and observed outcome. Fabricating a metric is worse than accurately describing what was known.
14. What adds leadership depth without losing technical detail?
Tie ownership, dependencies, milestones and release decisions to the actual architecture. Name who owns source quality, access, evaluation and incident response, and what artifact or signal each responsibility produces.
15. What should closing remarks contain?
The proposed baseline, the decisions driven by the main constraints, the largest compromise, the evidence required for rollout and the next unresolved question. Do not introduce an entirely new architecture in the final minute.
Final summary and notes
| Recall card | Practical action |
|---|---|
| Define before elaborating | Start a concept answer with its accepted meaning |
| Number the requirements | Separate behavior, constraints and assumptions |
| Draw before optimizing | Make the complete request and data paths visible |
| Compare before choosing | Use the same workload and hard gates |
| Diagnose before changing | Seek evidence that separates competing causes |
| Measure the complete outcome | Include failures, uncertainty and human work |
| Close with a decision | State the compromise and the next validation |
Practice with the question bank, whiteboard exercises and common pitfalls.