Learnastra AI SYSTEM DESIGNAnup Rai

Interview toolkit

Answer Frameworks for AI System Design Interviews

By Anup Rai16 min readReviewed September 2026

An answer framework is a way to organize reasoning so another person can follow it. For system design, use the conventional sequence: clarify requirements, estimate scale, propose a baseline, define contracts, investigate bottlenecks, compare changes, and explain evaluation and operations.

Remember: requirements → baseline → evidence → improvements → closing decision.

The practice mnemonics SPIDER and ETA on this page are optional memory aids used in this guide, not globally standardized interview methods. STAR is an established structure for behavioral answers; “L” makes reflection explicit. A framework should support the conversation rather than dictate every minute.

Choose the structure that fits the question

Question Useful structure What the answer must establish
Design a system Requirements → scale → baseline → deep dive → evaluation/operations The components satisfy a coherent product contract
Explain a concept Definition → mechanism/example → limits You understand what the term means and when it applies
Compare options Constraints → alternatives → evidence → decision → reversal trigger The choice follows from the requirements
Diagnose a failure Impact → containment → hypotheses → discriminating evidence → repair The fix addresses the observed cause
Describe past work Situation → task → action → result → reflection Your actual contribution and what you learned

If a term is unfamiliar, say so. Explain the capability you need and reason from the known contract. Do not guess SDK behavior from a product name or present a hypothetical deployment as experience.

System design framework: requirements first

If you use SPIDER, read it as Scope, Prioritize, Initial design, Deep dive, Evaluation, Reliability. Reliability, access control and evaluation also influence earlier decisions; they are not deferred until the final letter.

1. Scope the outcome

For “design an employee assistant,” clarify:

  1. Who uses it, and what counts as successful completion?
  2. Does it answer questions, create drafts or change records?
  3. Which sources are authoritative, and who may access them?
  4. Which mistakes are tolerable, and which must block a release?
  5. What workload and operating constraints should the design support?

A public handbook search box and a payroll-changing agent need different permissions and recovery. Summarize the scope before drawing, and do not spend the whole session gathering every possible requirement.

2. Write functional and nonfunctional requirements separately

Functional requirements describe required behavior. Nonfunctional requirements describe quality attributes or operating constraints on that behavior. Teams sometimes classify a security rule differently; make the requirement explicit rather than arguing about the heading.

Functional requirement Corresponding measurable constraint
Answer a policy question with evidence Define supported-answer quality and final-answer latency
Ingest document changes Define update lag and behavior when freshness is unknown
Enforce source permissions No unauthorized source text in retrieval, model input or output
Let the user open a citation Resolve the exact edition and recheck current access
Escalate an unsupported question Define ownership, response expectation and visible status

Distinguish a hard constraint, which rules out an option if violated, from a preference, which can be traded against another benefit. A weighted quality/cost score cannot compensate for failing an access requirement.

3. Estimate only what changes a decision

Write the units and assumptions. Calculate request rate, token throughput, storage or concurrency where they constrain the architecture. Average traffic does not determine peak capacity. Inference sizing and cost models provide detailed examples.

4. Draw a complete baseline

Show the actor, input, services, stores and output. Separate data preparation from live requests, and mark authorization and state ownership. A generic “AI layer” box is not enough to explain where evidence comes from or which operation may be retried.

Then trace one request end to end. Choose one or two deep dives from the dominant risks or the interviewer's direction. Add components when a requirement or measured failure justifies them.

5. Close the design with evidence and operations

Term Meaning
Baseline Current solution or simpler alternative used for comparison
Evaluation Measurement of behavior against stated criteria
Release gate Evidence required before a particular deployment decision
Rollout Sequence of users, workloads or capabilities receiving the change
Rollback Tested action that restores a usable earlier deployment or mode

Good offline results do not by themselves authorize broad exposure. Name the on-call owner, dashboard signals, fallback behavior and unresolved assumptions.

Worked interview: employee policy assistant

This is an illustrative practice design, not a report of an actual deployment. A 45-minute session might allocate 5 minutes to requirements, 5 to scale, 8 to the baseline, 15 to deep dives, 7 to evaluation/cost and 5 to closing discussion. Adapt to the interviewer's direction.

Functional requirements

  1. Answer employee policy questions using authorized, current source material.
  2. Attach passage-level citations with source and edition identifiers.
  3. Return an explicit limitation when evidence is missing or contradictory.
  4. Ingest updates and deletions from an owned policy corpus.
  5. Provide an escalation route to the policy owner.

Out of scope: payroll writes, benefit enrollment, legal advice and autonomous HR decisions. This first release answers questions only.

Nonfunctional requirements

  1. No cross-user disclosure through retrieval, model prompts, citations, logs or caches.
  2. Updates visible within 15 minutes under normal operation; stale or unknown freshness must be visible to the serving policy.
  3. Revoked source access must take effect before new disclosure; an unavailable permission check must not allow access.
  4. p95 time to a verified final answer below 6 seconds at the agreed peak workload; first-token latency is measured separately.
  5. At least 95% supported correct answers on the predefined answerable evaluation population, with uncertainty and slice results reported separately. Unanswerable cases and critical access failures have separate checks.
  6. A 99.9% monthly service-availability target for the defined query operation; an error or dependency-unavailable response does not count as a successful answer merely because it uses HTTP 200.

The numerical targets are interview assumptions. Define the final acceptance criteria with the stakeholder rather than suggesting that these are universal defaults. A finite test suite cannot prove the absence of all security failures.

Scale and initial estimates

Assumption or calculation Value Design implication
Employees 50,000 Identity and source permissions already exist
Daily active users 20% 10,000 active users
Questions per active user/day 5 50,000 requests/workday
Eight-hour active window 50,000 / 28,800 Mean ≈ 1.74 requests/second
Provisional burst factor 5 × mean ≈ 8.68 requests/second; test at 10 and above
Mean in-flight request duration 3 seconds at 10 requests/second Little's Law estimates 30 concurrent requests, not a p95 guarantee
Indexed corpus 100,000 documents × 10 chunks 1 million chunks
Dense vector storage 1M × 1,536 dimensions × 4 bytes 6.144 GB decimal, before index, metadata, replicas and source bytes
Input/output assumption 4,000 / 500 tokens per request At 10 requests/second: 2.4M input and 0.3M output tokens/minute

Confirm model quotas, queueing and realistic output lengths in load tests. A concurrency estimate is not a provider quota or capacity commitment.

Start with a baseline

Architecture / visual model
flowchart LR S[Versioned policy sources] --> I[Parse and index] I --> X[Search index with source IDs] U[Employee query] --> A[Authenticate and check access] A --> R[Retrieve permitted passages] X --> R R --> P[Pack evidence] P --> M[Model drafts answer] M --> V[Validate citations and support] V --> O[Answer or limitation]
Read diagram source
flowchart LR
    S[Versioned policy sources] --> I[Parse and index]
    I --> X[Search index with source IDs]
    U[Employee query] --> A[Authenticate and check access]
    A --> R[Retrieve permitted passages]
    X --> R
    R --> P[Pack evidence]
    P --> M[Model drafts answer]
    M --> V[Validate citations and support]
    V --> O[Answer or limitation]

Begin with one corpus, lexical retrieval or a simple retrieval method appropriate to its content, and a model deployment that meets data constraints. Compare to the existing employee search experience. Do not assume a vector database, reranker and agent loop are all necessary.

Find the flaws before adding components

Failure in the baseline How to distinguish it Improvement Cost or limitation
Current policy never reaches the index Compare source and indexed edition Versioned ingestion jobs, replay and freshness status More state, reconciliation and operator work
Correct passage is absent from candidates Measure candidate recall on labeled cases Improve parsing, lexical/dense retrieval or query handling Embedding/index costs and new evaluation
Correct passage is present but ranked too low Inspect candidate list versus packed context Evaluate a reranker on those cases Extra latency and call cost; cannot fix missing sources
Permission changes after indexing Check current authoritative policy at exposure boundaries Revision-aware permission enforcement and dependency invalidation Extra dependency and fail-closed behavior
Citation looks valid but does not support the answer Human or calibrated evidence-support grading Require specific evidence and constrain unsupported claims Validation cost; no automatic proof of truth
Dependency stalls consume all workers Queue age, deadlines and saturation Bounded admission, backpressure and circuit breaking Some requests receive a clear unavailable response

Detailed architecture and contracts

Architecture / visual model
flowchart TD S[Source change or deletion] --> Q[Durable ingestion queue] Q --> W[Versioned parser and index worker] W --> X[Search index and source revision catalog] S --> P[Authoritative permissions and revocations] U[Authenticated employee] --> G[Gateway: identity, admission and deadline] G --> R[Retrieval with current access enforcement] X --> R P --> R R --> K[Evidence packing and optional reranking] K --> M[Approved model deployment] M --> V[Output and evidence checks] P --> V V --> A[Answer with citation dependencies] V --> H[Limitation or policy-owner escalation] G --> T[Restricted traces, metrics and usage ledger] W --> T V --> T
Read diagram source
flowchart TD
    S[Source change or deletion] --> Q[Durable ingestion queue]
    Q --> W[Versioned parser and index worker]
    W --> X[Search index and source revision catalog]
    S --> P[Authoritative permissions and revocations]
    U[Authenticated employee] --> G[Gateway: identity, admission and deadline]
    G --> R[Retrieval with current access enforcement]
    X --> R
    P --> R
    R --> K[Evidence packing and optional reranking]
    K --> M[Approved model deployment]
    M --> V[Output and evidence checks]
    P --> V
    V --> A[Answer with citation dependencies]
    V --> H[Limitation or policy-owner escalation]
    G --> T[Restricted traces, metrics and usage ledger]
    W --> T
    V --> T

Explain these records before selecting a storage product:

Record Essential fields and rules
Query Server-derived employee identity, request ID, question, deadline and locale
Source revision Source ID, edition, content digest, effective time, ingestion status and current permission reference
Retrieved passage Source/edition/span, retrieval score and authorization evidence
Answer Request ID, model/prompt/index versions, status and cited source dependencies
Ingestion job Source ID + revision as logical identity, attempt ID, status and bounded retry state

A request-supplied employee ID does not establish identity. A cached group list does not prove that access is current. If source APIs cannot provide the promised revocation semantics, narrow the requirement or redesign the access boundary explicitly; do not promise instantaneous source truth from stale tags.

For this first release, avoid shared generated-answer caching. If later added, key by the relevant security scope and versions, record source dependencies, and invalidate/recheck on revocation. Do not disclose an old answer before its current authorization is established. See access control.

Recovery and degradation

  1. Ingestion worker crash: reclaim an expired lease and retry the same source revision. Publish only a complete, current revision; a late worker cannot overwrite a newer one.
  2. Source deletion: record a tombstone, remove searchable content and invalidate dependent artifacts. Serving checks the authoritative revision state during propagation.
  3. Model timeout: classify the outcome and remaining deadline. A read-only inference retry still consumes cost; any future tool write would require operation reconciliation before replay.
  4. Permission service outage: return unavailable for protected content. An older cached answer is not a safe shortcut around the failed check.
  5. Insufficient evidence: report the limitation and offer the escalation path. Do not manufacture an answer to improve the completion metric.
  6. Overload: bound the queue, apply per-user limits and shed work explicitly rather than allowing unbounded latency.

See reliability patterns and durable execution.

Evaluate the decision

Use the same held-out tasks for the search baseline and candidate. Include authoritative answers, unanswerable questions, stale editions, conflicting passages, access changes, injection attempts and relevant language slices. Measure retrieval separately from the final answer.

Report supported correct answers over the predefined eligible population, explicit limitations, incorrect answers and unresolved measurement cases. Do not silently exclude timeouts or bad answers from the denominator. Calibrate any LLM judge against appropriate human labels. See capability assessment.

Load-test end-to-end latency, including queueing and validation, and test degraded dependencies. Adding component p95 values does not calculate the system's p95. A pilot should also measure whether employees finish their task and how much human escalation work remains.

Calculate full cost and benefit

Assume 22 workdays/month: 1.1 million questions. The following figures are illustrative budget assumptions, not provider quotes or promised savings. An escalated question takes two minutes at a loaded rate of $60/hour, so each escalation costs $2.

Monthly cost Current search baseline Proposed assistant
Common search/service operations $6,000 $6,000
Model calls at assumed $0.008/question $0 $8,800
Added indexing/evaluation variable costs $0 $1,000
Human escalations 5% × 1.1M × $2 = $110,000 2% × 1.1M × $2 = $44,000
Additional assistant operations $0 $5,000
Implementation amortization $0 $2,000
Total $116,000 $66,800

The proposal saves $49,200/month only if the assumed escalation reduction occurs without unacceptable quality loss. If both systems produce 1 million independently verified accepted outcomes, the cost per 1,000 accepted outcomes is $116 versus $66.80. Count ordinary employee time separately if it differs; it is not included in this worksheet.

The candidate's non-escalation costs are $22,800. Its break-even escalation rate is (116,000 − 22,800) / (1,100,000 × 2) ≈ 4.24%. At 5%, the candidate costs $132,800 and is more expensive than the baseline. Measure the time per escalation and its tail, not just the rate.

Closing remarks

“I would start with a read-only assistant over one owned policy corpus. The main decisions are versioned evidence, current access checks and measurable answer support. Retrieval improvements are experiments against the baseline. I would pilot the system only after permission and recovery checks pass, then expand if final-answer quality, latency and full operating cost meet the agreed criteria. The largest unproven economic assumption is the reduction in human escalations.”

Explain a concept: definition, example, limits

ETA is a local memory aid for Explain simply, Technical details, Applications and tradeoffs. Start with the standard definition rather than an analogy alone.

For a KV cache:

  1. Definition: a KV cache retains previously computed attention keys and values for reuse during compatible autoregressive decoding.
  2. Mechanism: a new token produces its own states; attention can reuse earlier keys and values. Other computations and access to prior cache entries remain necessary.
  3. Decision: cache memory may limit concurrent requests even when the model weights fit in memory.

For uniform full-attention layers, an uncompressed cache estimate is:

bytes = 2 × batch × cached_tokens × layers × KV_heads × head_width × bytes_per_element

For one sequence, 8,192 cached tokens, 32 layers, 8 KV heads, head width 128 and BF16 (2 bytes), that is 1 GiB. With 32 KV heads, it is 4 GiB. Hybrid, sliding-window or compressed architectures require their own formulas; parameter count alone does not specify cache size.

Paging reduces allocation waste; grouped-query attention changes the KV-head count. Neither makes all decoding work free. Provider prompt-cache pricing and eligibility are separate contracts from the mathematical cache estimate.

Defend a tradeoff with comparable evidence

Use options → hard gates → measurements → recommendation → reversal trigger.

Step API versus self-hosting example
Options Approved managed deployment; supported self-hosted deployment
Hard gates Data location, license, feature support and minimum quality
Measurements Task outcomes, peak load, latency and total cost including staff
Recommendation Choose the feasible deployment with the stronger measured fit
Reversal trigger Sustained utilization, a contract change or a demonstrated cost/quality advantage

An interface abstraction can reduce integration work, but cannot remove an embedding migration, tokenization difference or behavior change. Compare actual current deployments in the model-selection guide. Avoid arbitrary star scores and unexplained weights.

Debug with hypotheses and discriminating evidence

First define the user impact and contain it. Then compare failing and known-good cases with versions and timestamps. A recent change is a hypothesis, not automatic proof of cause.

Architecture / visual model
flowchart TD A[Wrong or unsupported answer] --> B{Authoritative source correct?} B -->|No| C[Repair source quality or coverage] B -->|Yes| D{Current source parsed and indexed?} D -->|No| E[Repair ingestion and replay version] D -->|Yes| F{Required evidence reaches context?} F -->|No| G[Inspect permissions, retrieval and packing] F -->|Yes| H[Inspect interpretation, output and grading] C --> I[Verify repair and adjacent regression cases] E --> I G --> I H --> I
Read diagram source
flowchart TD
    A[Wrong or unsupported answer] --> B{Authoritative source correct?}
    B -->|No| C[Repair source quality or coverage]
    B -->|Yes| D{Current source parsed and indexed?}
    D -->|No| E[Repair ingestion and replay version]
    D -->|Yes| F{Required evidence reaches context?}
    F -->|No| G[Inspect permissions, retrieval and packing]
    F -->|Yes| H[Inspect interpretation, output and grading]
    C --> I[Verify repair and adjacent regression cases]
    E --> I
    G --> I
    H --> I
Hypothesis Evidence that distinguishes it Suitable response
Retrieval regressed Needed passage exists but is absent from candidates Inspect parser, filters, index and query behavior
Context packing dropped evidence Passage is retrieved but missing from model input Fix selection/truncation and retest the budget
Generator changed Same evidence and grading, different output quality Compare model/prompt settings on controlled cases
Judge changed Same outputs receive different scores Recalibrate or restore the measurement before judging a release
Timeout population grew More unresolved calls appear in the score denominator Investigate quotas, queueing and dependency latency

Avoid changing the model and retrieval pipeline together when trying to attribute the failure. In an incident, containment can precede complete diagnosis; record what changed so you can still investigate.

Behavioral answers: STAR with reflection

STAR means Situation, Task, Action and Result. It helps organize an actual experience. The National Careers Service also includes learning in its guidance; this guide's STAR-L makes that reflection explicit. STAR guidance.

Part Include Avoid
Situation Relevant context and constraints A long company history
Task Your specific responsibility Taking credit for the whole team
Action Decisions, alternatives, collaboration and implementation “We used AI” without your contribution
Result Supported quantitative or qualitative outcome Invented metrics or overstated causality
Learning What you changed in later practice A generic lesson unrelated to the event

A concrete qualitative result is valid when a trustworthy metric was not available. Say what was measured, what was observed and what remains uncertain. See behavioral preparation.

Handle interruptions and unknowns

Interviewer challenge Useful response
“Why not put everything in a long context?” Compare the same authorized corpus, freshness, quality and cost; direct context may be simpler for a small packet
“Just add a reranker.” First establish that the needed evidence is in the candidates; a reranker cannot recover absent evidence
“We need to launch next week.” Offer a smaller scope with named owners and release evidence; do not silently weaken a hard constraint
“You made an authorization mistake.” State the exposure path, repair it and check related caches/artifacts; preserve unaffected decisions
“Explain a tool you have not used.” State what you know, ask for the relevant contract and reason from it; do not invent production experience

An interruption changes the next explanation. It does not erase the requirements already agreed.

Interview questions with developed answers

1. How would you open an ambiguous system design question?

Identify the user outcome, action boundary, authoritative data and consequential failures. State bounded assumptions, number the key requirements and draw a baseline. Gather enough information to make decisions without consuming the entire session.

2. Can a weighted score compensate for a failed residency requirement?

No. If residency is a hard constraint, eliminate the deployment before comparing preferences. A high quality or low price score does not make it eligible.

3. Why distinguish time to first token from final-answer latency?

Streaming can produce early text while the complete answer remains unavailable or unverified. Measure the boundary the user requirement actually specifies, including queueing and validation.

4. Does adding component p95 latencies produce the end-to-end p95?

No. Percentiles do not generally add. Dependencies, parallelism and correlation matter. Use component measurements for diagnosis and measure the complete request distribution for the SLO.

5. How can an apparently cheaper model increase total cost?

It can require more retries, fallbacks, review or longer handling time. Compare the same workload and accepted outcomes with all incremental operating and implementation costs.

6. What does a concurrency estimate from Little's Law establish?

For a stable system, mean in-flight work equals mean arrival rate times mean time in the system. It does not prove a tail-latency target, peak capacity or behavior during an unstable backlog.

7. When should you add a reranker?

When relevant evidence is in the candidate set but ordering or selection is inadequate, and testing shows enough improvement to justify cost and latency. First fix missing or incorrectly parsed sources.

8. Why is a cache key containing a permission set insufficient on its own?

The permission set may be stale, and source content or policy may change. Establish current authorization and invalidate or revalidate affected dependencies before disclosure.

9. What should happen when the evaluator changes during an experiment?

Version it and distinguish a measurement change from a product change. Re-score comparable stored outputs where permitted, calibrate the new evaluator, and avoid attributing the whole score movement to the model.

10. Where does a formula belong in a concept answer?

After defining the concept and variables, when it explains a mechanism or decision. The KV-cache formula is useful because it connects context length and KV-head count to serving memory.

11. Can you promise all security requirements are proved by a passing suite?

No. Tests provide evidence for the cases and boundaries exercised. Combine implementation controls, threat analysis, adversarial testing and monitoring, and state the limits of the evidence.

12. How do you recover from an incorrect design decision during the interview?

Acknowledge the error and consequence, update the affected decision, and explain how the repair closes the failure path. Do not defend the mistake or restart unrelated parts of the design.

13. What if your behavioral result has no reliable number?

Give concrete qualitative evidence and its limits. Explain your contribution and observed outcome. Fabricating a metric is worse than accurately describing what was known.

14. What adds leadership depth without losing technical detail?

Tie ownership, dependencies, milestones and release decisions to the actual architecture. Name who owns source quality, access, evaluation and incident response, and what artifact or signal each responsibility produces.

15. What should closing remarks contain?

The proposed baseline, the decisions driven by the main constraints, the largest compromise, the evidence required for rollout and the next unresolved question. Do not introduce an entirely new architecture in the final minute.

Final summary and notes

Recall card Practical action
Define before elaborating Start a concept answer with its accepted meaning
Number the requirements Separate behavior, constraints and assumptions
Draw before optimizing Make the complete request and data paths visible
Compare before choosing Use the same workload and hard gates
Diagnose before changing Seek evidence that separates competing causes
Measure the complete outcome Include failures, uncertainty and human work
Close with a decision State the compromise and the next validation

Practice with the question bank, whiteboard exercises and common pitfalls.

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← AI Engineering and System Design Question Bank
NEXT LESSONCommon Pitfalls in AI System Design Interviews →

Explore the diagram