Learnastra AI SYSTEM DESIGNAnup Rai

Interview toolkit

AI System Design Interview Practice

By Anup Rai7 min readReviewed September 2026

An AI system design interview asks you to turn an ambiguous product request into an implementable, measurable system. Explain what the system must do, how data moves through it, what can fail, and why the proposed tradeoffs fit the requirements. Model selection is one decision within that design.

This practice section is part of Learnastra's The Design Room. Use it to rehearse explanations and decisions, then test whether you can adapt them when the interviewer changes a constraint. The questions are authored practice material; they are not claimed to be confidential or verified questions from particular employers.

Start with one observable baseline

Choose one prompt from the question bank. Work without looking at its answer for 30–45 minutes, depending on the format you are preparing for. Save your diagram and notes. Mark where you guessed, where you could not explain a component, and where you ran out of time.

If you struggled with… Study next Evidence of improvement
Defining tokens, attention or embeddings LLM internals and embeddings Explain the concept and calculate a small example without notes
Turning the prompt into a design Answer frameworks State numbered requirements and draw an end-to-end baseline
Defending a choice Common pitfalls Compare a viable alternative using the same constraints
Finding failures and recovery behavior Whiteboard exercises Trace an outage, duplicate request and permission change
Explaining your own contribution Behavioral preparation Give a factual account of your decisions, evidence and lessons
Understanding an unfamiliar job description Role and market research Map its responsibilities to concrete preparation tasks
Recalling a short definition FAQ and glossary Define the term, distinguish its nearest neighbor, give an example

There is no required number of weeks before beginning practice. When a gap blocks your explanation, study that concept and return to the same prompt.

Use a repeatable practice loop

Architecture / visual model
flowchart LR P[Choose a prompt] --> A[Answer without notes] A --> R[Review diagram and reasoning] R --> G[Study the specific gaps] G --> T[Retry with a changed constraint] T --> P
Read diagram source
flowchart LR
    P[Choose a prompt] --> A[Answer without notes]
    A --> R[Review diagram and reasoning]
    R --> G[Study the specific gaps]
    G --> T[Retry with a changed constraint]
    T --> P

A useful retry changes something material: the tenant model, deadline, data freshness, access policy or consequence of a wrong action. Repeating a memorized diagram without adapting its decisions does not establish understanding.

  1. Read: identify the concept or failure you need to understand.
  2. Recall: define it and work a small example without the page open.
  3. Draw: show actors, services, stores and labeled data flows.
  4. Critique: find an unmet requirement and explain how you would detect it.
  5. Repair: propose a change and account for its new cost and failure modes.
  6. Retry: answer a related prompt after a break and compare your reasoning.

Keep this interview sheet beside the whiteboard

Stage What to write Example for an internal policy assistant
Functional requirements Numbered user-visible capabilities and explicit exclusions Answer questions with current, authorized citations; no policy changes
Nonfunctional requirements Measurable quality, latency, availability and data constraints State the target, workload and how each will be measured
Assumptions and scale Users, peak requests, document sizes, update rates Distinguish observed requirements from illustrative estimates
Baseline Smallest complete design, with request and ingestion paths Versioned ingestion → authorized retrieval → grounded draft
Contracts and state API fields, identities, versions, ownership and statuses A citation names the exact document edition and passage
Deep dive One or two risks that dominate the problem Revocation, stale content, unsupported answers or overload
Failure and repair Trigger → visible effect → detection → recovery Deleted document still cached → invalidate affected results
Tradeoffs Benefit, cost, alternative and decision trigger Reranking helps measured retrieval quality but adds a network call
Evaluation and rollout Holdout cases, slices, operational measures and rollback Test permissions and evidence support before a limited release
Closing remarks Chosen design, main compromise and next validation State what is supported by evidence and what remains assumed

These are conventional system design activities, not a mandatory acronym or a universal hiring rubric. Adapt their order to the interviewer and the problem. If the interviewer asks for an implementation detail, answer it directly and reconnect it to the design.

Match practice to the responsibilities

Job titles are inconsistent across employers. Use the actual role description and recruiter guidance to choose depth.

Responsibility in the role Spend more practice time on Do not neglect
Build AI application features Retrieval, tool contracts, APIs, evaluation and product behavior Data authorization and ordinary service reliability
Own model serving or infrastructure Scheduling, memory, capacity, observability and cost Quality criteria and customer-visible latency
Train or adapt models Data, objectives, experiments, leakage and deployment Baselines, reproducibility and serving constraints
Lead architecture across teams Interfaces, migration, ownership and risk decisions Concrete request flows and implementation feasibility
Lead product or delivery User needs, experiments, metrics and launch dependencies Technical limits and the consequences of a wrong answer

An experienced candidate may be new to one subsystem. Say where your experience ends, state assumptions, and reason from the contract instead of inventing a production story.

Example practice schedules

Treat these as adjustable plans, not promises of readiness.

Time available Suggested allocation Deliverable
Three focused sessions Baseline mock; study the two largest gaps; retry plus behavioral story Two diagrams and a short correction log
One week Alternate concept recall and design practice; finish with a mock under the expected time limit Several independently explained designs and a prioritized gap list
Two weeks First week establishes fundamentals; second varies constraints and repeats weak topics Evidence that improvements carry over to unfamiliar prompts

For a 60-minute practice session, one possible split is 35 minutes answering, 15 minutes reviewing and 10 minutes rewriting the weakest explanation. A real interview may use a different length or include coding; confirm its format.

Review your work with evidence

Use this local practice scale: 0 = missing, 1 = asserted, 2 = explained, 3 = tested against a counterexample. It is a learning aid, not an employer's scoring system or an offer predictor.

Dimension A useful counterexample
Requirements What changes if actions may move money rather than create drafts?
Data flow Where does the identity used for retrieval authorization come from?
Scale What happens at the stated peak, including retries?
Quality Can a fluent but unsupported answer pass your proposed metric?
Recovery What happens after a timeout if the remote write may have succeeded?
Economics Does the cheaper model increase review or fallback work?
Communication Can someone trace a request without your verbal explanation?

Record the specific weakness and your repair, not just a score. For example: “I treated a timeout as a failed payment. On retry I will distinguish unknown outcome, reconcile the provider operation and preserve the original idempotency key.”

Practice questions to ask yourself

1. Is a diagram with many components stronger evidence than a small diagram?

No. Each component needs a responsibility justified by a requirement. A small design with correct state, contracts and recovery may be more defensible than unexplained infrastructure.

2. Can a proposed SLO stand without a workload?

It is incomplete. Specify the operation, measurement boundary, population, time window and expected load. A latency target for a short draft does not automatically cover a long document or queue delay.

3. What should you do when an interviewer rejects your assumption?

Update the requirement, identify which decisions depend on it, and revise those parts. Treat the change as information about the problem.

4. When is it appropriate to mention a framework?

After explaining the behavior you need. Describe how the framework provides it, what the application still owns, and the alternative if that dependency fails or changes.

5. What makes a behavioral answer credible?

A truthful situation, your actual responsibility, specific actions, supported results and honest limits. Do not invent metrics or describe a hypothetical case study as personal experience.

6. What distinguishes recall from understanding?

Understanding survives a changed constraint. You can explain why a decision worked, identify when it would fail, and choose a suitable alternative.

Final summary and notes

  • Define unfamiliar terms before relying on them.
  • Number requirements; put comparisons in tables; label diagram arrows.
  • Begin with a complete baseline and improve it for a demonstrated reason.
  • Separate facts, assumptions and measurements.
  • Include quality, recovery and full operating cost in your decision.
  • Finish with the main compromise and the next test you would run.

For complete worked designs, start with enterprise RAG, customer support automation or evaluation-gated delivery. Use the transition guide to connect existing experience to the preparation plan.

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

NEXT LESSONAI Engineering and System Design Question Bank →

Explore the diagram