AI System Design Interview Practice
An AI system design interview asks you to turn an ambiguous product request into an implementable, measurable system. Explain what the system must do, how data moves through it, what can fail, and why the…
The next chapter of your engineering career
Go from first principles to decisions you can defend. Learn how AI works, build a complete architecture, and rehearse the questions that make you think.
Build understanding. Then prove it.
Clear definitions, worked examples, diagrams, and the limits of each technique.
02 / DESIGNRequirements, estimates, a baseline, its failures, justified repairs, and a closing decision.
03 / PRACTISEQuiz yourself, rehearse out loud, and revisit the gaps in your reasoning.
One home for AI engineering
149 lessons
An AI system design interview asks you to turn an ambiguous product request into an implementable, measurable system. Explain what the system must do, how data moves through it, what can fail, and why the…
This page has three levels: 40 quick checks for recall, 128 developed answers for understanding, and five worked design scenarios for synthesis. The ten leadership follow-ups apply when the role includes those…
An answer framework is a way to organize reasoning so another person can follow it. For system design, use the conventional sequence: clarify requirements, estimate scale, propose a baseline, define contracts,…
A weak design answer often contains correct terms but omits the reasoning that connects them. This chapter helps you identify the missing requirement, mechanism or evidence, then repair the explanation. It is a…
A system design diagram explains how a defined user outcome is produced under stated constraints. Use these nine authored exercises to practice the full interview: requirements → baseline → flaws → detailed…
Behavioral interviews ask for evidence from past experience. Explain the situation, your responsibility, the action you took and the result. Match the depth to the role; individual contribution, technical…
Source check: September 24, 2026. Use current employer postings to identify the work, then prepare evidence that you can do it. A title, salary headline or model launch does not establish what every employer wants.
An AI engineer builds software systems that use AI models to perform useful tasks. Work can include data preparation, model integration, evaluation, serving, security and operations. A role may focus on…
Three ways to study the same subject: a full, paced explanation; a self-contained rapid revision handbook; and questions with concealed answers.
Text tokenization segments text into discrete units called tokens. For a language model, the tokenizer also maps those units to integer token IDs in its vocabulary. A decoder reconstructs text from IDs according…
Attention computes a weighted aggregation of value vectors, using compatibility scores between a query and the available keys. In standard scaled dot-product attention, the scores are scaled, masked and…
A Transformer is a neural-network architecture based on attention and position-wise feed-forward layers for processing sequences. Attention combines information from allowed positions; feed-forward layers…
An embedding is a numerical representation of an item in a vector space, usually learned so useful relationships are reflected in the representation. For dense text retrieval, the item is text and the output is…
Inference is using a trained model to compute outputs from new inputs. This chapter covers the serving pipeline for a decoder-only autoregressive language model: how a request becomes generated text, how the…
A taxonomy is a systematic classification using stated characteristics. In model selection, those characteristics include the training objective, input/output contract, architecture, available weights, license…
Capability assessment is the systematic evaluation of a model or system against specified tasks and criteria. A capability is what it can do; the assessment measures how reliably it does it under stated…
Pricing is the schedule of charges for a service; cost is the charge incurred by a measured workload. Total cost of ownership (TCO) includes the infrastructure, people and operating work within an explicitly…
Model selection is choosing a model and its operating configuration to meet a defined task, objective and constraints. In a production application, the decision includes the prompt, retrieval, tools, runtime and…
Pretraining is the initial large-scale training of a model on a broad data distribution before adaptation to a particular application. For a causal language model, training usually learns to predict the next…
Fine-tuning continues training a pretrained model on additional data to adapt its parameters to a task or domain. Supervised fine-tuning (SFT) learns from labeled examples, commonly instructions paired with…
Parameter-efficient fine-tuning (PEFT) adapts a model by training a subset of parameters or additional small components. Low-Rank Adaptation (LoRA) freezes a base weight matrix and learns an additive update…
Reinforcement learning from human feedback (RLHF) uses human feedback to guide a model through reinforcement learning. A common language-model recipe trains a reward model from human preferences and then…
Knowledge distillation trains a student model using supervision produced by a teacher model. That supervision can be predictions, probability distributions, generated responses, or intermediate representations.…
Synthetic data is data generated or constructed rather than directly collected as observations of the target process. For language-model training it may include generated instructions, answers, preference pairs,…
Quantization represents numerical values using a restricted set of levels, commonly to reduce the bits used for model weights, activations, or cached attention tensors. It introduces approximation. Whether that…
Reinforcement learning with verifiable rewards (RLVR) updates a model using rewards computed by a checker, such as answer comparison, code tests, or a proof verifier. The model is the policy that produces…
Inference is using a trained model to compute predictions or outputs. In an autoregressive language model, the service processes the prompt and then repeatedly predicts the next token from the available prefix.…
A KV cache stores attention keys and values for already processed tokens so autoregressive generation can reuse them. A prefix cache reuses compatible prompt-prefix computation across requests. An answer cache…
Speculative decoding accelerates autoregressive generation by proposing candidate tokens with a cheaper mechanism and verifying them with the target model. An exact sampling algorithm can preserve the target…
Batching groups requests or token work for more efficient execution. Continuous batching changes the active request set at generation scheduling boundaries as requests finish and new work can be admitted. It…
PagedAttention is an attention and memory-management approach that stores a sequence's KV cache in blocks that need not be physically contiguous. A block table maps logical token blocks to their physical cache…
Model serving is the system that accepts authorized inference requests, schedules model computation, returns results, and operates the workload within defined reliability and cost limits. An inference engine is…
Cost optimization improves resource spending for a defined outcome while retaining the required quality and service constraints. In this Learnastra practice, the unit is a completed, successful task—not merely a…
A diffusion language model generates text through a learned denoising process. In a common masked-discrete formulation, training corrupts token sequences by masking positions and teaches a model to recover…
On-device inference runs the model on the user's device. Edge inference runs near the data source, such as a site gateway or local appliance. Self-hosted inference means the organization operates the serving…
Prompt engineering is the design and evaluation of inputs that guide a language model toward a specified task. The input may contain instructions, examples, source material and an output contract. Prompting…
In-context learning (ICL) is a model's adaptation to a task through information supplied in its context, without updating its parameters. Few-shot prompting supplies a small number of demonstrations of the…
Chain-of-thought (CoT) prompting elicits intermediate reasoning steps before a final answer. The original few-shot approach included worked reasoning demonstrations in the prompt. It improved results on several…
Tree of Thoughts (ToT) is a search framework that explores and evaluates alternative intermediate reasoning states produced by a language model. A state represents partial progress toward a solution. The system…
Context engineering is the design of how information is selected, organized and supplied to a model for each call. That information can include instructions, the current request, conversation history, tool…
Structured generation produces model output that follows a defined machine-readable format or schema. It can use ordinary prompting, constrained decoding, or a provider's structured-output interface. The…
DSPy is a framework for expressing language-model programs and optimizing them against examples and a metric. It separates the program's input/output contracts and control flow from some of the instructions and…
Prompt injection is an attack in which supplied content attempts to override or redirect a language-model application's intended instructions. In a direct attack, the attacker supplies the user-facing input. In…
Retrieval-augmented generation (RAG) combines retrieval of external information with generation conditioned on that information. In a question-answering application, the system finds relevant evidence at request…
Chunking divides source material into units that can be indexed, retrieved and supplied as evidence. The search unit and the unit shown to the answering model do not have to be identical. A small matching…
An embedding model maps an input into a numerical representation learned for tasks such as similarity search. For retrieval, compatible query and document representations are scored to rank candidate evidence.…
A vector database stores vector representations and supports similarity search, usually alongside identifiers, metadata, updates and operational controls. Vector search can also be a capability inside a…
Hybrid search combines multiple retrieval signals, commonly lexical and dense-vector relevance, to produce a ranked candidate set. Its purpose is to cover different ways a query and document can match. It is a…
Reranking reorders a retrieved candidate set using an additional relevance model or scoring method. It spends more work on a bounded set of candidates than is usually practical across the entire corpus. It can…
Graph-based retrieval-augmented generation uses a graph of entities or relationships to select or organize information supplied to a generative model. It can support local relationship queries, multi-hop…
Agentic RAG uses model-driven decisions to choose retrieval actions as part of an answering workflow. The model may select a source, formulate a follow-up query or decide that more evidence is needed.…
Retrieval enhancements change the query, indexed representation or candidate-selection process to address a specific evidence-retrieval failure. Start by naming that failure. A more elaborate pipeline is useful…
Contextual retrieval adds information about a passage's surrounding document to its searchable representation. The aim is to make a chunk understandable when separated from its original location. In the approach…
Late interaction encodes queries and documents separately, then compares their finer-grained representations during scoring. ColBERT is a text-retrieval architecture that uses contextualized token vectors and a…
Multimodal retrieval-augmented generation retrieves evidence involving more than one type of information, such as text, images, audio or video, and uses that evidence to produce an answer. A table's structure…
RAG evaluation measures how well a retrieval-augmented system finds appropriate evidence, uses it accurately, and satisfies the user's task. Evaluate those stages separately so a failure leads to a specific repair.
A production RAG system combines maintained knowledge sources, authorized retrieval and answer generation under explicit quality and operating requirements. Scaling it means sustaining useful, supported answers…
Data engineering for AI builds and operates the pipelines that turn source data into reliable inputs for retrieval, training, prediction and evaluation. It includes ingestion, transformation, validation, lineage…
An AI agent uses a model to select actions toward a goal, observes their results and can adapt its next action. A workflow fixes more of that control flow in application code. Real products often combine both: a…
An AI agent is a system that selects actions based on observations to pursue an objective. In an LLM-based agent, a language model helps choose actions such as searching, calling an API or revising a plan.…
An agent control loop repeatedly chooses an action, executes it, observes the result and decides what to do next. Reasoning can guide those choices, while the application controls execution and stopping rules. A…
Tool use lets a model propose a structured operation that application code validates and executes. The result becomes input for a later model decision or answer. The model does not gain database, filesystem or…
Multi-agent orchestration coordinates two or more agents: it assigns work, controls communication, manages shared state and combines results. Each agent may have its own model, instructions, tools and context.…
Agent memory is information retained from earlier interactions or observations and made available for later tasks. State is the information needed to represent and continue the current computation. They overlap,…
Planning selects actions and their ordering to reach a goal under constraints. Decomposition breaks a task into smaller tasks with explicit dependencies. A plan can be a short checklist, a dependency graph or a…
Error handling detects and classifies failures, then selects an appropriate response. Recovery restores the task to a known valid state or ends it with an accurate account of what remains unresolved. A…
Human-in-the-loop (HITL) systems include human input, judgment or authorization at defined points in an automated process. The person may supply missing information, review a proposal or resolve an exception.…
Agentic security protects data, systems and users when an AI application can choose and execute actions. The familiar goals of confidentiality, integrity and availability still apply. The additional challenge is…
Agent evaluation measures whether the complete system achieves specified goals while respecting constraints and resource limits. The evaluated system includes the model, instructions, tools, permissions, memory…
Durable execution preserves enough execution state and completed results for work to continue after a process failure. Its guarantees depend on the runtime, persistence configuration and the contracts of the…
An agent loop repeatedly assembles context, selects an action, executes permitted work, observes the result and decides whether to continue. The surrounding application code is often called the agent harness.…
An AI application's memory architecture defines what information it retains, how it changes and how the application selects it for later use. Some memory is conversation-specific; some persists across sessions.…
Context management is the application's selection, organization and budgeting of information supplied to a model interaction. It includes instructions, messages, tool schemas, retrieved evidence and tool…
Long-term memory retains information for use beyond the current interaction or session. It can contain explicit preferences, past events, derived assertions and reviewed procedures. Persistence does not make a…
Mem0 is a memory layer that extracts and retrieves information for AI applications, with managed Platform and self-hosted open-source offerings. It can reduce the work needed to build conversation-derived…
Semantic caching reuses a previously computed result when a new request is judged equivalent enough for that result to remain valid. It usually uses embeddings to find candidates, then applies additional…
Application state is the information needed to describe a system's current condition and determine its next valid actions. For an agent task, that includes its goal, stage, artifact versions, completed…
LangChain is a framework ecosystem for composing model integrations, tools and agent behavior. Its current high-level agent entry point is create_agent; LangGraph provides the lower-level orchestration runtime…
LangGraph is an orchestration framework for stateful workflows and agents. It represents work through state, nodes and control flow, with facilities for persistence, interrupts and streaming. A graph can contain…
Observability is the ability to understand a system's behavior from its emitted telemetry. In an AI application, useful telemetry connects the user's task to retrieval, model calls, tool actions, state…
LlamaIndex is a framework for connecting applications to external data, especially for retrieval-augmented generation (RAG). It provides document ingestion, indexing, retrieval, query engines, and agent…
DSPy is a Python framework for composing language-model programs and optimizing parts of those programs against a chosen metric. You specify inputs, outputs, and program structure. An optimizer can search for…
Semantic Kernel is Microsoft's open-source SDK for integrating models, prompts, and callable functions into applications. A kernel connects configured services and functions; plugins group related functions. The…
A multi-agent framework coordinates multiple model-driven components, their tools, and their shared work. It can help organize delegation, state, and execution. It cannot make several agents' answers…
Framework selection is the choice of reusable software that fits an application's execution, data, and operational requirements. It is not a ranking of which library is most advanced. The right choice depends on…
Claude Code is Anthropic's coding agent: a model-driven tool loop that can inspect a repository, edit files, run commands, and use configured integrations. It is available through several interfaces, including…
A coding model predicts useful code or tool calls; a coding agent combines a model with context, tools, an execution loop, and verification. A product adds a user interface, account policies, integrations, and…
Pydantic AI is a Python framework for building model-driven applications with typed dependencies, tools, and outputs. Mastra is a TypeScript framework with agents, tools, workflows, storage integrations, and…
Dependency churn is the ongoing change in libraries, integrations, runtimes and hosted services that an application depends on. It can break imports, alter behavior, or retire a service while your own source…
Optical character recognition (OCR) converts images of text into machine-readable characters. Layout analysis identifies document regions and their relationships: paragraphs, headings, columns, tables, figures,…
LLM infrastructure is the compute, networking, storage, scheduling and operational machinery that serves model-backed application requests. Its job is to meet the application's quality, latency, availability and…
Continuous integration (CI) regularly integrates changes and checks them automatically. Continuous delivery keeps a tested release ready for deployment. Continuous deployment automatically releases changes that…
An AI gateway is an intermediary on the request path between applications and model services. It can centralize authentication, policy, routing, usage accounting and traffic controls. Its data plane handles…
FinOps is an operating discipline that connects technology spending to business value through shared engineering, product and finance accountability. For AI, the unit of work may include model calls, retrieval,…
LLM application security protects the confidentiality, integrity and availability of an application that uses a language model. It includes ordinary web, identity, data and infrastructure security, plus risks…
Authentication verifies an identity. Authorization decides whether a principal may perform an action on a resource. Isolation enforces the boundaries between users, tenants or workloads. Authentication alone…
A guardrail is a check or constraint intended to keep an AI application within defined behavior or operating limits. It may be a deterministic rule, a statistical detector or a workflow control. It does not mean…
An ensemble combines information from multiple predictors or model outputs to produce a decision. In LLM applications, this can mean voting over answers, selecting a candidate using scores, combining drafts or…
Reliability is the ability of a system to perform its required function under stated conditions over time. For an AI application, a fast HTTP response is insufficient if the answer violates the task contract or…
AI governance is the system of accountability, policies, decision rights and oversight used to direct and control AI across its lifecycle. Compliance means meeting the requirements that apply to the organization…
Evaluation is the systematic assessment of a system against specified criteria. An LLM evaluation runs defined tasks, records outputs and observable effects, grades them using a stated method, and summarizes the…
Observability is the ability to understand a system's behavior from the signals it emits. Instrumentation produces telemetry; monitoring checks selected signals against expected conditions; diagnosis uses that…
A benchmark is a defined set of tasks and an evaluation protocol used to compare systems. A leaderboard orders submitted or measured results under a stated scoring method. A result describes a particular model…
A design pattern is a reusable approach to a recurring design problem, with assumptions and consequences. It is not a framework, a mandatory architecture, or a guarantee of quality. In an interview, name the…
An anti-pattern is a recurring approach that appears useful but produces harmful consequences in a particular context. The lesson is not “never use a long prompt” or “always add agents.” Identify the failed…
Interview problem: design a read-only assistant that answers employee questions from internal policies, procedures and research, with inspectable citations and current access control.
Interview problem: design a support assistant for a multi-tenant SaaS product. It must resolve a bounded set of issues, maintain useful conversation state, read authorized account information, and transfer…
Interview problem: build a system that helps analysts prepare company research reports from filings, earnings material and licensed research. Every published report requires an authorized analyst to approve its…
Interview problem: provide fast inline completions, explanations and developer-requested edits inside an IDE, while protecting repository data and preserving the developer's work.
Interview problem: moderate text and media at social-platform scale while limiting harmful exposure, avoiding wrongful restrictions, and providing timely human review and appeals.
Interview problem: answer questions about company filings, licensed news and market commentary with timely evidence, explicit coverage and a short response deadline.
Interview problem: turn a scoped repository task into a complete proposed patch, demonstrate what was checked, and publish it only through the authorized review workflow.
Interview problem: let competing businesses upload private contracts and ask grounded questions on shared infrastructure, while enforcing access, predictable service and accountable data lifecycle operations.
Interview problem: automate routine e-commerce support across 12 languages, integrate with existing help-desk systems, and transfer sensitive or unresolved work to people without duplicate actions or false promises.
Interview problem: convert native and scanned contracts into searchable, structured records whose important fields can be traced to the exact source evidence.
Interview problem: recommend eligible movies quickly, explain the supporting reason, and learn from user outcomes without exposing another person's viewing history.
This is a hypothetical interview scenario. Workload, performance, staffing and costs are planning assumptions, not measured results. Regulatory references describe the US scope of this example and must be…
This is a hypothetical interview scenario. Workload, latency, staffing and costs are planning assumptions, not measured results. Model and interoperability facts have primary-source links. Clinical review and…
This is a hypothetical interview scenario. Traffic, thresholds, latency and costs are planning assumptions, not measured results. Provider-specific behavior has source links; the proposed service has its own…
This is a hypothetical interview scenario. Traffic, latency, storage, staffing and costs are planning assumptions, not measured results. Connector behavior and model pricing have primary-source links.
This is a hypothetical interview scenario. Volumes, timings, error rates and costs are planning assumptions, not measured production results. Model and platform facts have primary-source links.
This is a hypothetical interview scenario. Tenant counts, workloads, latency, training times and costs are planning assumptions, not measured results. Model, framework and research references are linked to their…
Hypothetical interview scenario. Workloads, tolerances, costs and service targets are assumptions for this design. They are not Learnastra operating results or universal standards.
Hypothetical interview scenario. Traffic, quality margins, GPU throughput, staffing and project costs are planning assumptions. Model documentation and quoted API prices have primary-source links.
Hypothetical interview scenario. Workloads, latency targets, staffing and costs are planning assumptions. Connector and protocol details are checked against primary documentation.
A tool-use agent is a system in which a model selects actions, software executes allowed actions, and the resulting observations guide subsequent decisions. A tool might search a database, edit a file, run a…
A tool-use architecture separates model decisions from authorized execution and verified outcomes. The model proposes what to do. The application validates the request, enforces access and limits, executes it,…
OpenClaw is an open-source, self-hostable assistant platform that connects chat channels and other interfaces to agents, models, tools, and persistent state. Its gateway coordinates these capabilities. The…
A computer-use agent is a software system that uses a model to choose actions in a graphical interface, executes permitted actions through a controller, and observes the resulting state. Its observations may…
A tool-use agent combines model-directed decisions with application-controlled operations. The model selects a tool and proposes arguments. Your application decides whether those arguments are valid and…
A useful tool-agent use case has a concrete task, a permitted action surface, a verifiable result, and an acceptable recovery path. “Add an agent to operations” is not a requirement. “Prepare a replenishment…
Agent safety means preventing or limiting harmful outcomes from an agent's decisions and actions. Security protects the system and its data against unauthorized access, manipulation, and disruption. Governance…
A real-time voice agent is a system that receives spoken input, interprets it, performs permitted work, and produces spoken responses during an ongoing conversation. It combines a media pipeline, conversation…
Multimodal generation produces content in one or more media types—such as images, video or audio—conditioned on text or other media. A text-to-image model is one example. A system that generates a video with…
Learn the concept, build a defensible design, and explain the decision clearly. This guide combines technical lessons with interview questions, diagrams, quantitative examples and complete system-design…
A design pattern is a reusable approach to a recurring problem under stated conditions. It gives you a starting structure and known tradeoffs. It does not establish that your application needs the pattern or…
The model can propose an action. Application code decides whether that action is allowed and records its outcome.
Build a small evaluation workflow you can explain in an interview and inspect in code. Start with the evaluation foundations guide for definitions and statistical assumptions. This companion applies those ideas…
An AI evaluation is a systematic assessment of a model or application against defined criteria using specified data, procedures, and measurements. Its purpose is to support a decision: whether a behavior meets a…
A useful interview guide must be accurate enough to build understanding and clear enough to recall under pressure. This page explains how to report a problem and what a publishable correction should establish.
Choose a resource to close a specific skill gap, then demonstrate the skill in a small project and a design explanation. Completing more courses is not the same as becoming ready for an interview.
Start with the work you want to do, then connect your existing experience to evidence that you can do it. An “AI” title can describe product development, model research, infrastructure, evaluation, customer…
Research is useful in an interview when it helps you explain a mechanism, question an assumption or design a better experiment. A paper's result is evidence about its tested setting. It is not automatically a…
Interview problem: build a tutor that explains approved course material, gives hints, checks answers and recommends the next exercise. Optimize learning and independent problem solving rather than conversation…
Interview problem: several product teams use different model providers. Create one governed API that enforces identity, budget, model eligibility, regional processing constraints and reliable usage accounting.
Interview problem: teams need to classify, enrich or summarize millions of records overnight. Build a platform that validates input, schedules model work within quotas, resumes after failure and produces…
Interview problem: let a creative team generate, revise and export images and short videos from text and approved reference assets. Jobs can take seconds to minutes, so the product must expose progress,…
Interview problem: provide an internal inference API for interactive assistants and offline jobs. Teams choose from approved model versions; the platform enforces tenant budgets, predictable latency, isolation…
Interview problem: offer filtered semantic search over customer document passages. Support ingestion, updates, deletions, index upgrades and predictable search latency without leaking another tenant's records.
No lessons match these filters. Try a broader search.

A thoughtful next step
Work through a design, identify a gap, or shape your preparation around your next role.