Learnastra AI SYSTEM DESIGNAnup Rai

The next chapter of your engineering career

Understand the model.
Design the system.

Go from first principles to decisions you can defend. Learn how AI works, build a complete architecture, and rehearse the questions that make you think.

106concept lessons
26complete system designs
22learning domains
Your role.Your goals. Your study path.

Build understanding. Then prove it.

A library built for the
way you actually prepare.

01 / UNDERSTAND

First principles.

Clear definitions, worked examples, diagrams, and the limits of each technique.

02 / DESIGN

The full interview.

Requirements, estimates, a baseline, its failures, justified repairs, and a closing decision.

03 / PRACTISE

Make your case.

Quiz yourself, rehearse out loud, and revisit the gaps in your reasoning.

One home for AI engineering

Find your next breakthrough.

Choose by job profile ↗

149 lessons

INTERVIEW TOOLKIT
Interview Prep

AI System Design Interview Practice

An AI system design interview asks you to turn an ambiguous product request into an implementable, measurable system. Explain what the system must do, how data moves through it, what can fail, and why the…

7 min read↗
INTERVIEW TOOLKIT
Interview Prep

AI Engineering and System Design Question Bank

This page has three levels: 40 quick checks for recall, 128 developed answers for understanding, and five worked design scenarios for synthesis. The ten leadership follow-ups apply when the role includes those…

2 hr 27 min read↗
INTERVIEW TOOLKIT
Interview Prep

Answer Frameworks for AI System Design Interviews

An answer framework is a way to organize reasoning so another person can follow it. For system design, use the conventional sequence: clarify requirements, estimate scale, propose a baseline, define contracts,…

16 min read↗
INTERVIEW TOOLKIT
Interview Prep

Common Pitfalls in AI System Design Interviews

A weak design answer often contains correct terms but omits the reasoning that connects them. This chapter helps you identify the missing requirement, mechanism or evidence, then repair the explanation. It is a…

13 min read↗
INTERVIEW TOOLKIT
Interview Prep

Whiteboard Exercises for AI System Design

A system design diagram explains how a defined user outcome is produced under stated constraints. Use these nine authored exercises to practice the full interview: requirements → baseline → flaws → detailed…

34 min read↗
INTERVIEW TOOLKIT
Interview Prep

AI Engineering FAQ

An AI engineer builds software systems that use AI models to perform useful tasks. Work can include data preparation, model integration, evaluation, serving, security and operations. A role may focus on…

15 min read↗
CONCEPT
Foundations

Tokenization Deep Dive: How Text Becomes Model Input

Text tokenization segments text into discrete units called tokens. For a language model, the tokenizer also maps those units to integer token IDs in its vocabulary. A decoder reconstructs text from IDs according…

31 min read↗
CONCEPT
Foundations

Attention Mechanisms: How Tokens Share Information

Attention computes a weighted aggregation of value vectors, using compatibility scores between a query and the available keys. In standard scaled dot-product attention, the scores are scaled, masked and…

34 min read↗
CONCEPT
Foundations

Embeddings and Vector Spaces

An embedding is a numerical representation of an item in a vector space, usually learned so useful relationships are reflected in the representation. For dense text retrieval, the item is text and the output is…

50 min read↗
CONCEPT
Foundations

Inference Pipeline

Inference is using a trained model to compute outputs from new inputs. This chapter covers the serving pipeline for a decoder-only autoregressive language model: how a request becomes generated text, how the…

21 min read↗
CONCEPT
Model Landscape

Model Taxonomy

A taxonomy is a systematic classification using stated characteristics. In model selection, those characteristics include the training objective, input/output contract, architecture, available weights, license…

32 min read↗
CONCEPT
Model Landscape

Capability Assessment

Capability assessment is the systematic evaluation of a model or system against specified tasks and criteria. A capability is what it can do; the assessment measures how reliably it does it under stated…

25 min read↗
CONCEPT
Model Landscape

Pricing and Costs

Pricing is the schedule of charges for a service; cost is the charge incurred by a measured workload. Total cost of ownership (TCO) includes the infrastructure, people and operating work within an explicitly…

32 min read↗
CONCEPT
Model Landscape

Model Selection Guide

Model selection is choosing a model and its operating configuration to meet a defined task, objective and constraints. In a production application, the decision includes the prompt, retrieval, tools, runtime and…

33 min read↗
CONCEPT
Training And Adaptation

Pretraining: learning a reusable language model

Pretraining is the initial large-scale training of a model on a broad data distribution before adaptation to a particular application. For a causal language model, training usually learns to predict the next…

5 min read↗
CONCEPT
Training And Adaptation

Fine-tuning: change behavior for a measured reason

Fine-tuning continues training a pretrained model on additional data to adapt its parameters to a task or domain. Supervised fine-tuning (SFT) learns from labeled examples, commonly instructions paired with…

5 min read↗
CONCEPT
Training And Adaptation

LoRA and QLoRA: understand what becomes smaller

Parameter-efficient fine-tuning (PEFT) adapts a model by training a subset of parameters or additional small components. Low-Rank Adaptation (LoRA) freezes a base weight matrix and learns an additive update…

5 min read↗
CONCEPT
Training And Adaptation

Preference learning: RLHF and DPO

Reinforcement learning from human feedback (RLHF) uses human feedback to guide a model through reinforcement learning. A common language-model recipe trains a reward model from human preferences and then…

5 min read↗
CONCEPT
Training And Adaptation

Synthetic data: create examples for a specific learning gap

Synthetic data is data generated or constructed rather than directly collected as observations of the target process. For language-model training it may include generated instructions, answers, preference pairs,…

5 min read↗
CONCEPT
Training And Adaptation

Quantization: budget memory without guessing quality

Quantization represents numerical values using a restricted set of levels, commonly to reduce the bits used for model weights, activations, or cached attention tensors. It introduces approximation. Whether that…

5 min read↗
CONCEPT
Training And Adaptation

RLVR and GRPO: train against a checkable outcome

Reinforcement learning with verifiable rewards (RLVR) updates a model using rewards computed by a checker, such as answer comparison, code tests, or a proof verifier. The model is the policy that produces…

7 min read↗
CONCEPT
Inference Optimization

Inference: follow the request before optimizing it

Inference is using a trained model to compute predictions or outputs. In an autoregressive language model, the service processes the prompt and then repeatedly predicts the next token from the available prefix.…

4 min read↗
CONCEPT
Inference Optimization

Speculative decoding: propose cheaply, verify correctly

Speculative decoding accelerates autoregressive generation by proposing candidate tokens with a cheaper mechanism and verifying them with the target model. An exact sampling algorithm can preserve the target…

4 min read↗
CONCEPT
Inference Optimization

Batching: schedule useful work without hiding the wait

Batching groups requests or token work for more efficient execution. Continuous batching changes the active request set at generation scheduling boundaries as requests finish and new work can be admitted. It…

4 min read↗
CONCEPT
Inference Optimization

PagedAttention: allocate the cache as the sequence grows

PagedAttention is an attention and memory-management approach that stores a sequence's KV cache in blocks that need not be physically contiguous. A block table maps logical token blocks to their physical cache…

4 min read↗
CONCEPT
Inference Optimization

Serving infrastructure: build a service around the model

Model serving is the system that accepts authorized inference requests, schedules model computation, returns results, and operates the workload within defined reliability and cost limits. An inference engine is…

10 min read↗
CONCEPT
Prompting And Context

Prompt Engineering Fundamentals

Prompt engineering is the design and evaluation of inputs that guide a language model toward a specified task. The input may contain instructions, examples, source material and an output contract. Prompting…

5 min read↗
CONCEPT
Prompting And Context

Few-Shot and In-Context Learning

In-context learning (ICL) is a model's adaptation to a task through information supplied in its context, without updating its parameters. Few-shot prompting supplies a small number of demonstrations of the…

5 min read↗
CONCEPT
Prompting And Context

Chain-of-Thought Prompting and Reasoning

Chain-of-thought (CoT) prompting elicits intermediate reasoning steps before a final answer. The original few-shot approach included worked reasoning demonstrations in the prompt. It improved results on several…

5 min read↗
CONCEPT
Prompting And Context

Tree of Thoughts and Deliberate Search

Tree of Thoughts (ToT) is a search framework that explores and evaluates alternative intermediate reasoning states produced by a language model. A state represents partial progress toward a solution. The system…

5 min read↗
CONCEPT
Prompting And Context

Context Engineering

Context engineering is the design of how information is selected, organized and supplied to a model for each call. That information can include instructions, the current request, conversation history, tool…

9 min read↗
CONCEPT
Prompting And Context

Structured Generation

Structured generation produces model output that follows a defined machine-readable format or schema. It can use ordinary prompting, constrained decoding, or a provider's structured-output interface. The…

5 min read↗
CONCEPT
Prompting And Context

Prompt Optimization with DSPy

DSPy is a framework for expressing language-model programs and optimizing them against examples and a metric. It separates the program's input/output contracts and control flow from some of the instructions and…

5 min read↗
CONCEPT
Prompting And Context

Prompt Injection and Defense

Prompt injection is an attack in which supplied content attempts to override or redirect a language-model application's intended instructions. In a direct attack, the attacker supplies the user-facing input. In…

9 min read↗
CONCEPT
Retrieval Systems

RAG Fundamentals

Retrieval-augmented generation (RAG) combines retrieval of external information with generation conditioned on that information. In a question-answering application, the system finds relevant evidence at request…

9 min read↗
CONCEPT
Retrieval Systems

Chunking Strategies

Chunking divides source material into units that can be indexed, retrieved and supplied as evidence. The search unit and the unit shown to the answering model do not have to be identical. A small matching…

5 min read↗
CONCEPT
Retrieval Systems

Embedding Models

An embedding model maps an input into a numerical representation learned for tasks such as similarity search. For retrieval, compatible query and document representations are scored to rank candidate evidence.…

6 min read↗
CONCEPT
Retrieval Systems

Vector Databases and Search Indexes

A vector database stores vector representations and supports similarity search, usually alongside identifiers, metadata, updates and operational controls. Vector search can also be a capability inside a…

9 min read↗
CONCEPT
Retrieval Systems

Hybrid Search

Hybrid search combines multiple retrieval signals, commonly lexical and dense-vector relevance, to produce a ranked candidate set. Its purpose is to cover different ways a query and document can match. It is a…

7 min read↗
CONCEPT
Retrieval Systems

Reranking Strategies

Reranking reorders a retrieved candidate set using an additional relevance model or scoring method. It spends more work on a bounded set of candidates than is usually practical across the entire corpus. It can…

7 min read↗
CONCEPT
Retrieval Systems

GraphRAG

Graph-based retrieval-augmented generation uses a graph of entities or relationships to select or organize information supplied to a generative model. It can support local relationship queries, multi-hop…

7 min read↗
CONCEPT
Retrieval Systems

Agentic RAG

Agentic RAG uses model-driven decisions to choose retrieval actions as part of an answering workflow. The model may select a source, formulate a follow-up query or decide that more evidence is needed.…

5 min read↗
CONCEPT
Retrieval Systems

Advanced Retrieval Patterns

Retrieval enhancements change the query, indexed representation or candidate-selection process to address a specific evidence-retrieval failure. Start by naming that failure. A more elaborate pipeline is useful…

6 min read↗
CONCEPT
Retrieval Systems

Contextual Retrieval

Contextual retrieval adds information about a passage's surrounding document to its searchable representation. The aim is to make a chunk understandable when separated from its original location. In the approach…

8 min read↗
CONCEPT
Retrieval Systems

Late Interaction and ColBERT

Late interaction encodes queries and documents separately, then compares their finer-grained representations during scoring. ColBERT is a text-retrieval architecture that uses contextualized token vectors and a…

7 min read↗
CONCEPT
Retrieval Systems

Multimodal RAG

Multimodal retrieval-augmented generation retrieves evidence involving more than one type of information, such as text, images, audio or video, and uses that evidence to produce an answer. A table's structure…

10 min read↗
CONCEPT
Retrieval Systems

RAG Evaluation Patterns

RAG evaluation measures how well a retrieval-augmented system finds appropriate evidence, uses it accurately, and satisfies the user's task. Evaluate those stages separately so a failure leads to a specific repair.

12 min read↗
CONCEPT
Retrieval Systems

Production RAG at Scale

A production RAG system combines maintained knowledge sources, authorized retrieval and answer generation under explicit quality and operating requirements. Scaling it means sustaining useful, supported answers…

15 min read↗
CONCEPT
Retrieval Systems

Data Engineering for AI

Data engineering for AI builds and operates the pipelines that turn source data into reliable inputs for retrieval, training, prediction and evaluation. It includes ingestion, transformation, validation, lineage…

12 min read↗
CONCEPT
Agentic Systems

Agentic Systems: Design, Control and Verification

An AI agent uses a model to select actions toward a goal, observes their results and can adapt its next action. A workflow fixes more of that control flow in application code. Real products often combine both: a…

4 min read↗
CONCEPT
Agentic Systems

Agent Fundamentals

An AI agent is a system that selects actions based on observations to pursue an objective. In an LLM-based agent, a language model helps choose actions such as searching, calling an API or revising a plan.…

6 min read↗
CONCEPT
Agentic Systems

Reasoning Loops: ReAct and Beyond

An agent control loop repeatedly chooses an action, executes it, observes the result and decides what to do next. Reasoning can guide those choices, while the application controls execution and stopping rules. A…

7 min read↗
CONCEPT
Agentic Systems

Tool Use and MCP

Tool use lets a model propose a structured operation that application code validates and executes. The result becomes input for a later model decision or answer. The model does not gain database, filesystem or…

12 min read↗
CONCEPT
Agentic Systems

Multi-Agent Orchestration

Multi-agent orchestration coordinates two or more agents: it assigns work, controls communication, manages shared state and combines results. Each agent may have its own model, instructions, tools and context.…

8 min read↗
CONCEPT
Agentic Systems

Agent Memory and State

Agent memory is information retained from earlier interactions or observations and made available for later tasks. State is the information needed to represent and continue the current computation. They overlap,…

12 min read↗
CONCEPT
Agentic Systems

Planning and Decomposition

Planning selects actions and their ordering to reach a goal under constraints. Decomposition breaks a task into smaller tasks with explicit dependencies. A plan can be a short checklist, a dependency graph or a…

7 min read↗
CONCEPT
Agentic Systems

Error Handling and Recovery

Error handling detects and classifies failures, then selects an appropriate response. Recovery restores the task to a known valid state or ends it with an accurate account of what remains unresolved. A…

10 min read↗
CONCEPT
Agentic Systems

Human-in-the-Loop Patterns

Human-in-the-loop (HITL) systems include human input, judgment or authorization at defined points in an automated process. The person may supply missing information, review a proposal or resolve an exception.…

10 min read↗
CONCEPT
Agentic Systems

Agentic Security and Sandboxing

Agentic security protects data, systems and users when an AI application can choose and execute actions. The familiar goals of confidentiality, integrity and availability still apply. The additional challenge is…

10 min read↗
CONCEPT
Agentic Systems

Evaluating Agentic Systems

Agent evaluation measures whether the complete system achieves specified goals while respecting constraints and resource limits. The evaluated system includes the model, instructions, tools, permissions, memory…

11 min read↗
CONCEPT
Agentic Systems

Durable Execution for Long-Running Agents

Durable execution preserves enough execution state and completed results for work to continue after a process failure. Its guarantees depend on the runtime, persistence configuration and the contracts of the…

15 min read↗
CONCEPT
Agentic Systems

Loop Engineering

An agent loop repeatedly assembles context, selects an action, executes permitted work, observes the result and decides whether to continue. The surrounding application code is often called the agent harness.…

10 min read↗
CONCEPT
Memory And State

Memory Architectures

An AI application's memory architecture defines what information it retains, how it changes and how the application selects it for later use. Some memory is conversation-specific; some persists across sessions.…

7 min read↗
CONCEPT
Memory And State

Short-Term Context Management

Context management is the application's selection, organization and budgeting of information supplied to a model interaction. It includes instructions, messages, tool schemas, retrieved evidence and tool…

7 min read↗
CONCEPT
Memory And State

Long-Term Memory

Long-term memory retains information for use beyond the current interaction or session. It can contain explicit preferences, past events, derived assertions and reviewed procedures. Persistence does not make a…

7 min read↗
CONCEPT
Memory And State

Agentic Memory with Mem0

Mem0 is a memory layer that extracts and retrieves information for AI applications, with managed Platform and self-hosted open-source offerings. It can reduce the work needed to build conversation-derived…

7 min read↗
CONCEPT
Memory And State

Semantic Caching

Semantic caching reuses a previously computed result when a new request is judged equivalent enough for that result to remain valid. It usually uses embeddings to find candidates, then applies additional…

7 min read↗
CONCEPT
Memory And State

State Management Patterns

Application state is the information needed to describe a system's current condition and determine its next valid actions. For an agent task, that includes its goal, stage, artifact versions, completed…

7 min read↗
CONCEPT
Frameworks And Tools

LangChain Deep Dive

LangChain is a framework ecosystem for composing model integrations, tools and agent behavior. Its current high-level agent entry point is create_agent; LangGraph provides the lower-level orchestration runtime…

7 min read↗
CONCEPT
Frameworks And Tools

LangGraph Orchestration

LangGraph is an orchestration framework for stateful workflows and agents. It represents work through state, nodes and control flow, with facilities for persistence, interrupts and streaming. A graph can contain…

7 min read↗
CONCEPT
Frameworks And Tools

LangSmith Observability

Observability is the ability to understand a system's behavior from its emitted telemetry. In an AI application, useful telemetry connects the user's task to retrieval, model calls, tool actions, state…

8 min read↗
CONCEPT
Frameworks And Tools

LlamaIndex: document retrieval and event-driven workflows

LlamaIndex is a framework for connecting applications to external data, especially for retrieval-augmented generation (RAG). It provides document ingestion, indexing, retrieval, query engines, and agent…

8 min read↗
CONCEPT
Frameworks And Tools

DSPy: programming and evaluating model behavior

DSPy is a Python framework for composing language-model programs and optimizing parts of those programs against a chosen metric. You specify inputs, outputs, and program structure. An optimizer can search for…

7 min read↗
CONCEPT
Frameworks And Tools

Semantic Kernel and Microsoft Agent Framework

Semantic Kernel is Microsoft's open-source SDK for integrating models, prompts, and callable functions into applications. A kernel connects configured services and functions; plugins group related functions. The…

8 min read↗
CONCEPT
Frameworks And Tools

Multi-agent frameworks: CrewAI, AutoGen, and current SDKs

A multi-agent framework coordinates multiple model-driven components, their tools, and their shared work. It can help organize delegation, state, and execution. It cannot make several agents' answers…

8 min read↗
CONCEPT
Frameworks And Tools

Claude Code: designing a dependable coding workflow

Claude Code is Anthropic's coding agent: a model-driven tool loop that can inspect a repository, edit files, run commands, and use configured integrations. It is available through several interfaces, including…

8 min read↗
CONCEPT
Frameworks And Tools

Coding models and agents: choose by evidence

A coding model predicts useful code or tool calls; a coding agent combines a model with context, tools, an execution loop, and verification. A product adds a user interface, account policies, integrations, and…

8 min read↗
CONCEPT
Frameworks And Tools

Framework changes: reproduce, diagnose, migrate

Dependency churn is the ongoing change in libraries, integrations, runtimes and hosted services that an application depends on. It can break imports, alter behavior, or retire a service while your own source…

8 min read↗
CONCEPT
Infrastructure And MLOps

CI/CD for LLM applications: release the complete behavior

Continuous integration (CI) regularly integrates changes and checks them automatically. Continuous delivery keeps a tested release ready for deployment. Continuous deployment automatically releases changes that…

9 min read↗
CONCEPT
Infrastructure And MLOps

FinOps and token economics: measure cost per useful outcome

FinOps is an operating discipline that connects technology spending to business value through shared engineering, product and finance accountability. For AI, the unit of work may include model calls, retrieval,…

8 min read↗
CONCEPT
Security And Access

LLM security: protect data, authority and execution

LLM application security protects the confidentiality, integrity and availability of an application that uses a language model. It includes ordinary web, identity, data and infrastructure security, plus risks…

10 min read↗
CONCEPT
Security And Access

Access control for AI applications

Authentication verifies an identity. Authorization decides whether a principal may perform an action on a resource. Isolation enforces the boundaries between users, tenants or workloads. Authentication alone…

10 min read↗
CONCEPT
Reliability And Safety

Guardrails: enforce a specific rule at the right boundary

A guardrail is a check or constraint intended to keep an AI application within defined behavior or operating limits. It may be a deterministic rule, a statistical detector or a workflow control. It does not mean…

11 min read↗
CONCEPT
Evaluation And Observability

LLM evaluation: measure the behavior required by the product

Evaluation is the systematic assessment of a system against specified criteria. An LLM evaluation runs defined tasks, records outputs and observable effects, grades them using a stated method, and summarizes the…

14 min read↗
CONCEPT
Evaluation And Observability

AI observability: explain what happened to a user task

Observability is the ability to understand a system's behavior from the signals it emits. Instrumentation produces telemetry; monitoring checks selected signals against expected conditions; diagnosis uses that…

12 min read↗
SYSTEM DESIGN
Case Studies

Design an Enterprise Knowledge Assistant

Interview problem: design a read-only assistant that answers employee questions from internal policies, procedures and research, with inspectable citations and current access control.

12 min read↗
SYSTEM DESIGN
Case Studies

Design a Conversational Customer-Support Agent

Interview problem: design a support assistant for a multi-tenant SaaS product. It must resolve a bounded set of issues, maintain useful conversation state, read authorized account information, and transfer…

11 min read↗
SYSTEM DESIGN
Case Studies

Design a Source-Verified Financial Research Assistant

Interview problem: build a system that helps analysts prepare company research reports from filings, earnings material and licensed research. Every published report requires an authorized analyst to approve its…

12 min read↗
SYSTEM DESIGN
Case Studies

Design an IDE Code Assistant

Interview problem: provide fast inline completions, explanations and developer-requested edits inside an IDE, while protecting repository data and preserving the developer's work.

12 min read↗
SYSTEM DESIGN
Case Studies

Design a Content-Moderation Platform

Interview problem: moderate text and media at social-platform scale while limiting harmful exposure, avoiding wrongful restrictions, and providing timely human review and appeals.

13 min read↗
SYSTEM DESIGN
Case Studies

Design a Multi-Tenant Contract-Analysis Platform

Interview problem: let competing businesses upload private contracts and ask grounded questions on shared infrastructure, while enforcing access, predictable service and accountable data lifecycle operations.

13 min read↗
SYSTEM DESIGN
Case Studies

Design Support Automation That Resolves the Right Issue

Interview problem: automate routine e-commerce support across 12 languages, integrate with existing help-desk systems, and transfer sensitive or unresolved work to people without duplicate actions or false promises.

12 min read↗
SYSTEM DESIGN
Case Studies

Case Study: Pharmaceutical Promotion Review

This is a hypothetical interview scenario. Workload, performance, staffing and costs are planning assumptions, not measured results. Regulatory references describe the US scope of this example and must be…

12 min read↗
SYSTEM DESIGN
Case Studies

Case Study: Clinical Voice Documentation

This is a hypothetical interview scenario. Workload, latency, staffing and costs are planning assumptions, not measured results. Model and interoperability facts have primary-source links. Clinical review and…

13 min read↗
SYSTEM DESIGN
Case Studies

Case Study: Real-Time Payment Fraud Decisions

This is a hypothetical interview scenario. Traffic, thresholds, latency and costs are planning assumptions, not measured results. Provider-specific behavior has source links; the proposed service has its own…

13 min read↗
SYSTEM DESIGN
Case Studies

Case Study: Enterprise Knowledge Assistant

This is a hypothetical interview scenario. Traffic, latency, storage, staffing and costs are planning assumptions, not measured results. Connector behavior and model pricing have primary-source links.

14 min read↗
SYSTEM DESIGN
Case Studies

Case Study: Customer-Specific Fine-Tuning Platform

This is a hypothetical interview scenario. Tenant counts, workloads, latency, training times and costs are planning assumptions, not measured results. Model, framework and research references are linked to their…

17 min read↗
SYSTEM DESIGN
Case Studies

Design an Evaluation Gate for AI Releases

Hypothetical interview scenario. Workloads, tolerances, costs and service targets are assumptions for this design. They are not Learnastra operating results or universal standards.

18 min read↗
SYSTEM DESIGN
Case Studies

Design a Customer-Specific Distillation Pipeline

Hypothetical interview scenario. Traffic, quality margins, GPU throughput, staffing and project costs are planning assumptions. Model documentation and quoted API prices have primary-source links.

20 min read↗
SYSTEM DESIGN
Case Studies

Design an Enterprise Knowledge Agent with MCP

Hypothetical interview scenario. Workloads, latency targets, staffing and costs are planning assumptions. Connector and protocol details are checked against primary documentation.

20 min read↗
CONCEPT
Tool Use And Computer Agents

Architecture patterns for dependable tool-use agents

A tool-use architecture separates model decisions from authorized execution and verified outcomes. The model proposes what to do. The application validates the request, enforces access and limits, executes it,…

14 min read↗
CONCEPT
Tool Use And Computer Agents

Building tool-use agents: contracts, execution, and evidence

A tool-use agent combines model-directed decisions with application-controlled operations. The model selects a tool and proposes arguments. Your application decides whether those arguments are valid and…

16 min read↗
CONCEPT
Tool Use And Computer Agents

Tool-agent use cases: choose the workflow, prove the value

A useful tool-agent use case has a concrete task, a permitted action surface, a verifiable result, and an acceptable recovery path. “Add an agent to operations” is not a requirement. “Prepare a replenishment…

20 min read↗
CONCEPT
Tool Use And Computer Agents

Safety and governance for tool-using agents

Agent safety means preventing or limiting harmful outcomes from an agent's decisions and actions. Security protects the system and its data against unauthorized access, manipulation, and disruption. Governance…

24 min read↗
CONCEPT
Multimodal Generation

Multimodal Generation: From Prompt to Publishable Media

Multimodal generation produces content in one or more media types—such as images, video or audio—conditioned on text or other media. A text-to-image model is one example. A system that generates a video with…

27 min read↗
INTERVIEW TOOLKIT
Resources

Learnastra AI Interview Guide

Learn the concept, build a defensible design, and explain the decision clearly. This guide combines technical lessons with interview questions, diagrams, quantitative examples and complete system-design…

5 min read↗
INTERVIEW TOOLKIT
Resources

AI Architecture Pattern Reference

A design pattern is a reusable approach to a recurring problem under stated conditions. It gives you a starting structure and known tradeoffs. It does not establish that your application needs the pattern or…

15 min read↗
INTERVIEW TOOLKIT
Resources

AI Engineering Glossary

The model can propose an action. Application code decides whether that action is allowed and records its outcome.

23 min read↗
INTERVIEW TOOLKIT
Resources

AI Evaluation Lab: Build, Inspect, and Compare

Build a small evaluation workflow you can explain in an interview and inspect in code. Start with the evaluation foundations guide for definitions and statistical assumptions. This companion applies those ideas…

25 min read↗
INTERVIEW TOOLKIT
Resources

AI Evaluation: Evidence, Metrics, and Release Decisions

An AI evaluation is a systematic assessment of a model or application against defined criteria using specified data, procedures, and measurements. Its purpose is to support a decision: whether a behavior meets a…

36 min read↗
INTERVIEW TOOLKIT
Resources

Corrections and Editorial Standards

A useful interview guide must be accurate enough to build understanding and clear enough to recall under pressure. This page explains how to report a problem and what a publishable correction should establish.

4 min read↗
INTERVIEW TOOLKIT
Resources

Learning Resources and Practice Paths

Choose a resource to close a specific skill gap, then demonstrate the skill in a small project and a design explanation. Completing more courses is not the same as becoming ready for an interview.

10 min read↗
INTERVIEW TOOLKIT
Resources

Preparing for an AI Engineering Role

Start with the work you want to do, then connect your existing experience to evidence that you can do it. An “AI” title can describe product development, model research, infrastructure, evaluation, customer…

17 min read↗
INTERVIEW TOOLKIT
Resources

Research Reading for AI System Design

Research is useful in an interview when it helps you explain a mechanism, question an assumption or design a better experiment. A paper's result is evidence about its tested setting. It is not automatically a…

22 min read↗
SYSTEM DESIGN
Platform & Product Designs

Design an Adaptive AI Learning Tutor

Interview problem: build a tutor that explains approved course material, gives hints, checks answers and recommends the next exercise. Optimize learning and independent problem solving rather than conversation…

8 min read↗
SYSTEM DESIGN
Platform & Product Designs

Design an AI Gateway and Model-Routing Service

Interview problem: several product teams use different model providers. Create one governed API that enforces identity, budget, model eligibility, regional processing constraints and reliable usage accounting.

8 min read↗
SYSTEM DESIGN
Platform & Product Designs

Design a Deadline-Aware Batch Inference Platform

Interview problem: teams need to classify, enrich or summarize millions of records overnight. Build a platform that validates input, schedules model work within quotas, resumes after failure and produces…

8 min read↗
SYSTEM DESIGN
Platform & Product Designs

Design an Image and Video Generation Platform

Interview problem: let a creative team generate, revise and export images and short videos from text and approved reference assets. Jobs can take seconds to minutes, so the product must expose progress,…

8 min read↗
SYSTEM DESIGN
Platform & Product Designs

Design a Multi-Tenant Model-Serving Platform

Interview problem: provide an internal inference API for interactive assistants and offline jobs. Teams choose from approved model versions; the platform enforces tenant budgets, predictable latency, isolation…

8 min read↗
SYSTEM DESIGN
Platform & Product Designs

Design a Multi-Tenant Vector Search Service

Interview problem: offer filtered semantic search over customer document passages. Support ingestion, updates, deletions, index upgrades and predictable search latency without leaking another tenant's records.

8 min read↗
Anup Rai

A thoughtful next step

Turn practice into a conversation.

Work through a design, identify a gap, or shape your preparation around your next role.

Explore personal tutoring →