Learnastra AI SYSTEM DESIGNAnup Rai

Interview toolkit

AI Engineering FAQ

By Anup Rai15 min readReviewed September 2026

Use this page for a direct definition and a practical distinction. Follow the linked lesson for diagrams, calculations and interview practice. Product availability changes; the model, pricing and framework chapters carry dated checks. There is no universal best stack independent of workload and constraints.

General

What is an AI engineer?

An AI engineer builds software systems that use AI models to perform useful tasks. Work can include data preparation, model integration, evaluation, serving, security and operations. A role may focus on generative AI, traditional machine learning, or both.

Interview example: a support assistant needs identity, authorized retrieval, answer evaluation and failure handling in addition to a model call. See the role preparation guide.

What is the difference between an AI engineer and an ML engineer?

The titles overlap. “AI engineer” often emphasizes applications and model integration; “ML engineer” often emphasizes training, data pipelines or model deployment. Neither title has a universal boundary, and either role can own the full lifecycle.

Read the responsibilities rather than assuming one role never trains models or the other never builds applications. See role and market research.

How do I become an AI engineer?

Build the missing skills for a target responsibility and demonstrate them in a small, complete project:

  1. Implement the ordinary software and data path.
  2. Add a justified model capability.
  3. Evaluate it against a simpler baseline.
  4. Test permissions, failure recovery and cost.
  5. Explain the results and limitations.

Existing backend, frontend, data or testing experience can transfer, but does not remove the need to learn model behavior and evaluation. Use the transition guide.

What programming language should I learn for AI engineering?

Choose the language required by the work and its supported libraries. Python is useful for model experiments and data workflows; TypeScript is useful for web applications and supported agent SDKs. Serving infrastructure may also involve C++, Rust, Go or Java.

A language choice alone is not a career strategy. Learn to inspect API contracts, test failures and operate the resulting system. Check a framework's current SDK support before choosing it. See framework selection.

Is AI engineering a good career?

That depends on your interests, local opportunities and the responsibilities you want. Investigate actual postings, interview expectations, team stability and the work behind the title. No guide can promise demand, compensation or an offer for an individual.

A useful interview question for an employer is: “Which production outcomes will this role own, and how are they measured?” See job-market research.

RAG

What is RAG?

Retrieval-augmented generation combines retrieval of external information with generation conditioned on that information. A policy assistant might retrieve the applicable policy edition and use its passages to answer with citations.

RAG can improve access to relevant information. It does not guarantee correct retrieval, faithful interpretation or a supported answer. The original research combines parametric and retrieved knowledge; today's applications use several retrieval architectures. RAG paper; RAG fundamentals.

How does RAG work?

Separate the data preparation path from the request path:

Architecture / visual model
flowchart LR S[Source documents and permissions] --> I[Parse, version and index] I --> X[Search index] Q[Authenticated question] --> R[Authorized retrieval] X --> R R --> C[Select evidence within budget] C --> G[Generate answer] G --> V[Validate support and citations] V --> A[Answer or explicit limitation]
Read diagram source
flowchart LR
    S[Source documents and permissions] --> I[Parse, version and index]
    I --> X[Search index]
    Q[Authenticated question] --> R[Authorized retrieval]
    X --> R
    R --> C[Select evidence within budget]
    C --> G[Generate answer]
    G --> V[Validate support and citations]
    V --> A[Answer or explicit limitation]

Dense embeddings are one retrieval method; lexical or structured retrieval can also supply evidence. Preserve document versions and permissions through ingestion, caches and citations. See production RAG.

Is RAG dead because of long context windows?

No. A larger context window permits more input but does not decide which material is relevant, current or authorized. Retrieval can reduce unnecessary input and target evidence. Directly providing a small, stable document set may be simpler than operating an index.

Compare answer quality, omitted evidence, latency and total cost on the same task. Long context and retrieval can be combined. See context engineering.

What is the difference between RAG and fine-tuning?

Technique What changes Useful question
RAG Information provided at inference time Does the answer need current, accessible external evidence?
Fine-tuning Model parameters learned from training examples Does adaptation improve the target behavior on held-out cases?

Fine-tuning can affect knowledge and behavior; it is not merely a formatting tool. RAG does not update weights. Neither guarantees strict format, correctness or lower latency. They can be used together. See fine-tuning strategies.

What is the best vector database?

There is no universal winner. Test the actual embedding dimension, dataset size, filters, update/delete load, recall target, latency and deployment requirements. Include operational ownership and the option of adding vector search to an existing database.

A published latency number without its hardware, recall, filters and concurrency is not a sizing guarantee. See vector databases.

What is contextual retrieval?

In this retrieval technique, a short explanation of a chunk's document context is added before indexing it. For example, a passage saying “the limit is 30 days” needs context about the product and policy to which the limit applies.

The added context can help retrieval and can also introduce incorrect information. Keep the original passage, version the derived context, and measure against a plain-chunk baseline. See contextual retrieval.

Hybrid search combines retrieval signals, commonly lexical matching and dense-vector similarity. An exact error code may benefit from lexical matching, while a paraphrased question may benefit from semantic retrieval.

Fusion methods such as reciprocal rank fusion combine ranked results; they do not make unrelated raw scores directly comparable. Product support and filtering behavior vary. See hybrid search.

What is GraphRAG?

GraphRAG is a family of approaches that use graph structure in retrieval-augmented generation. Nodes and edges may represent entities, relationships, documents or extracted claims. Some implementations also summarize communities for broader questions.

A graph query can reliably follow stored edges while those edges remain wrong, stale or incomplete. Graph construction and permissions add costs. Use a graph when the task needs relationships that simpler retrieval does not adequately provide. See GraphRAG.

What is the best chunk size for RAG?

Choose it by measuring evidence retrieval and answer quality for the content. A table, legal clause and source-code function have different boundaries. Small chunks can lose context; large chunks can dilute relevance and consume the input budget.

Test a few sensible structural strategies, including how headings, tables and neighboring passages are retained. Count chunks and their overlap when estimating indexing cost. See chunking strategies.

Agents

What is an AI agent?

An AI agent is a system that selects actions based on observations to pursue an objective. In an LLM-based agent, the model helps choose actions; application code executes them within defined permissions, budgets and stopping rules.

An agent does not require a specific model brand, framework or vector database. In a predefined workflow, code determines the sequence and branches; an agentic workflow gives the model some control over the path. Agent fundamentals; architectural distinction.

What is the difference between an agent and a chatbot?

“Chatbot” describes a conversational interface. “Agent” describes how actions are selected. A chatbot can answer directly, run a fixed workflow or expose an agent. An agent can operate through a chat interface or a background task.

Ask what the system can read or change, how it selects actions, and which results require verification. See agent fundamentals.

What is MCP (Model Context Protocol)?

MCP is an open protocol for connecting AI applications with tools and context provided by servers. It standardizes integration contracts; it does not itself establish that a tool is trustworthy or that the user is authorized to perform a business action.

Protocol versions and SDK releases are separate. Choose compatible versions, enforce authorization at the service boundary and treat returned content as data. Official introduction; MCP lesson.

What is the best agent framework?

Choose by the behavior required: typed inputs, resumable execution, human approval, tracing, testing, supported runtimes and deployment ownership. A small workflow may need only ordinary application code and an SDK.

Framework popularity does not prove correct recovery or tenant isolation. Test an interrupted run and a duplicate tool request before relying on the abstraction. See framework selection.

What is agentic RAG?

Agentic RAG lets a model direct some retrieval decisions, such as rewriting a query, choosing a source or deciding that more evidence is needed. A fixed retrieve-then-generate pipeline does not make those decisions dynamically.

Bound the search and generation work. More iterations can increase cost without finding better evidence; measure stopping quality and failed searches. Different research methods use different control loops. See agentic RAG.

How do computer-use agents work?

They observe an interface, propose actions such as clicking or typing, execute permitted actions, and inspect the result. Observations may use screenshots, accessibility information or browser structure, depending on the system.

A clicked “Submit” button is not proof of a completed transaction. Reconcile the resulting application state, and separate observation permissions from consequential writes. See computer-use agents.

What is context engineering?

Context engineering is the design of the information supplied to model invocations: instructions, task state, retrieved evidence, conversation, tools and their results. It includes selection, formatting, budgeting and updating that information.

The term overlaps with prompt engineering and does not replace a precise explanation of what changed. A larger prompt is not automatically better context. See context engineering.

What are Agent Skills?

In the Agent Skills format, a skill packages instructions and optional resources that an agent can load for a task. Metadata can support discovery before fuller instructions or resources are loaded.

A skill is not an authorization boundary, and an instruction file does not sandbox scripts. Review its provenance and execution requirements under the application's permission model. See building tools and skills.

Models

What is the best LLM right now?

Select a deployment that passes the task's hard constraints and performs well on representative held-out cases. Evaluate quality, feature compatibility, latency, reliability and complete operating cost. A public benchmark is useful evidence for its stated task, not a universal ranking for your application.

Record the model identifier, endpoint, region and settings. Re-evaluate material changes. See model selection.

How much does Claude / GPT / Gemini / DeepSeek cost?

The answer depends on the exact model, input/output usage, cache categories, service tier, context thresholds, region and separately billed tools. Subscription prices for a chat application are not API rates.

Use the dated pricing tables and worked cost models. Count retries and review work as well as model calls. Do not copy a single price into a business case without its unit and conditions.

What is the difference between Claude Opus and Claude Sonnet?

They are separate model families with different product positioning, pricing and model-specific capabilities. The names alone do not establish a fixed quality percentage, latency ratio or feature contract.

Compare the exact current models on your task. Check endpoint, tool, reasoning and context support independently of the family name. See model taxonomy and selection.

Should I use an open-source model?

First distinguish open weights from a model released under terms that meet an open-source definition. Inspect the actual license, allowed uses and available artifacts.

Self-hosting can provide deployment control, but you then own capacity, security, upgrades and operations. It does not automatically provide lower cost or stronger privacy than every managed alternative. See model taxonomy and serving infrastructure.

What is prompt caching?

Provider prompt caching reuses processing for a compatible repeated input prefix, subject to the provider's rules. Cached input may still be billed. Writes, reads, minimum sizes, expiry and routing differ by model and provider.

If ordinary input costs 2 units, a cache write costs 2.5 and a later read costs 0.2, two uses cost 2.7 rather than 4 units. That example assumes an actual cache hit and excludes unrelated costs. See cache semantics and cache economics.

Evaluation

How do you evaluate an LLM?

Evaluate the behavior required by the application, not just whether the response sounds plausible.

  1. Define success, critical failures and the target population.
  2. Create representative held-out cases and important slices.
  3. Run comparable configurations with controlled versions.
  4. Use suitable deterministic checks, human labels or calibrated judges.
  5. Report uncertainty, missing outcomes, latency and cost.
  6. Validate the complete application under limited production exposure.

For RAG, evidence support is different from relevance and retrieval quality. “Reference-free” grading may still require the supplied context. See LLM evaluation.

What is LLM-as-judge?

An LLM judge grades an output against a rubric, reference or evidence. It is a measurement instrument with possible bias and errors, not independent ground truth.

Validate it against appropriate human judgments, inspect disagreement, and test sensitivity to answer order, style and injected instructions. A more expensive judge is not automatically reliable for every domain. See capability assessment.

What is the best LLM observability tool?

Choose by the required traces, metrics, evaluations, privacy controls, retention, deployment model and export/interoperability support. Confirm that instrumentation captures the application boundaries and versions you need.

A trace can reveal latency or a failed tool call; it does not by itself prove factual correctness. Avoid recording raw sensitive prompts by default. See observability.

What is RAGAS?

Ragas is an evaluation framework with metrics and workflows for AI applications, including RAG. Select metrics according to the question being measured and the required inputs.

A faithfulness score assesses support relative to supplied context; that context may itself be stale or wrong. Version the metric, prompts and judge configuration, and calibrate against human review. See RAG evaluation. Do not treat one framework as a mandatory standard.

How do you detect and handle model drift in production?

Monitor input distributions, task outcomes, slice performance, cost and latency; compare a stable regression set and newly labeled production cases. Changes can come from the model, data, retrieval, prompts, tools or the user population.

A falling score signals a problem but does not identify its cause. Trace versions, reproduce failures, and use a tested rollback or constrained fallback if it meets the same requirements. See evaluation-gated delivery.

Inference

What is vLLM?

vLLM is an open-source model inference and serving engine. Its features include mechanisms for scheduling requests and managing attention memory; supported models, hardware and APIs depend on the release.

Compatibility with an API shape does not guarantee identical behavior to another provider. Test the exact model/runtime/hardware combination, load and failure paths. See serving infrastructure.

What is the difference between vLLM and SGLang?

They are distinct inference/serving projects with overlapping capabilities and different implementations and release support. Neither is universally faster or safer for every model and workload.

Compare quality compatibility, time to first token, inter-token latency, throughput, memory use and operational requirements under the same conditions. Check supported versions and security advisories when deploying. See inference fundamentals.

What is TensorRT-LLM?

TensorRT-LLM provides NVIDIA-focused components for optimizing and serving language-model inference. Model support, precision formats and execution paths depend on the release and GPU architecture.

Evaluate the deployment and build workflow as well as measured performance. There is no universal speed multiplier or fixed setup duration relative to another engine. See serving infrastructure.

How do you optimize LLM inference cost?

Measure cost per accepted outcome first. Then test reductions in unnecessary input, output, retries, model size, redundant calls or idle serving capacity. Caching, batching, routing and quantization help under different conditions.

Change Possible benefit What can make it worse?
Smaller model Lower call cost More failures or human review
Batching Higher utilization or lower API rate Longer wait or unsupported features
Caching Less repeated processing Low hit rate, stale answers or isolation errors
Quantization Lower memory use Quality loss or unsupported kernels

See cost optimization.

What is speculative decoding?

A draft mechanism proposes tokens that the target model verifies, potentially reducing expensive sequential target steps. Correct exact speculative sampling can preserve the target distribution under its algorithmic assumptions.

Speed depends on acceptance rate, draft overhead, hardware and workload. Heuristic variants need separate quality analysis; there is no guaranteed speedup. See speculative decoding.

What is a token budget and how do you enforce it?

A token budget limits input or generated token usage for an operation. A spending budget also needs model rates and other billable work, such as tools and retries.

Count with the correct tokenizer or provider method, reserve estimated spend atomically before admission, cap supported outputs, and reconcile actual usage. Hold uncertain in-flight charges until resolved. A prompt asking the model to “stay under budget” is not enforcement. See FinOps and token economics.

How do you design fallbacks across multiple LLM providers?

Select alternatives that satisfy the same data, feature and quality constraints. Normalize interfaces while preserving provider-specific limits and errors. Classify the failure before retrying.

Architecture / visual model
flowchart TD E[Primary call did not complete normally] --> O{Outcome known?} O -->|No, side effect possible| R[Reconcile operation before replay] O -->|Yes| C{Alternative meets task constraints?} C -->|Yes, retry permitted| F[Bounded fallback with shared deadline] C -->|No| H[Return limitation or review path] R --> S[Record confirmed or unresolved status]
Read diagram source
flowchart TD
    E[Primary call did not complete normally] --> O{Outcome known?}
    O -->|No, side effect possible| R[Reconcile operation before replay]
    O -->|Yes| C{Alternative meets task constraints?}
    C -->|Yes, retry permitted| F[Bounded fallback with shared deadline]
    C -->|No| H[Return limitation or review path]
    R --> S[Record confirmed or unresolved status]

A fallback must not bypass an authorization denial or safety requirement. Independent providers can still share correlated failures or dependencies. See AI gateways.

Memory

What is the best AI agent memory framework?

Start with the memory contract: what may be retained, who owns it, how it is retrieved, how claims are corrected, and when deletion must take effect. Then choose storage and tooling that implement those rules.

Evaluate usefulness, false memories, access isolation and lifecycle behavior. A vector database or published benchmark score does not prove that remembered information is true. See memory architectures.

What is the difference between short-term and long-term memory in agents?

Short-term memory commonly refers to working information for the current task or conversation. Long-term memory retains selected information across tasks or sessions. These are application concepts, not a required mapping to one storage technology.

Working state can live outside the current model context. A KV cache is computation reuse, not durable semantic memory. See short-term context and long-term memory.

How does a knowledge graph help an AI agent?

It represents entities and relationships that can support structured lookup, relationship traversal and provenance. For example, a graph can record which document edition supports a particular claim.

The graph is only as reliable as its data and update process. Exact traversal does not imply correct or complete facts. Apply access rules to nodes, edges and derived answers. See GraphRAG.

Security

What is prompt injection?

Prompt injection occurs when instructions embedded in input influence a model in ways that conflict with the application's intended instruction or trust boundaries. An indirect injection may arrive in a retrieved page or tool result.

For example, a document may tell a support agent to send private account data to an outside URL. The document is task data; it cannot authorize that action. See prompt injection.

How do you prevent prompt injection?

No prompt format or detector guarantees prevention in an arbitrary application. Reduce exposure and consequences through enforceable boundaries:

  1. Keep authorization and allowed actions in application code.
  2. Use least-privilege tools and scoped credentials.
  3. Preserve the origin of untrusted content.
  4. Restrict execution, file access and network destinations as appropriate.
  5. Validate outputs and bind required approval to the exact action.
  6. Test attacks, monitor failures and limit recovery costs.

Delimiters and model-based classifiers can help but do not replace these controls. See LLM security.

What is OWASP LLM Top 10?

It is OWASP's community guidance on major risks in LLM applications, with explanations and mitigation approaches. It is not a certification or proof that a system is secure.

The 2026 edition was published August 3, 2026. Always name the edition when using an identifier because categories and rankings change. Use the official 2026 resource and the security lesson's mapped examples.

What is sandboxing in AI agents?

Sandboxing restricts code or tool execution within a controlled environment. Isolation mechanisms may include operating-system controls, containers or virtual machines; these provide different boundaries.

A disposable workspace can still expose mounted secrets, send data over the network or consume resources. Define filesystem, credential, network, tenant and lifetime restrictions, then test them. A container is not automatically equivalent to a microVM. See agent sandboxing.

Final summary and notes

Remember Avoid assuming
RAG supplies retrieved evidence Retrieved or cited means correct
Fine-tuning changes parameters It guarantees format or fresh knowledge
An agent chooses actions from observations A chat interface implies autonomy
MCP standardizes connections It authorizes every connected tool
A judge provides a measurement Its score is ground truth
Caches reuse work All cached input is free or safe to share
Sandboxes restrict execution Throwaway environments cannot cause harm
Selection depends on the task A leaderboard determines the whole architecture

Interview tip: define the term in one sentence, give one concrete example, and state the limitation that matters to the proposed design. If asked for a product choice, connect it to a measurable requirement rather than repeating a popularity ranking.

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← AI Roles, Hiring Evidence and Interview Preparation — September 2026
NEXT LESSONLLM Internals: How a Model Learns and Answers →

Explore the diagram