Learnastra AI SYSTEM DESIGNAnup Rai

Interview toolkit

AI Engineering Glossary

By Anup Rai23 min readReviewed September 2026

Use this reference to prepare a clear first sentence, then explain the example and tradeoff in the linked lesson. Definitions describe mechanisms; product names and informal practitioner labels are identified separately. Reviewed September 24, 2026.

Distinctions to remember

Often confused The distinction Interview example
Model / application / agent A model computes outputs; an application adds data and business logic; an agent can select actions within that application. A refund proposal still needs authorization and an idempotent payment operation.
RAG / fine-tuning / memory Retrieval supplies evidence; fine-tuning changes learned parameters; memory retains information across interactions. Retrieve today's policy; fine-tune a stable extraction behavior; store an approved preference.
JSON / schema / meaning Parsing, structural conformance, and business correctness are separate checks. {"amount":-100} can be valid JSON and still violate a payment rule.
Factuality / faithfulness Factuality concerns truth; faithfulness concerns consistency with the supplied source. Repeating an obsolete policy can be faithful and factually wrong today.
KV cache / response cache A KV cache saves attention computation; a response cache returns an earlier application answer. Neither a shared prefix nor a similar question establishes permission to share private data.
Durability / idempotency Durability preserves recorded state; idempotency prevents an operation from taking additional effect when repeated. A crash-safe workflow can still duplicate a charge if the payment boundary lacks deduplication.
pass@k / pass^k At least one success / success on every attempt, under the stated evaluation protocol. For independent attempts with fixed success probability 0.6, at k = 8: 99.934464% / 1.679616%. These are illustrative probabilities, not estimates from an aggregate benchmark score.
CAP consistency / ACID consistency CAP uses linearizability; ACID consistency concerns preservation of application/database invariants. A replicated read order and a nonnegative inventory invariant are different properties.
Architecture / visual model
flowchart LR U[Authenticated request] --> A[Application policy and context] M[Approved memory] --> A R[Authorized retrieval] --> A A --> L[Model inference] L --> O[Answer validation] L --> P[Proposed tool call] P --> G[Authorization and business rules] G --> T[Idempotent operation] T --> A O --> D[User result]
Read diagram source
flowchart LR
  U[Authenticated request] --> A[Application policy and context]
  M[Approved memory] --> A
  R[Authorized retrieval] --> A
  A --> L[Model inference]
  L --> O[Answer validation]
  L --> P[Proposed tool call]
  P --> G[Authorization and business rules]
  G --> T[Idempotent operation]
  T --> A
  O --> D[User result]

The model can propose an action. Application code decides whether that action is allowed and records its outcome.

A

Term Definition Distinction or example
A2A Agent2Agent is a protocol for communication and task coordination between agent applications. It does not provide a shared authorization policy or make remote agents trustworthy.
ABAC — Attribute-Based Access Control Authorization evaluates attributes of the subject, object, requested operation and, where relevant, environment against policy. A document's tenant, classification and the user's clearance can all matter.
ACID Atomicity, consistency, isolation and durability are transaction properties. Consistency means a valid transaction preserves declared invariants. ACID does not specify a single isolation level or guarantee linearizable distributed reads.
Accuracy In classification, the fraction of evaluated examples classified correctly. At 0.1% fraud prevalence, predicting “not fraud” everywhere scores 99.9% accuracy and detects no fraud.
Advisor / executor An orchestration pattern in which an executor consults another model for advice at selected decision points. Measure consultation frequency, added latency and full task cost; savings are not inherent.
Agentic coding Software work in which an agent selects and performs steps such as reading files, proposing edits and running checks. Passing the checks it chose does not establish that every requirement is satisfied.
Agentic system An application in which a model selects actions or workflow steps toward a goal, subject to runtime limits. A fixed classification pipeline need not be an agent. Autonomy is a design choice.
Agent Plugins A named packaging specification for distributing agent skills and MCP server configurations with a root plugin.json. Version 1.0 standardizes those components; vendor-specific extensions require compatible clients. Specification.
AI control A research approach that evaluates deployment protocols under the assumption that some models may behave adversarially or be misaligned. A monitor's success under one threat model is not a general safety guarantee.
AI gateway An intermediary that applies common policies to requests sent to model services, such as routing, authentication, quotas and telemetry. Request handling is a data-plane function; policy management is a control-plane function. One API shape is optional.
Attention mechanism A computation that combines value vectors using weights derived from query–key compatibility. Scaled dot-product attention commonly normalizes masked scores with softmax. Self-attention derives queries, keys and values from the same sequence; causal masking excludes future positions.
Authentication Establishing that a claimant controls credentials associated with an identity. An authenticated employee may still lack permission to read payroll.
Authorization Deciding whether a principal may perform an operation on a resource under the applicable policy. Check at retrieval and tool execution, not only at login.
Availability In CAP, every request received by a non-failing node eventually receives the response required by the operation, rather than an error caused by sacrificing availability. Operational availability is separately measured using a stated success and time criterion. CAP's formal property is not a monthly uptime percentage or a bounded-latency SLO.

B

Term Definition Distinction or example
Batching Processing several requests or items together to share computation or scheduling overhead. Continuous batching changes active requests during generation. Offline batch APIs and GPU scheduling batches are different contracts.
Benchmark saturation A benchmark has little remaining ability to distinguish the candidates of interest near its measurement ceiling. Check uncertainty, task difficulty and current candidates before calling a benchmark saturated.
BM25 A lexical ranking function combining query-term matches, inverse document frequency, term-frequency saturation and document-length normalization. Strong exact identifiers can be missed by an embedding-only search.
Budget tokens A provider-specific token allocation or limit for a model's reasoning or generation. It is not a universal API field. Reasoning effort may be a qualitative setting rather than a hard token or monetary cap.
Bulkhead Resource isolation that limits how one overloaded or failing workload affects others. Separate interactive and offline concurrency pools; a shared bottleneck can still couple them.

C

Term Definition Distinction or example
C2PA / Content Credentials Standards and associated credentials for cryptographically verifiable media provenance assertions. Valid provenance does not prove that a scene is true. Metadata can be removed; recovery mechanisms have limits.
Calibration Agreement between predicted probabilities and observed frequencies across appropriate groups of predictions. Among predictions scored 0.8, about 80% should be correct under good calibration. Fluent confidence is not calibration.
CAP theorem In an asynchronous distributed system that may partition, it is impossible to guarantee both linearizability and availability for every execution. During a partition, refusing an unsafe read preserves consistency at the expense of availability.
Capability index / composite benchmark An aggregate of several evaluation scores using a defined normalization and weighting method. A ranking can change when task weights change; the index is not a universal capability measure.
Chain-of-thought — CoT Intermediate reasoning steps produced before a final answer, either through prompting or a model's trained behavior. An explanation can be incorrect or fail to faithfully describe how an answer was produced.
Chunking Dividing content into units for indexing, retrieval or downstream processing while retaining useful context and provenance. Embedding chunks and the larger passages supplied to generation need not have equal size.
Circuit breaker A mechanism that temporarily rejects calls to a failing dependency and later probes recovery. It reduces repeated pressure; it does not recover a lost transaction outcome.
Claude Code Anthropic's coding-agent product, with documented development surfaces and tool permissions. CLAUDE.md provides project instructions; it is not an operating-system security boundary.
Claude Fable 5 An earlier model in Anthropic's Fable family, retained here to identify references in older material. Use the current model catalog and version-specific contract; this name does not mean “current best model.”
Cline An open-source coding-agent project with editor, terminal and integration surfaces. Product access, model credentials and execution permissions are separate choices.
Computer use Agent interaction with graphical applications through supported observations and actions, such as screenshots, accessibility information, clicks and typing. Interface actions are less direct evidence of business completion than an authoritative operation status.
Consistency In distributed data systems, a consistency model specifies which histories of reads and writes are permitted. State the model: linearizable, sequential, causal and eventual consistency have different guarantees. CAP uses linearizability.
Context7 A documentation-retrieval service and MCP integration for looking up library documentation. Retrieved material still needs the correct version and relevance; installing it does not ensure it is used.
Context engineering Designing how instructions, evidence, conversation and tool results are selected and arranged for model calls. Include permissions, freshness and token budgets, not just prompt wording.
Context window The maximum token context supported under a model's request contract, often including input and generated output. Separate maximum input, output and total limits. Usable quality at long lengths must be measured.
Context rot An informal label for quality degradation associated with accumulated, distracting, stale or poorly organized context. It is not a universal failure threshold at a fixed token count.
Cosine similarity For nonzero vectors, their dot product divided by the product of their lengths. Undefined for a zero vector. High similarity is not proof of entailment or authorization.
Cursor A development environment with AI-assisted editing and coding-agent capabilities. Evaluate the actual execution surface and plan, rather than assuming every feature is local.

D

Term Definition Distinction or example
Data contamination Overlap or leakage between evaluation material and information used to train, tune or select a system, compromising the intended assessment. A private test set can also leak through repeated optimization against its scores.
Data drift A change in the distribution of input data relative to a reference period or dataset. Changed inputs do not necessarily imply worse performance; measure the resulting outcomes.
Diffusion language model A language model that generates through iterative denoising or refinement of a sequence representation. Masked, continuous and blockwise approaches differ; fully parallel generation and faster service are not universal.
DPO — Direct Preference Optimization A preference-optimization objective that directly trains a policy from preferred and rejected responses, without a separately fitted reward model in the original formulation. DPO is not the same optimization procedure as PPO-based RLHF.
DSPy A framework for composing language-model programs and optimizing supported components against examples and a metric. Optimization can overfit an unreliable metric or a reused development set.
Durable execution Execution supported by persisted progress so a workflow can recover across process failures and long waits under the runtime's guarantees. Replay and checkpoint designs differ. External effects still need idempotency or reconciliation.

E

Term Definition Distinction or example
Effective context length A task- and test-dependent length over which a model meets a stated quality criterion. It depends on evidence position, distractors and task difficulty, not just the advertised window.
Embedding A mapping of items into a vector representation used to encode useful relationships. Items can include text, images, audio and other data. Dense neural embeddings are common; dimensions and distance metrics belong to the model contract.
Endpointing / turn detection Estimating when a speaker's turn has ended so a conversational system can respond. Silence, semantic completion and explicit signals have different false-cutoff and delay tradeoffs.
Ensemble Combining predictions or outputs from multiple models or runs using a specified aggregation method. Correlated errors can defeat voting; additional calls also add cost and failure points.
Eval awareness A model recognizing or behaving differently in an evaluation setting. A behavioral difference requires evidence; benchmark success alone does not demonstrate deployment behavior.
Eventual consistency A model in which replicas eventually converge if updates stop, subject to the system's delivery and conflict-resolution assumptions. It does not by itself promise a maximum staleness interval or read-your-writes.
Extended thinking Anthropic's terminology for model reasoning before an answer; configuration and exposed reasoning information depend on the model. Current adaptive modes cannot all be configured through the older budget_tokens field.
EU AI Act EU Regulation 2024/1689 establishes obligations for specified AI systems and general-purpose AI models, with scope- and role-dependent requirements and application dates. “All obligations started in 2026” is incorrect. Consult the dated governance chapter and official legal text.

F

Term Definition Distinction or example
F1 score The harmonic mean of precision and recall: 2PR / (P + R) when defined. Specify averaging and zero-denominator conventions. F1 does not encode every business cost.
Faithfulness In grounded generation, whether an answer's claims are supported by the supplied evidence. Other evaluation tasks may use a different operational definition. Report unsupported and contradictory claims separately when useful. A citation alone is not support.
Few-shot prompting Supplying a small set of demonstrations in the input to guide a task. No parameter update is required; examples still consume context and can bias results.
Fine-tuning Further training an existing model by updating its parameters or added trainable parameters using a chosen objective and dataset. It can target behavior or a domain; it is not a reliable live database update mechanism.
FinOps for AI Collaborative financial accountability and value-based management of AI usage and costs. Include infrastructure, human review and operations alongside token charges.
Framework churn An informal description of frequent framework, API or dependency changes that create maintenance work. Version locks improve reproducibility; upgrades still need compatibility and behavior checks.
Function calling A model interface for returning a proposed function name and structured arguments. The application validates and executes the call; a proposed invocation is not a completed action.

G

Term Definition Distinction or example
GGUF A binary format for model tensors and metadata used in the llama.cpp ecosystem. It can contain different tensor precisions; a filename alone does not identify quality or runtime compatibility.
Graph engineering A practitioner label for designing workflow nodes, transitions and shared state, including static or dynamically routed graphs. It is not a universally standardized discipline or evidence that every application needs multiple agents.
Grounding Connecting generated statements or decisions to relevant external evidence or an environment. Evidence must be current, authorized and actually support the statement.
Grok 4.3 An earlier xAI Grok model version referenced in older comparisons. Consult current model-selection and pricing chapters rather than treating an old version as the current flagship.
GRPO — Group Relative Policy Optimization An RL method that estimates relative advantages from rewards within groups of sampled outputs, avoiding a separate value model in its original formulation. Reward design, group diversity and implementation affect both cost and training behavior.
Guardrails Controls that constrain or check an AI application's inputs, outputs, actions and execution. Deterministic authorization, schemas, limits and model-based checks have different guarantees.

H

Term Definition Distinction or example
Hallucination Generated content that is nonsensical or unfaithful to its source; in factual question answering, the term also covers fabricated or false claims. Define the task's criterion. Schema errors, execution failures and unsupported claims need different fixes. Research definitions.
Harness / scaffold variance Changes in measured performance caused by prompts, tools, budgets, execution environment or other evaluation scaffolding. Report the configuration and uncertainty; there is no universal score swing.
Harness engineering A practitioner term for building the runtime around a model: context, tools, state, verification, limits and telemetry. Runtime controls must function even when model output is invalid.
HNSW — Hierarchical Navigable Small World A graph-based approximate nearest-neighbor search method using a hierarchy of proximity graphs. Tuning changes memory, index construction, latency and recall; approximate retrieval can miss neighbors.
Human-in-the-loop — HITL Human participation in specified review, approval, correction or escalation steps. A review queue needs staffing and expiry rules; a checkbox does not guarantee informed approval.
Hybrid search Combining different retrieval signals, commonly lexical and vector retrieval, into a candidate ranking. Raw scores from unlike retrievers are not automatically comparable.

I

Term Definition Distinction or example
Idempotency Repeating an operation has the same intended effect as applying it once. A payment key needs correct scope, payload conflict checks and retention; retries alone do not provide this property.
In-context learning A model exhibiting task adaptation from information or demonstrations in its input without updating model parameters. It changes behavior during that context, not the stored weights.
Indirect prompt injection Instructions embedded in external content attempt to redirect an application's model away from the authorized task. Treat retrieved documents and tool results as untrusted data; enforce permissions outside the model.
Inference Computing outputs with a trained model for supplied inputs. Serving also includes queueing, preprocessing, networking and application checks.

J

Term Definition Distinction or example
JSON mode A provider output mode aimed at producing syntactically valid JSON under its supported conditions. It does not necessarily enforce a schema. Handle refusals, truncation and transport failures separately.

K

Term Definition Distinction or example
KV cache Stored attention key and value tensors reused during decoding to avoid recomputing prior token states. It is not a key–value application database. Hybrid architectures may also maintain other state.

L

Term Definition Distinction or example
LangChain An open-source framework and integrations for constructing model-powered applications and agents. Its related hosted products and orchestration packages have separate roles and contracts.
Leaderboard Illusion The title of a research critique of incentives and evaluation practices in model preference leaderboards. It is not a formal metric. Separate preference, factual correctness and uncertainty; Arena uses Bradley–Terry modeling.
Linearizability Each completed operation appears to take effect at one point between its invocation and response, in a legal sequential history that respects real-time order. A read started after a write completes must reflect that write or a later one; overlapping operations allow more than one ordering.
LlamaIndex A framework for connecting data to model applications through ingestion, indexing, retrieval and workflow components. Choosing it does not determine storage consistency, authorization or answer quality.
LLM-as-judge Using a language model to assess an output or behavior against a defined criterion or comparison. Validate against expert labels; judge errors, position bias and missing outcomes need measurement.
LiveCodeBench A coding benchmark drawing time-stamped problems from programming contests, with tasks and splits defined by its release. Contest performance does not directly measure repository maintenance or production reliability.
LoRA — Low-Rank Adaptation Parameter-efficient adaptation that represents trainable weight updates using low-rank matrix factors while freezing the base weights. Training fewer parameters reduces some costs; serving and activation memory still matter.
Loop engineering A practitioner label for designing an agent's repeated observation, action, verification and stopping behavior. Define progress and external bounds; extra iterations need evidence of benefit.
Loopmaxxing Informal shorthand for increasing iterations as if repetition alone ensured success. A larger budget cannot repair missing authority, contradictory requirements or an invalid success test.

M

Term Definition Distinction or example
Managed agents Provider-hosted services that operate some of an agent application's execution infrastructure. Check exactly which persistence, identity, sandbox and recovery guarantees are included.
MCP — Model Context Protocol An open protocol for connecting AI applications with tools, resources and prompts through client–server interactions. Use the dated protocol revision, currently 2026-07-28; an SDK's major version is separate.
Memory poisoning Malicious or false information is inserted into retained agent memory so it can influence later work. Validate provenance and write authority; also enforce permissions, expiry and deletion when reading.
Mixture of agents — MoA Combining contributions from multiple model agents, often through an aggregation or synthesis stage. It is an application-level ensemble, distinct from a model's mixture-of-experts architecture.
Mixture of experts — MoE A model architecture with multiple expert subnetworks and a routing mechanism, often activating a subset for each token. Active parameter count and total weight memory are different quantities.
Model routing Selecting a model or service for a request according to policy or measured characteristics. A fallback must preserve required features and data restrictions; cheap-first is only one design.
Multi-tenancy Serving multiple customer or organizational tenants using some shared application or infrastructure resources. Isolation must cover data, cache, queues, credentials, logs, budgets and execution.

O

Term Definition Distinction or example
o3 An earlier OpenAI reasoning model, retained as historical terminology. o3-mini is a separate model name, not a configuration flag on o3. Check current availability before using either.
OCR — Optical Character Recognition Recognizing textual characters in images and producing machine-readable text. Layout, tables, reading order and field extraction are related but distinct tasks.
OpenHands An open-source software-agent platform and SDK for code-related work. Runtime, sandbox and credential configuration determine its actual execution boundary.

P

Term Definition Distinction or example
pass@k Success on at least one of k attempts under the benchmark's sampling and scoring procedure. An oracle selecting a successful attempt may be unavailable to the deployed application.
pass^k Success on all k repeated attempts under the stated procedure, used to examine consistency of completion. Repeated tasks can have correlated failures. Do not compute it by exponentiating an overall measured success rate. τ-bench.
Precision The fraction of predicted positives that are truly positive: TP / (TP + FP). Ask how many moderation flags deserve action, and define zero-positive handling.
Prefill Processing the input tokens to establish model state before output-token decoding. Time to first token also includes queueing, routing and network delays.
Prefix caching Reusing compatible computed state for an identical leading token sequence across requests. Model, tokenization, positions, configuration and isolation scope must match.
Prompt caching A provider or runtime feature that reuses prompt computation, commonly for repeated prefixes. Eligibility, expiration, storage/write fees and read discounts vary. No fixed saving applies everywhere.
Prompt injection An attack that uses instructions in input or external content to subvert an application's intended model behavior. A detector can reduce risk; authorization and action limits must hold if detection fails.

Q

Term Definition Distinction or example
QLoRA A method for training low-rank adapters through a frozen, quantized base model; the original method uses 4-bit NormalFloat, double quantization and paged optimizers. Trainable adapters and computation need not be 4-bit.
Quantization Representing values with a smaller discrete set of levels, often to reduce model memory, transfer or computation cost. Speed depends on kernels and hardware; lower precision can reduce task quality.

R

Term Definition Distinction or example
RAG — Retrieval-Augmented Generation Generating an answer using information retrieved from an external collection as additional context. Retrieval can be lexical, vector, structured or hybrid; a vector database is not mandatory.
RBAC — Role-Based Access Control Assigning permissions to roles and authorizing users through their assigned roles. A role usually still needs resource or tenant scope.
ReAct An approach that interleaves reasoning and actions, incorporating observations from those actions into subsequent steps. The runtime executes tools and enforces limits; exposing private reasoning is not required.
Recall The fraction of actual positives identified: TP / (TP + FN). Retrieval recall measures recovered relevant items under a stated relevance set and cutoff. High precision can coexist with poor recall.
Reranking Rescoring an initial set of retrieval candidates to change their order. A reranker cannot recover a relevant item absent from its candidate set.
Response cache Storage of completed application outputs for reuse when a new request satisfies the cache's matching and validity rules. Exact and semantic matching need permission scope, versioning and freshness checks.
RLHF — Reinforcement Learning from Human Feedback Reinforcement learning that uses a reward signal derived from human feedback, often through a learned reward model. Human preference is not identical to factual truth or safe behavior.
RLVR — Reinforcement Learning with Verifiable Rewards Reinforcement learning using rewards computed by checks such as answer verifiers or executable tests. Incomplete or exploitable verifiers can reward incorrect behavior.

S

Term Definition Distinction or example
Self-consistency Sampling several reasoning paths and aggregating their final answers, commonly by voting. Repeated agreement does not prove correctness when errors share a cause.
Semantic search Retrieving information using representations or methods intended to capture meaning beyond exact term overlap. Dense embeddings are common; literal identifiers and authorization still matter.
SLI — Service-Level Indicator A quantitative measure of a service's behavior, such as the fraction of valid requests completed within a deadline. State numerator, denominator, window and exclusions.
SLO — Service-Level Objective A target for an SLI over a specified period. A p99 latency objective is not a guarantee that every request meets that latency.
Speculative decoding Proposing candidate tokens through a cheaper mechanism and verifying them with the target model to accelerate decoding. Exact algorithms preserve the target sampling distribution under their assumptions; acceptance and speedup depend on workload.
Speech-to-speech — S2S A model or pipeline that takes speech input and produces speech output; direct audio models differ from cascaded recognition, text reasoning and synthesis. Naturalness, latency and control require measurement, not an architecture label.
State-handle hijacking Unauthorized access to application state by reusing or guessing an identifier that addresses another user's state. Bind handles to authenticated principals and operations; unpredictability alone is not authorization.
Structured outputs Output constrained to a supported structural contract, commonly a JSON Schema, using the provider or runtime's documented enforcement. Schema conformance does not establish factual or business correctness; refusals and incomplete outputs need separate handling.
SWE-bench Verified A human-validated, 500-instance subset of SWE-bench repository issue-resolution tasks. Identify dataset, harness, budget and contamination concerns; it is not a complete measure of software engineering.
System prompt Application-supplied instructions that set intended model behavior using the provider's message or instruction interface. Instructions are not secrets or an enforceable access-control boundary.

T

Term Definition Distinction or example
Temperature A sampling parameter that rescales logits before probability normalization; lower positive values concentrate probability on higher-scoring choices. A zero setting commonly selects greedily, but does not guarantee end-to-end determinism.
Test-time compute / inference-time scaling Allocating computation during prediction, for example to longer reasoning, repeated samples or search. In the usual frozen-model comparison, more compute need not improve every task; distinguish adaptation that updates parameters or learned state.
Test-time training — TTT Adapting trainable parameters or learned state using a test input or a related self-supervised objective during inference. The adapted object and persistence depend on the method; it is not always a temporary LoRA discarded after one answer.
Token A unit represented by a tokenizer's vocabulary, often a subword, byte sequence or special symbol; multimodal interfaces also account for non-text inputs. Characters and words do not convert to tokens by a universal fixed ratio.
Tool use A model-powered application's use of external functions, services or environments to retrieve information or perform actions. Separate requested action, authorized execution and confirmed result.
Transformer A neural architecture built around attention, feed-forward transformations and residual connections, with positional information and normalization in its blocks. Encoder, decoder, encoder–decoder and hybrid designs differ. Original paper.
TTFT — Time to First Token Time from a defined request start to receipt of the first generated token. State the measurement boundary; a fast first token does not mean a fast completed answer.

V

Term Definition Distinction or example
Vector database A database or database capability for storing vectors and performing similarity or nearest-neighbor queries with associated data. Index choice, filters, updates and consistency are separate design decisions.

W

Term Definition Distinction or example
Windsurf An earlier coding-product name still found in tutorials and comparisons. Its documentation entry now directs readers to Devin Desktop. Use the current product documentation for supported features and migration details.
Workflow A defined arrangement of steps and transitions that carries out a process. It can include model calls without allowing the model to choose every transition.

Z

Term Definition Distinction or example
Zero-shot prompting Asking a model to perform a task without task demonstrations in the prompt. Instructions and external evidence can still be present.

Interview recall checks

Try each answer aloud before opening the explanation.

  1. Can a schema-valid response still be unsafe?
    AnswerYes. Structure cannot establish identity, permission, factual truth or business-rule compliance. Validate those separately.
  2. Why does an agent need more than a tool-calling model?
    AnswerIt needs execution logic, state, permissions, budgets, recovery and a completion test. A proposed call is only one part.
  3. Does a durable workflow prevent duplicate refunds?
    AnswerNot on its own. Deduplicate at the payment boundary and reconcile unknown outcomes before issuing another operation.
  4. What does CAP's C mean?
    AnswerLinearizability: operations behave as though performed atomically in an order compatible with real time. It is not ACID's invariant-preservation meaning of consistency.
  5. What is wrong with “60% pass@1 implies 25% pass^8”?
    AnswerThat conclusion does not follow. Under a fixed independent 0.6 success probability, all eight succeed with probability 0.6⁸ = 1.679616%. Real benchmarks need their own repeated-trial estimator.
  6. Does high cosine similarity establish factual support?
    AnswerNo. Similarity helps retrieve candidates. Whether a passage entails a particular claim is a different test.
  7. Can a correctly quoted answer be wrong?
    AnswerYes. Its source can be stale or incorrect. Check both faithfulness to the source and the source's validity for the question.
  8. Must RAG use a vector database?
    AnswerNo. Lexical, relational, graph, vector and hybrid retrieval can provide evidence for generation.
  9. Why is temperature zero insufficient for reproducibility?
    AnswerModel versions, hardware computation, routing, tools and changing inputs can still alter results. Record the full execution contract.
  10. Does a small adapter imply a small serving footprint?
    AnswerNo. Base-model weights, activations, caches and runtime overhead remain. Count the full deployed configuration.
  11. What must a cache key capture besides the question?
    AnswerThe relevant tenant/principal scope, permissions or policy version, content and model versions, and other answer-affecting inputs. Freshness also needs explicit rules.
  12. When does adding a second model hurt reliability?
    AnswerWhen shared errors survive aggregation or the extra dependencies, latency, inconsistent features and operating complexity exceed the measured benefit.
  13. Does a model's declared confidence establish calibration?
    AnswerNo. Compare predicted probabilities with observed outcomes on an appropriate held-out population.
  14. Can a prompt file enforce least privilege?
    AnswerNo. It can communicate instructions. Credential scopes, authorization checks and execution boundaries enforce privileges.
  15. How should you use an informal term in an interview?
    AnswerDefine the concrete mechanism first. A term such as “loop engineering” is useful shorthand only when both people understand its scope.

Final notes

  1. Give the standard definition before choosing a product or drawing an architecture.
  2. State which property you mean when a term is overloaded: consistency, reliability, memory and grounding all need context.
  3. Separate a mechanism from an expected benefit. Caching can save work; it does not guarantee a particular saving.
  4. Link a definition to an observable test, such as allowed read histories, valid tool outcomes or measured claim support.
  5. For product APIs, use the dated model guide and framework guide. For design choices, continue to the pattern reference.

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← AI Architecture Pattern Reference
NEXT LESSONAI Evaluation Lab: Build, Inspect, and Compare →

Explore the diagram