Use this reference to prepare a clear first sentence, then explain the example and tradeoff in the linked lesson. Definitions describe mechanisms; product names and informal practitioner labels are identified separately. Reviewed September 24, 2026.
Distinctions to remember
| Often confused | The distinction | Interview example |
|---|---|---|
| Model / application / agent | A model computes outputs; an application adds data and business logic; an agent can select actions within that application. | A refund proposal still needs authorization and an idempotent payment operation. |
| RAG / fine-tuning / memory | Retrieval supplies evidence; fine-tuning changes learned parameters; memory retains information across interactions. | Retrieve today's policy; fine-tune a stable extraction behavior; store an approved preference. |
| JSON / schema / meaning | Parsing, structural conformance, and business correctness are separate checks. | {"amount":-100} can be valid JSON and still violate a payment rule. |
| Factuality / faithfulness | Factuality concerns truth; faithfulness concerns consistency with the supplied source. | Repeating an obsolete policy can be faithful and factually wrong today. |
| KV cache / response cache | A KV cache saves attention computation; a response cache returns an earlier application answer. | Neither a shared prefix nor a similar question establishes permission to share private data. |
| Durability / idempotency | Durability preserves recorded state; idempotency prevents an operation from taking additional effect when repeated. | A crash-safe workflow can still duplicate a charge if the payment boundary lacks deduplication. |
| pass@k / pass^k | At least one success / success on every attempt, under the stated evaluation protocol. | For independent attempts with fixed success probability 0.6, at k = 8: 99.934464% / 1.679616%. These are illustrative probabilities, not estimates from an aggregate benchmark score. |
| CAP consistency / ACID consistency | CAP uses linearizability; ACID consistency concerns preservation of application/database invariants. | A replicated read order and a nonnegative inventory invariant are different properties. |
Architecture / visual model
flowchart LR
U[Authenticated request] --> A[Application policy and context]
M[Approved memory] --> A
R[Authorized retrieval] --> A
A --> L[Model inference]
L --> O[Answer validation]
L --> P[Proposed tool call]
P --> G[Authorization and business rules]
G --> T[Idempotent operation]
T --> A
O --> D[User result]
Read diagram source
flowchart LR
U[Authenticated request] --> A[Application policy and context]
M[Approved memory] --> A
R[Authorized retrieval] --> A
A --> L[Model inference]
L --> O[Answer validation]
L --> P[Proposed tool call]
P --> G[Authorization and business rules]
G --> T[Idempotent operation]
T --> A
O --> D[User result]
The model can propose an action. Application code decides whether that action is allowed and records its outcome.
A
| Term | Definition | Distinction or example |
|---|---|---|
| A2A | Agent2Agent is a protocol for communication and task coordination between agent applications. | It does not provide a shared authorization policy or make remote agents trustworthy. |
| ABAC — Attribute-Based Access Control | Authorization evaluates attributes of the subject, object, requested operation and, where relevant, environment against policy. | A document's tenant, classification and the user's clearance can all matter. |
| ACID | Atomicity, consistency, isolation and durability are transaction properties. Consistency means a valid transaction preserves declared invariants. | ACID does not specify a single isolation level or guarantee linearizable distributed reads. |
| Accuracy | In classification, the fraction of evaluated examples classified correctly. | At 0.1% fraud prevalence, predicting “not fraud” everywhere scores 99.9% accuracy and detects no fraud. |
| Advisor / executor | An orchestration pattern in which an executor consults another model for advice at selected decision points. | Measure consultation frequency, added latency and full task cost; savings are not inherent. |
| Agentic coding | Software work in which an agent selects and performs steps such as reading files, proposing edits and running checks. | Passing the checks it chose does not establish that every requirement is satisfied. |
| Agentic system | An application in which a model selects actions or workflow steps toward a goal, subject to runtime limits. | A fixed classification pipeline need not be an agent. Autonomy is a design choice. |
| Agent Plugins | A named packaging specification for distributing agent skills and MCP server configurations with a root plugin.json. |
Version 1.0 standardizes those components; vendor-specific extensions require compatible clients. Specification. |
| AI control | A research approach that evaluates deployment protocols under the assumption that some models may behave adversarially or be misaligned. | A monitor's success under one threat model is not a general safety guarantee. |
| AI gateway | An intermediary that applies common policies to requests sent to model services, such as routing, authentication, quotas and telemetry. | Request handling is a data-plane function; policy management is a control-plane function. One API shape is optional. |
| Attention mechanism | A computation that combines value vectors using weights derived from query–key compatibility. Scaled dot-product attention commonly normalizes masked scores with softmax. | Self-attention derives queries, keys and values from the same sequence; causal masking excludes future positions. |
| Authentication | Establishing that a claimant controls credentials associated with an identity. | An authenticated employee may still lack permission to read payroll. |
| Authorization | Deciding whether a principal may perform an operation on a resource under the applicable policy. | Check at retrieval and tool execution, not only at login. |
| Availability | In CAP, every request received by a non-failing node eventually receives the response required by the operation, rather than an error caused by sacrificing availability. Operational availability is separately measured using a stated success and time criterion. | CAP's formal property is not a monthly uptime percentage or a bounded-latency SLO. |
B
| Term | Definition | Distinction or example |
|---|---|---|
| Batching | Processing several requests or items together to share computation or scheduling overhead. Continuous batching changes active requests during generation. | Offline batch APIs and GPU scheduling batches are different contracts. |
| Benchmark saturation | A benchmark has little remaining ability to distinguish the candidates of interest near its measurement ceiling. | Check uncertainty, task difficulty and current candidates before calling a benchmark saturated. |
| BM25 | A lexical ranking function combining query-term matches, inverse document frequency, term-frequency saturation and document-length normalization. | Strong exact identifiers can be missed by an embedding-only search. |
| Budget tokens | A provider-specific token allocation or limit for a model's reasoning or generation. It is not a universal API field. | Reasoning effort may be a qualitative setting rather than a hard token or monetary cap. |
| Bulkhead | Resource isolation that limits how one overloaded or failing workload affects others. | Separate interactive and offline concurrency pools; a shared bottleneck can still couple them. |
C
| Term | Definition | Distinction or example |
|---|---|---|
| C2PA / Content Credentials | Standards and associated credentials for cryptographically verifiable media provenance assertions. | Valid provenance does not prove that a scene is true. Metadata can be removed; recovery mechanisms have limits. |
| Calibration | Agreement between predicted probabilities and observed frequencies across appropriate groups of predictions. | Among predictions scored 0.8, about 80% should be correct under good calibration. Fluent confidence is not calibration. |
| CAP theorem | In an asynchronous distributed system that may partition, it is impossible to guarantee both linearizability and availability for every execution. | During a partition, refusing an unsafe read preserves consistency at the expense of availability. |
| Capability index / composite benchmark | An aggregate of several evaluation scores using a defined normalization and weighting method. | A ranking can change when task weights change; the index is not a universal capability measure. |
| Chain-of-thought — CoT | Intermediate reasoning steps produced before a final answer, either through prompting or a model's trained behavior. | An explanation can be incorrect or fail to faithfully describe how an answer was produced. |
| Chunking | Dividing content into units for indexing, retrieval or downstream processing while retaining useful context and provenance. | Embedding chunks and the larger passages supplied to generation need not have equal size. |
| Circuit breaker | A mechanism that temporarily rejects calls to a failing dependency and later probes recovery. | It reduces repeated pressure; it does not recover a lost transaction outcome. |
| Claude Code | Anthropic's coding-agent product, with documented development surfaces and tool permissions. | CLAUDE.md provides project instructions; it is not an operating-system security boundary. |
| Claude Fable 5 | An earlier model in Anthropic's Fable family, retained here to identify references in older material. | Use the current model catalog and version-specific contract; this name does not mean “current best model.” |
| Cline | An open-source coding-agent project with editor, terminal and integration surfaces. | Product access, model credentials and execution permissions are separate choices. |
| Computer use | Agent interaction with graphical applications through supported observations and actions, such as screenshots, accessibility information, clicks and typing. | Interface actions are less direct evidence of business completion than an authoritative operation status. |
| Consistency | In distributed data systems, a consistency model specifies which histories of reads and writes are permitted. | State the model: linearizable, sequential, causal and eventual consistency have different guarantees. CAP uses linearizability. |
| Context7 | A documentation-retrieval service and MCP integration for looking up library documentation. | Retrieved material still needs the correct version and relevance; installing it does not ensure it is used. |
| Context engineering | Designing how instructions, evidence, conversation and tool results are selected and arranged for model calls. | Include permissions, freshness and token budgets, not just prompt wording. |
| Context window | The maximum token context supported under a model's request contract, often including input and generated output. | Separate maximum input, output and total limits. Usable quality at long lengths must be measured. |
| Context rot | An informal label for quality degradation associated with accumulated, distracting, stale or poorly organized context. | It is not a universal failure threshold at a fixed token count. |
| Cosine similarity | For nonzero vectors, their dot product divided by the product of their lengths. | Undefined for a zero vector. High similarity is not proof of entailment or authorization. |
| Cursor | A development environment with AI-assisted editing and coding-agent capabilities. | Evaluate the actual execution surface and plan, rather than assuming every feature is local. |
D
| Term | Definition | Distinction or example |
|---|---|---|
| Data contamination | Overlap or leakage between evaluation material and information used to train, tune or select a system, compromising the intended assessment. | A private test set can also leak through repeated optimization against its scores. |
| Data drift | A change in the distribution of input data relative to a reference period or dataset. | Changed inputs do not necessarily imply worse performance; measure the resulting outcomes. |
| Diffusion language model | A language model that generates through iterative denoising or refinement of a sequence representation. | Masked, continuous and blockwise approaches differ; fully parallel generation and faster service are not universal. |
| DPO — Direct Preference Optimization | A preference-optimization objective that directly trains a policy from preferred and rejected responses, without a separately fitted reward model in the original formulation. | DPO is not the same optimization procedure as PPO-based RLHF. |
| DSPy | A framework for composing language-model programs and optimizing supported components against examples and a metric. | Optimization can overfit an unreliable metric or a reused development set. |
| Durable execution | Execution supported by persisted progress so a workflow can recover across process failures and long waits under the runtime's guarantees. | Replay and checkpoint designs differ. External effects still need idempotency or reconciliation. |
E
| Term | Definition | Distinction or example |
|---|---|---|
| Effective context length | A task- and test-dependent length over which a model meets a stated quality criterion. | It depends on evidence position, distractors and task difficulty, not just the advertised window. |
| Embedding | A mapping of items into a vector representation used to encode useful relationships. Items can include text, images, audio and other data. | Dense neural embeddings are common; dimensions and distance metrics belong to the model contract. |
| Endpointing / turn detection | Estimating when a speaker's turn has ended so a conversational system can respond. | Silence, semantic completion and explicit signals have different false-cutoff and delay tradeoffs. |
| Ensemble | Combining predictions or outputs from multiple models or runs using a specified aggregation method. | Correlated errors can defeat voting; additional calls also add cost and failure points. |
| Eval awareness | A model recognizing or behaving differently in an evaluation setting. | A behavioral difference requires evidence; benchmark success alone does not demonstrate deployment behavior. |
| Eventual consistency | A model in which replicas eventually converge if updates stop, subject to the system's delivery and conflict-resolution assumptions. | It does not by itself promise a maximum staleness interval or read-your-writes. |
| Extended thinking | Anthropic's terminology for model reasoning before an answer; configuration and exposed reasoning information depend on the model. | Current adaptive modes cannot all be configured through the older budget_tokens field. |
| EU AI Act | EU Regulation 2024/1689 establishes obligations for specified AI systems and general-purpose AI models, with scope- and role-dependent requirements and application dates. | “All obligations started in 2026” is incorrect. Consult the dated governance chapter and official legal text. |
F
| Term | Definition | Distinction or example |
|---|---|---|
| F1 score | The harmonic mean of precision and recall: 2PR / (P + R) when defined. |
Specify averaging and zero-denominator conventions. F1 does not encode every business cost. |
| Faithfulness | In grounded generation, whether an answer's claims are supported by the supplied evidence. Other evaluation tasks may use a different operational definition. | Report unsupported and contradictory claims separately when useful. A citation alone is not support. |
| Few-shot prompting | Supplying a small set of demonstrations in the input to guide a task. | No parameter update is required; examples still consume context and can bias results. |
| Fine-tuning | Further training an existing model by updating its parameters or added trainable parameters using a chosen objective and dataset. | It can target behavior or a domain; it is not a reliable live database update mechanism. |
| FinOps for AI | Collaborative financial accountability and value-based management of AI usage and costs. | Include infrastructure, human review and operations alongside token charges. |
| Framework churn | An informal description of frequent framework, API or dependency changes that create maintenance work. | Version locks improve reproducibility; upgrades still need compatibility and behavior checks. |
| Function calling | A model interface for returning a proposed function name and structured arguments. | The application validates and executes the call; a proposed invocation is not a completed action. |
G
| Term | Definition | Distinction or example |
|---|---|---|
| GGUF | A binary format for model tensors and metadata used in the llama.cpp ecosystem. | It can contain different tensor precisions; a filename alone does not identify quality or runtime compatibility. |
| Graph engineering | A practitioner label for designing workflow nodes, transitions and shared state, including static or dynamically routed graphs. | It is not a universally standardized discipline or evidence that every application needs multiple agents. |
| Grounding | Connecting generated statements or decisions to relevant external evidence or an environment. | Evidence must be current, authorized and actually support the statement. |
| Grok 4.3 | An earlier xAI Grok model version referenced in older comparisons. | Consult current model-selection and pricing chapters rather than treating an old version as the current flagship. |
| GRPO — Group Relative Policy Optimization | An RL method that estimates relative advantages from rewards within groups of sampled outputs, avoiding a separate value model in its original formulation. | Reward design, group diversity and implementation affect both cost and training behavior. |
| Guardrails | Controls that constrain or check an AI application's inputs, outputs, actions and execution. | Deterministic authorization, schemas, limits and model-based checks have different guarantees. |
H
| Term | Definition | Distinction or example |
|---|---|---|
| Hallucination | Generated content that is nonsensical or unfaithful to its source; in factual question answering, the term also covers fabricated or false claims. | Define the task's criterion. Schema errors, execution failures and unsupported claims need different fixes. Research definitions. |
| Harness / scaffold variance | Changes in measured performance caused by prompts, tools, budgets, execution environment or other evaluation scaffolding. | Report the configuration and uncertainty; there is no universal score swing. |
| Harness engineering | A practitioner term for building the runtime around a model: context, tools, state, verification, limits and telemetry. | Runtime controls must function even when model output is invalid. |
| HNSW — Hierarchical Navigable Small World | A graph-based approximate nearest-neighbor search method using a hierarchy of proximity graphs. | Tuning changes memory, index construction, latency and recall; approximate retrieval can miss neighbors. |
| Human-in-the-loop — HITL | Human participation in specified review, approval, correction or escalation steps. | A review queue needs staffing and expiry rules; a checkbox does not guarantee informed approval. |
| Hybrid search | Combining different retrieval signals, commonly lexical and vector retrieval, into a candidate ranking. | Raw scores from unlike retrievers are not automatically comparable. |
I
| Term | Definition | Distinction or example |
|---|---|---|
| Idempotency | Repeating an operation has the same intended effect as applying it once. | A payment key needs correct scope, payload conflict checks and retention; retries alone do not provide this property. |
| In-context learning | A model exhibiting task adaptation from information or demonstrations in its input without updating model parameters. | It changes behavior during that context, not the stored weights. |
| Indirect prompt injection | Instructions embedded in external content attempt to redirect an application's model away from the authorized task. | Treat retrieved documents and tool results as untrusted data; enforce permissions outside the model. |
| Inference | Computing outputs with a trained model for supplied inputs. | Serving also includes queueing, preprocessing, networking and application checks. |
J
| Term | Definition | Distinction or example |
|---|---|---|
| JSON mode | A provider output mode aimed at producing syntactically valid JSON under its supported conditions. | It does not necessarily enforce a schema. Handle refusals, truncation and transport failures separately. |
K
| Term | Definition | Distinction or example |
|---|---|---|
| KV cache | Stored attention key and value tensors reused during decoding to avoid recomputing prior token states. | It is not a key–value application database. Hybrid architectures may also maintain other state. |
L
| Term | Definition | Distinction or example |
|---|---|---|
| LangChain | An open-source framework and integrations for constructing model-powered applications and agents. | Its related hosted products and orchestration packages have separate roles and contracts. |
| Leaderboard Illusion | The title of a research critique of incentives and evaluation practices in model preference leaderboards. | It is not a formal metric. Separate preference, factual correctness and uncertainty; Arena uses Bradley–Terry modeling. |
| Linearizability | Each completed operation appears to take effect at one point between its invocation and response, in a legal sequential history that respects real-time order. | A read started after a write completes must reflect that write or a later one; overlapping operations allow more than one ordering. |
| LlamaIndex | A framework for connecting data to model applications through ingestion, indexing, retrieval and workflow components. | Choosing it does not determine storage consistency, authorization or answer quality. |
| LLM-as-judge | Using a language model to assess an output or behavior against a defined criterion or comparison. | Validate against expert labels; judge errors, position bias and missing outcomes need measurement. |
| LiveCodeBench | A coding benchmark drawing time-stamped problems from programming contests, with tasks and splits defined by its release. | Contest performance does not directly measure repository maintenance or production reliability. |
| LoRA — Low-Rank Adaptation | Parameter-efficient adaptation that represents trainable weight updates using low-rank matrix factors while freezing the base weights. | Training fewer parameters reduces some costs; serving and activation memory still matter. |
| Loop engineering | A practitioner label for designing an agent's repeated observation, action, verification and stopping behavior. | Define progress and external bounds; extra iterations need evidence of benefit. |
| Loopmaxxing | Informal shorthand for increasing iterations as if repetition alone ensured success. | A larger budget cannot repair missing authority, contradictory requirements or an invalid success test. |
M
| Term | Definition | Distinction or example |
|---|---|---|
| Managed agents | Provider-hosted services that operate some of an agent application's execution infrastructure. | Check exactly which persistence, identity, sandbox and recovery guarantees are included. |
| MCP — Model Context Protocol | An open protocol for connecting AI applications with tools, resources and prompts through client–server interactions. | Use the dated protocol revision, currently 2026-07-28; an SDK's major version is separate. |
| Memory poisoning | Malicious or false information is inserted into retained agent memory so it can influence later work. | Validate provenance and write authority; also enforce permissions, expiry and deletion when reading. |
| Mixture of agents — MoA | Combining contributions from multiple model agents, often through an aggregation or synthesis stage. | It is an application-level ensemble, distinct from a model's mixture-of-experts architecture. |
| Mixture of experts — MoE | A model architecture with multiple expert subnetworks and a routing mechanism, often activating a subset for each token. | Active parameter count and total weight memory are different quantities. |
| Model routing | Selecting a model or service for a request according to policy or measured characteristics. | A fallback must preserve required features and data restrictions; cheap-first is only one design. |
| Multi-tenancy | Serving multiple customer or organizational tenants using some shared application or infrastructure resources. | Isolation must cover data, cache, queues, credentials, logs, budgets and execution. |
O
| Term | Definition | Distinction or example |
|---|---|---|
| o3 | An earlier OpenAI reasoning model, retained as historical terminology. | o3-mini is a separate model name, not a configuration flag on o3. Check current availability before using either. |
| OCR — Optical Character Recognition | Recognizing textual characters in images and producing machine-readable text. | Layout, tables, reading order and field extraction are related but distinct tasks. |
| OpenHands | An open-source software-agent platform and SDK for code-related work. | Runtime, sandbox and credential configuration determine its actual execution boundary. |
P
| Term | Definition | Distinction or example |
|---|---|---|
| pass@k | Success on at least one of k attempts under the benchmark's sampling and scoring procedure. | An oracle selecting a successful attempt may be unavailable to the deployed application. |
| pass^k | Success on all k repeated attempts under the stated procedure, used to examine consistency of completion. | Repeated tasks can have correlated failures. Do not compute it by exponentiating an overall measured success rate. τ-bench. |
| Precision | The fraction of predicted positives that are truly positive: TP / (TP + FP). |
Ask how many moderation flags deserve action, and define zero-positive handling. |
| Prefill | Processing the input tokens to establish model state before output-token decoding. | Time to first token also includes queueing, routing and network delays. |
| Prefix caching | Reusing compatible computed state for an identical leading token sequence across requests. | Model, tokenization, positions, configuration and isolation scope must match. |
| Prompt caching | A provider or runtime feature that reuses prompt computation, commonly for repeated prefixes. | Eligibility, expiration, storage/write fees and read discounts vary. No fixed saving applies everywhere. |
| Prompt injection | An attack that uses instructions in input or external content to subvert an application's intended model behavior. | A detector can reduce risk; authorization and action limits must hold if detection fails. |
Q
| Term | Definition | Distinction or example |
|---|---|---|
| QLoRA | A method for training low-rank adapters through a frozen, quantized base model; the original method uses 4-bit NormalFloat, double quantization and paged optimizers. | Trainable adapters and computation need not be 4-bit. |
| Quantization | Representing values with a smaller discrete set of levels, often to reduce model memory, transfer or computation cost. | Speed depends on kernels and hardware; lower precision can reduce task quality. |
R
| Term | Definition | Distinction or example |
|---|---|---|
| RAG — Retrieval-Augmented Generation | Generating an answer using information retrieved from an external collection as additional context. | Retrieval can be lexical, vector, structured or hybrid; a vector database is not mandatory. |
| RBAC — Role-Based Access Control | Assigning permissions to roles and authorizing users through their assigned roles. | A role usually still needs resource or tenant scope. |
| ReAct | An approach that interleaves reasoning and actions, incorporating observations from those actions into subsequent steps. | The runtime executes tools and enforces limits; exposing private reasoning is not required. |
| Recall | The fraction of actual positives identified: TP / (TP + FN). Retrieval recall measures recovered relevant items under a stated relevance set and cutoff. |
High precision can coexist with poor recall. |
| Reranking | Rescoring an initial set of retrieval candidates to change their order. | A reranker cannot recover a relevant item absent from its candidate set. |
| Response cache | Storage of completed application outputs for reuse when a new request satisfies the cache's matching and validity rules. | Exact and semantic matching need permission scope, versioning and freshness checks. |
| RLHF — Reinforcement Learning from Human Feedback | Reinforcement learning that uses a reward signal derived from human feedback, often through a learned reward model. | Human preference is not identical to factual truth or safe behavior. |
| RLVR — Reinforcement Learning with Verifiable Rewards | Reinforcement learning using rewards computed by checks such as answer verifiers or executable tests. | Incomplete or exploitable verifiers can reward incorrect behavior. |
S
| Term | Definition | Distinction or example |
|---|---|---|
| Self-consistency | Sampling several reasoning paths and aggregating their final answers, commonly by voting. | Repeated agreement does not prove correctness when errors share a cause. |
| Semantic search | Retrieving information using representations or methods intended to capture meaning beyond exact term overlap. | Dense embeddings are common; literal identifiers and authorization still matter. |
| SLI — Service-Level Indicator | A quantitative measure of a service's behavior, such as the fraction of valid requests completed within a deadline. | State numerator, denominator, window and exclusions. |
| SLO — Service-Level Objective | A target for an SLI over a specified period. | A p99 latency objective is not a guarantee that every request meets that latency. |
| Speculative decoding | Proposing candidate tokens through a cheaper mechanism and verifying them with the target model to accelerate decoding. | Exact algorithms preserve the target sampling distribution under their assumptions; acceptance and speedup depend on workload. |
| Speech-to-speech — S2S | A model or pipeline that takes speech input and produces speech output; direct audio models differ from cascaded recognition, text reasoning and synthesis. | Naturalness, latency and control require measurement, not an architecture label. |
| State-handle hijacking | Unauthorized access to application state by reusing or guessing an identifier that addresses another user's state. | Bind handles to authenticated principals and operations; unpredictability alone is not authorization. |
| Structured outputs | Output constrained to a supported structural contract, commonly a JSON Schema, using the provider or runtime's documented enforcement. | Schema conformance does not establish factual or business correctness; refusals and incomplete outputs need separate handling. |
| SWE-bench Verified | A human-validated, 500-instance subset of SWE-bench repository issue-resolution tasks. | Identify dataset, harness, budget and contamination concerns; it is not a complete measure of software engineering. |
| System prompt | Application-supplied instructions that set intended model behavior using the provider's message or instruction interface. | Instructions are not secrets or an enforceable access-control boundary. |
T
| Term | Definition | Distinction or example |
|---|---|---|
| Temperature | A sampling parameter that rescales logits before probability normalization; lower positive values concentrate probability on higher-scoring choices. | A zero setting commonly selects greedily, but does not guarantee end-to-end determinism. |
| Test-time compute / inference-time scaling | Allocating computation during prediction, for example to longer reasoning, repeated samples or search. | In the usual frozen-model comparison, more compute need not improve every task; distinguish adaptation that updates parameters or learned state. |
| Test-time training — TTT | Adapting trainable parameters or learned state using a test input or a related self-supervised objective during inference. | The adapted object and persistence depend on the method; it is not always a temporary LoRA discarded after one answer. |
| Token | A unit represented by a tokenizer's vocabulary, often a subword, byte sequence or special symbol; multimodal interfaces also account for non-text inputs. | Characters and words do not convert to tokens by a universal fixed ratio. |
| Tool use | A model-powered application's use of external functions, services or environments to retrieve information or perform actions. | Separate requested action, authorized execution and confirmed result. |
| Transformer | A neural architecture built around attention, feed-forward transformations and residual connections, with positional information and normalization in its blocks. | Encoder, decoder, encoder–decoder and hybrid designs differ. Original paper. |
| TTFT — Time to First Token | Time from a defined request start to receipt of the first generated token. | State the measurement boundary; a fast first token does not mean a fast completed answer. |
V
| Term | Definition | Distinction or example |
|---|---|---|
| Vector database | A database or database capability for storing vectors and performing similarity or nearest-neighbor queries with associated data. | Index choice, filters, updates and consistency are separate design decisions. |
W
| Term | Definition | Distinction or example |
|---|---|---|
| Windsurf | An earlier coding-product name still found in tutorials and comparisons. Its documentation entry now directs readers to Devin Desktop. | Use the current product documentation for supported features and migration details. |
| Workflow | A defined arrangement of steps and transitions that carries out a process. | It can include model calls without allowing the model to choose every transition. |
Z
| Term | Definition | Distinction or example |
|---|---|---|
| Zero-shot prompting | Asking a model to perform a task without task demonstrations in the prompt. | Instructions and external evidence can still be present. |
Interview recall checks
Try each answer aloud before opening the explanation.
- Can a schema-valid response still be unsafe?
Answer
Yes. Structure cannot establish identity, permission, factual truth or business-rule compliance. Validate those separately. - Why does an agent need more than a tool-calling model?
Answer
It needs execution logic, state, permissions, budgets, recovery and a completion test. A proposed call is only one part. - Does a durable workflow prevent duplicate refunds?
Answer
Not on its own. Deduplicate at the payment boundary and reconcile unknown outcomes before issuing another operation. - What does CAP's C mean?
Answer
Linearizability: operations behave as though performed atomically in an order compatible with real time. It is not ACID's invariant-preservation meaning of consistency. - What is wrong with “60% pass@1 implies 25% pass^8”?
Answer
That conclusion does not follow. Under a fixed independent 0.6 success probability, all eight succeed with probability 0.6⁸ = 1.679616%. Real benchmarks need their own repeated-trial estimator. - Does high cosine similarity establish factual support?
Answer
No. Similarity helps retrieve candidates. Whether a passage entails a particular claim is a different test. - Can a correctly quoted answer be wrong?
Answer
Yes. Its source can be stale or incorrect. Check both faithfulness to the source and the source's validity for the question. - Must RAG use a vector database?
Answer
No. Lexical, relational, graph, vector and hybrid retrieval can provide evidence for generation. - Why is temperature zero insufficient for reproducibility?
Answer
Model versions, hardware computation, routing, tools and changing inputs can still alter results. Record the full execution contract. - Does a small adapter imply a small serving footprint?
Answer
No. Base-model weights, activations, caches and runtime overhead remain. Count the full deployed configuration. - What must a cache key capture besides the question?
Answer
The relevant tenant/principal scope, permissions or policy version, content and model versions, and other answer-affecting inputs. Freshness also needs explicit rules. - When does adding a second model hurt reliability?
Answer
When shared errors survive aggregation or the extra dependencies, latency, inconsistent features and operating complexity exceed the measured benefit. - Does a model's declared confidence establish calibration?
Answer
No. Compare predicted probabilities with observed outcomes on an appropriate held-out population. - Can a prompt file enforce least privilege?
Answer
No. It can communicate instructions. Credential scopes, authorization checks and execution boundaries enforce privileges. - How should you use an informal term in an interview?
Answer
Define the concrete mechanism first. A term such as “loop engineering” is useful shorthand only when both people understand its scope.
Final notes
- Give the standard definition before choosing a product or drawing an architecture.
- State which property you mean when a term is overloaded: consistency, reliability, memory and grounding all need context.
- Separate a mechanism from an expected benefit. Caching can save work; it does not guarantee a particular saving.
- Link a definition to an observable test, such as allowed read histories, valid tool outcomes or measured claim support.
- For product APIs, use the dated model guide and framework guide. For design choices, continue to the pattern reference.