An AI agent is a system that selects actions based on observations to pursue an objective. In an LLM-based agent, a language model helps choose actions such as searching, calling an API or revising a plan. Application code executes those actions within defined permissions and limits, then returns observations for the next decision.
There is no requirement that an agent use a specially named reasoning model, a vector database or persistent memory. The important design question is how much control the model has and how the application checks the result.
Distinguish a model, a workflow and an agent
| Component | Responsibility | Example |
|---|---|---|
| Model | Produce a prediction, text or proposed tool call from input | Select a search query |
| Predefined workflow | Execute application-defined steps and branches | Parse a file, validate fields, store a record |
| Agentic workflow | Let the model choose some steps from permitted actions | Search different sources until required evidence is found |
| Runtime | Manage execution, state, authorization and resource limits | Reject an unauthorized tool call and record the outcome |
An application can combine fixed and model-directed steps. A chat interface may front either design. A graph node may be a deterministic function, a tool or a model call; it is not automatically a separate agent. The workflow/agent distinction is also used in Anthropic's architectural guidance. Building effective agents.
Ground the idea in one task
Suppose the task is to identify why a software deployment failed and prepare a repair proposal. Available tools can read deployment logs, inspect repository files and run tests in an isolated workspace.
- Read the failed job's logs.
- Select relevant files or tests based on the observed error.
- Form a candidate repair and verify it in the workspace.
- Report what changed, which checks passed and what remains unresolved.
The next file or test may depend on the previous result, which makes model-directed choice useful. Deploying the repair is a separate capability with its own authorization policy. A successful test command is evidence for a particular check, not proof that every production requirement is satisfied.
Define the control loop
Read diagram source
flowchart TD
G[Objective and success criteria] --> D[Model proposes next step]
S[Relevant task state and observations] --> D
D --> P{Policy and budget permit step?}
P -->|Yes| T[Runtime executes bounded action]
T --> O[Validate and record observation]
O --> V{Success criteria satisfied?}
V -->|Yes| F[Report verified result]
V -->|No, useful step remains| S
V -->|Missing input or no progress| H[Clarify or return incomplete outcome]
P -->|No| H
The model proposes; the runtime decides whether the action is permitted and executes it. Tool output can contain mistakes or malicious instructions, so preserve its provenance and interpret it under the task's rules. A tool result cannot grant itself additional authority.
Name the components precisely
| Component | Data or behavior it owns | Common mistake |
|---|---|---|
| Objective and acceptance criteria | Requested outcome and constraints | Treating any fluent answer as completion |
| Model policy | Selection among available actions | Assuming longer reasoning guarantees a correct choice |
| Tool interface | Typed arguments, results and error semantics | Passing unconstrained text to privileged execution |
| Task state | Completed actions, evidence, open issues and operation IDs | Using an unstructured transcript as the only recovery record |
| Evaluator | Tests or rubric for success and failure | Treating a model's self-score as ground truth |
| Runtime | Deadlines, retries, authorization, checkpoints and cancellation | Letting the model reset its own limits |
Useful tools have clear inputs, outputs, side effects, deadlines and retry behavior. MCP can standardize integration, but a protocol connection does not itself sandbox execution or authorize a business action.
Separate context, memory and compute caches
Context is information supplied to a particular model invocation. Task state is the application's record of the ongoing operation. Persistent memory retains selected information across interactions. A KV cache reuses transformer attention computations for a compatible token prefix; it is not the same thing as application memory or a durable task record.
An agent may need only current observations and a small state object. Long-running work may require durable checkpoints. Cross-session personalization may justify persistent memory, with provenance, access rules, retention and correction. Adding a vector database does not determine which facts are true or appropriate to remember. See memory and state.
Describe autonomy without inventing a universal ladder
Terms such as “L0 to L4 agency” are not a single globally accepted classification. Describe concrete dimensions instead:
| Dimension | Questions to answer |
|---|---|
| Action selection | Fixed sequence, model-chosen tool, or model-created plan? |
| Capability | Read, draft, modify, purchase, publish or deploy? |
| Duration | One request, resumable task or background monitoring? |
| Oversight | Which outcomes can proceed under existing authorization? |
| Recovery | What happens after interruption or an uncertain side effect? |
A background schedule describes when work starts, not how intelligent or autonomous the system is. A tightly bounded writing tool can have more consequential permissions than an elaborate read-only research agent.
Choose model effort using evidence
“System 1” and “System 2” are cognitive analogies, not precise model architecture specifications. Both ordinary and reasoning-oriented language models generate tokens; their training, inference policies and available compute can differ. Do not claim that one only predicts words while the other performs guaranteed logical reasoning.
Extra inference effort may improve difficult decisions and may also increase latency or produce more wrong reasoning. Compare task success, tool-call correctness, recovery, constraint adherence and total cost on representative trajectories. Training and post-training can develop reasoning behavior; it should not be universally attributed to pretraining alone. See reasoning training and reasoning loops.
Use cheaper or lower-effort decisions where measured performance is adequate. Policy enforcement remains in code even if a stronger model appears more reliable.
Prevent drift and recognize completion
- Keep the original objective and constraints in durable task state.
- Record evidence-backed progress and unresolved requirements.
- Preserve constraints during context compaction and replanning.
- Detect repeated failed actions and lack of new evidence.
- Enforce elapsed-time, token, action and spending limits outside the model.
- Check task-specific acceptance criteria before claiming completion.
A second model can review progress, but is an optional fallible evaluator with its own cost. It is not a universal remedy for drift. For a code task, compile/test results and inspected changes may be better evidence than another model's approval score.
Consider an illustrative five-step task where each necessary step succeeds independently with probability 0.95. All five succeed with probability 0.95^5 ≈ 0.774. This is a simplified model, not a measured agent success rate; real errors can be correlated and recovery can change the outcome. Evaluate the complete task as well as individual calls.
Interview practice
Q1: What makes this application agentic?
The model chooses some actions or control-flow steps based on observations while pursuing an objective. I would identify those decisions explicitly and contrast them with steps fixed by code. The number of prompts or the presence of a chat window is not sufficient.
Q2: Does an agent require persistent memory?
No. A short task can operate on current context and transient state. Persistent memory is useful when the task requires continuity, but introduces storage, provenance, correction and access responsibilities.
Q3: Is a reasoning model automatically safer for tool use?
No. It may perform better on an evaluated task, but the runtime must still validate arguments, authorize actions and control retries and budgets. More inference compute does not provide a security boundary.
Q4: How do you prevent an agent from wandering away from its task?
Preserve the objective and constraints, maintain explicit progress and open issues, check proposed actions against the allowed scope, and detect no-progress loops. Use task-specific tests and bounded execution rather than trusting a self-reported confidence number.
Q5: What is the difference between a failed action and an unknown outcome?
A confirmed failure establishes that the intended result did not occur under the tool's contract. A timeout may mean the result occurred but its acknowledgment was lost. Recovery must inspect operation state or use idempotent semantics before repeating a side effect.
Q6: How do you know the agent is done?
Match the result against explicit acceptance criteria and supporting evidence. Report partial completion and uncertainty when criteria remain unmet. A final model message, exhausted budget or lack of another tool call is not proof of success.
Final notes
Recall card: Objective → observation → proposed action → authorized execution → checked result. Build the smallest amount of model discretion that solves the task and measure the entire trajectory.