Learnastra AI SYSTEM DESIGNAnup Rai

Concept · Understand the mechanism

Agentic Systems: Design, Control and Verification

By Anup Rai4 min readReviewed September 2026

An AI agent uses a model to select actions toward a goal, observes their results and can adapt its next action. A workflow fixes more of that control flow in application code. Real products often combine both: a fixed approval process can contain an agent that investigates a support request.

This Learnastra section develops an interview answer from the smallest useful loop to a system that can recover, respect permissions and demonstrate completion. A multi-agent architecture is one option, not the required destination.

Study in six passes

  1. Explain the decision to use an agent. Start with a fixed workflow and identify which decisions require adaptation. Learn the loop vocabulary before choosing a framework.
  2. Define the execution boundary. Specify tool inputs, identities, permissions and outcomes. Distinguish function calling, MCP interoperability and cross-agent communication.
  3. Manage work and information. Separate task state from remembered facts, then design dependencies, parallel investigations and plan revision.
  4. Handle failure and authority. Classify errors, make retries safe, bind approval to the intended action and constrain execution.
  5. Prove behavior. Evaluate actual task outcomes, unacceptable actions, cost and latency across repeated runs.
  6. Operate longer tasks. Preserve progress, reconcile uncertain effects, enforce shared budgets and stop honestly when work cannot finish.

Chapter map

Chapter Core question to answer aloud
Agent fundamentals Why does this task need adaptive action selection?
Reasoning loops What changes after the agent receives an observation?
Tool use and MCP Who validates and authorizes a model's proposed call?
Multi-agent orchestration Which tasks are independent, and who integrates their results?
Agent memory and state Which facts persist, under whose scope, and with what evidence?
Planning and decomposition Which dependency must finish before another action is safe?
Error handling and recovery Did the action fail, or is its outcome merely unknown?
Human-in-the-loop patterns What exact proposal did the person approve?
Security and sandboxing What prevents retrieved instructions from gaining authority?
Evaluating agents How do we know the task succeeded without prohibited conduct?
Durable execution What survives a crash, and which effects are safe to repeat?
Loop engineering How do we control progress, cost and termination?

Place the components in a working architecture

Architecture / visual model
flowchart TD U[Authenticated request and scope] --> C[Controller: goal, plan, limits] C --> X[Build relevant context] S[(Task state and scoped memory)] --> X X --> M[Model proposes action or answer] M -->|Action proposal| G[Validate arguments and current permission] G -->|Action needs approval| H[Persist proposal and await decision] H -->|Approved proposal| G H -->|Rejected or expired| P[Return recorded partial outcome] G -->|Denied| P G -->|Permitted action| T[Tool service or isolated execution] T --> O[Record result or uncertain outcome] O --> S O --> C M -->|Completion claim| V[Verify required outcome and conduct] V -->|Incomplete| C V -->|Complete| D[Return result with evidence] C -->|Blocked, cancelled or limit reached| P
Read diagram source
flowchart TD
    U[Authenticated request and scope] --> C[Controller: goal, plan, limits]
    C --> X[Build relevant context]
    S[(Task state and scoped memory)] --> X
    X --> M[Model proposes action or answer]
    M -->|Action proposal| G[Validate arguments and current permission]
    G -->|Action needs approval| H[Persist proposal and await decision]
    H -->|Approved proposal| G
    H -->|Rejected or expired| P[Return recorded partial outcome]
    G -->|Denied| P
    G -->|Permitted action| T[Tool service or isolated execution]
    T --> O[Record result or uncertain outcome]
    O --> S
    O --> C
    M -->|Completion claim| V[Verify required outcome and conduct]
    V -->|Incomplete| C
    V -->|Complete| D[Return result with evidence]
    C -->|Blocked, cancelled or limit reached| P

This is a logical architecture, not a requirement for one microservice per box. An ordinary API call does not need a code sandbox, while untrusted generated code needs an appropriate isolation boundary. Memory supplies selected context; it is not a compulsory intermediate step after every tool call. Evaluation also happens offline during development, beyond the runtime checks shown here.

The approval decision must identify the proposal and change its recorded status; the gateway then rechecks that status and current permissions. A rejection or expiry stops the proposed action. It must not create an endless approval loop.

Practice a complete interview response

Use a support agent that may investigate a refund but can issue one only under a defined policy:

  1. List functional requirements: find the order, establish eligibility, propose the refund, obtain any required approval, execute and report the outcome.
  2. State non-functional requirements: tenant isolation, spending limits, response/recovery targets, audit retention and acceptable duplicate-effect risk.
  3. Start with one controller, scoped read tools and an authoritative order/payment service.
  4. Walk through a successful request with actual records and operation identities.
  5. Introduce a failure: the payment succeeds immediately before the worker crashes. Explain deduplication and reconciliation.
  6. Add durable waits, parallel investigation or a different runtime only when a requirement justifies the complexity.
  7. Close with measured success, cost per acceptable outcome, unresolved tradeoffs and the next validation experiment.

Interview tip: A diagram is useful only if you can trace a request, name each trust boundary and explain a failure window. More boxes do not establish a better design.

Quick recall

Remember Avoid claiming
Model proposes; trusted application enforces A prompt is an authorization mechanism
Interoperability and authorization are separate MCP makes every tool safe
Persistent state and model context differ A large context window is a complete memory system
A timeout can leave an unknown outcome No response means no business effect occurred
Replay and receiver deduplication solve different problems A checkpoint alone guarantees exactly-once external effects
Verify outcome, conduct and total cost A fluent final answer proves completion

For deeper implementation choices, continue with memory architecture, LangGraph orchestration, agentic retrieval and reliability patterns.

Final note: Keep the first design small, make its contracts explicit, then add complexity in response to a demonstrated requirement or failure.

Interview questions

Q1. When is a fixed workflow a better starting point than an agent?

When the valid steps and decision rules are known, implement them directly. This makes permissions, retries and testing easier to reason about. Introduce model-directed selection only for decisions whose variability justifies it, and compare completion quality, latency and cost against the fixed workflow.

Q2. A refund worker crashes after the payment succeeds. What should the next worker do?

Recover the task's stable operation identity and query the payment service's authoritative result. Retry with that same idempotency key only when the service contract makes it safe. A missing local checkpoint does not prove that no payment occurred. Record unresolved outcomes and reconcile them before reporting completion.

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← Data Engineering for AI
NEXT LESSONAgent Fundamentals →

Explore the diagram