Learnastra AI SYSTEM DESIGNAnup Rai

Concept · Understand the mechanism

Architecture patterns for dependable tool-use agents

By Anup Rai14 min readReviewed September 2026

A tool-use architecture separates model decisions from authorized execution and verified outcomes. The model proposes what to do. The application validates the request, enforces access and limits, executes it, and records what actually happened. A successful API response or a fluent final answer does not by itself establish task success.

Four useful patterns are typed tool calling, computer use, code execution, and multi-agent orchestration. The first three describe action interfaces; the fourth describes how work is divided. They can be combined. Start with a single controller and the narrowest practical tools, then add complexity for a concrete requirement.

Start with one running design problem

Interview prompt: “Build an internal support assistant that retrieves order information, prepares a refund, and can use a legacy console when the refund API is unavailable.” The following requirements and estimates are hypothetical assumptions to agree with the interviewer.

Functional requirements

  1. Identify the authenticated employee and the customer/order in scope.
  2. Retrieve order status and relevant policy with source references.
  3. Prepare a refund proposal with order, amount, currency, recipient, and reason.
  4. Execute only refunds permitted by business policy and any required approval.
  5. Support a controlled legacy-console path for a documented subset of cases.
  6. Report progress, cancellation, failure, or an uncertain result accurately.
  7. Preserve enough operation history to resume or reconcile interrupted work.

Non-functional requirements

  1. No access to another tenant's orders or credentials.
  2. No duplicate refund from retries, reconnects, or concurrent workers.
  3. A 30-second target for a read-only answer; mutation completion may be asynchronous.
  4. A 120-second active-run deadline, at most 12 tool attempts, and explicit spend limits.
  5. Bounded execution resources and auditable policy/approval decisions.
  6. Recoverable state after worker loss; no automatic success claim for unknown outcomes.

A requirement such as “no duplicate refund” demands cooperation from the authoritative payment system or a controlled transactional adapter. It cannot be guaranteed by asking the model to be careful.

Pattern 1: typed tool calling

Tool calling lets a model request a named operation with structured arguments. Application code interprets the request under a defined contract. The tool may be an in-process function, an HTTP adapter, or an MCP tool. MCP is useful for interoperable integrations; an internal function does not need a separate server merely to be production-ready.

Architecture / visual model
sequenceDiagram participant M as Model participant H as Agent host participant A as Authorization and policy participant T as Order tool participant D as Authoritative order store M->>H: lookup_order(order_id) H->>H: Validate schema and reserve attempt budget H->>A: Authenticated principal and requested order A-->>H: Permit or deny H->>T: Permitted request with deadline T->>D: Query within tenant and object scope D-->>T: Authorized fields or not found T-->>H: Bounded result, source revision and timestamp H-->>M: Observation linked to tool-call ID
Read diagram source
sequenceDiagram
    participant M as Model
    participant H as Agent host
    participant A as Authorization and policy
    participant T as Order tool
    participant D as Authoritative order store
    M->>H: lookup_order(order_id)
    H->>H: Validate schema and reserve attempt budget
    H->>A: Authenticated principal and requested order
    A-->>H: Permit or deny
    H->>T: Permitted request with deadline
    T->>D: Query within tenant and object scope
    D-->>T: Authorized fields or not found
    T-->>H: Bounded result, source revision and timestamp
    H-->>M: Observation linked to tool-call ID

The sequence shows the permitted path. A denied request does not reach the store. Identity comes from the authenticated session, never from an argument such as tenant_id invented by the model. When authorization and data retrieval occur separately, account for changes between the checks; enforce the effective policy again at the source boundary.

Design the contract before the wrapper

Contract element Refund assistant example
Operation lookup_order, prepare_refund, commit_refund are separate capabilities
Input Canonical order ID; money in explicitly defined minor units and currency
Authority Employee/tenant identity and policy context supplied by trusted host
Preconditions Order is refundable; amount within remaining refundable balance
Output Typed status, permitted fields, source version, evidence reference
Side effects Lookup is read-only; commit changes external state
Failure Invalid input, denied, not found, unavailable, conflict, unknown outcome
Recovery Stable operation key, deduplication retention, status lookup, reconciliation

A schema validates shape. It does not prove ownership, available balance, or user intent. A regular expression rejecting DROP TABLE is not a SQL security boundary. Use parameterized domain operations, database permissions, and bounded queries.

A normalized application observation might look like this; it is not a provider-specific message schema:

{
  "call_id": "call-14",
  "operation_id": "refund-8c12",
  "status": "unknown",
  "resource": "order-2031",
  "evidence_ref": "audit-927",
  "next_action": "reconcile"
}

The host translates model-provider calls and MCP messages at the integration boundary. Raw MCP tool schemas are not automatically valid Anthropic or OpenAI tool definitions. Process every returned tool call and preserve its ID. Append the model's message once, then the corresponding results in the provider's required format; do not duplicate the whole assistant message for each call. See tool use and MCP.

Tradeoff: narrow typed tools simplify authorization, testing, and recovery but require adapter development and maintenance. A generic shell or unrestricted SQL tool is easier to expose and much harder to constrain.

Pattern 2: computer use

Computer use is an observation/action loop over a graphical application. The observation can contain screen images, browser structure, accessibility information, or a supported combination. A visual model can propose keyboard/mouse input, but your runtime must apply it to the correct environment and verify the effect.

Architecture / visual model
flowchart TD S[Current application observation] --> M[Model proposes action] M --> G{Identity, policy, state<br/>and approval still valid?} G -->|No| H[Stop, reobserve or request required decision] G -->|Yes| X[Execute one bounded action] X --> O[Observe resulting application state] O --> V{Business postcondition verified?} V -->|Yes| D[Record evidence and completion] V -->|No, safe continuation| S V -->|Submission outcome uncertain| R[Reconcile with authoritative system]
Read diagram source
flowchart TD
    S[Current application observation] --> M[Model proposes action]
    M --> G{Identity, policy, state<br/>and approval still valid?}
    G -->|No| H[Stop, reobserve or request required decision]
    G -->|Yes| X[Execute one bounded action]
    X --> O[Observe resulting application state]
    O --> V{Business postcondition verified?}
    V -->|Yes| D[Record evidence and completion]
    V -->|No, safe continuation| S
    V -->|Submission outcome uncertain| R[Reconcile with authoritative system]

For the legacy refund console:

  1. Open the intended customer/order in the isolated browser session.
  2. Observe the order identity, refundable balance, currency, and current form state.
  3. Fill a draft and compare every material field with the prepared proposal.
  4. Bind any required approval to that exact proposal and current policy.
  5. Submit through one authorized worker; immediately record the returned reference.
  6. If the session disconnects during submission, inspect authoritative refund status before another submission.

A screenshot can be stale by the time input arrives. Screen scaling, scroll offsets, overlays, and multiple windows can change the meaning of coordinates. Browser locators and accessibility nodes can also become stale. Reobserve when state changes; prefer domain identifiers over a remembered coordinate.

Do not treat “button clicked” or “page displayed a success banner” as sufficient evidence for a high-impact action. Use an authoritative transaction/reference lookup when available. A VM protects the host; it does not undo a refund sent through the VM's valid login.

Tradeoff: computer use reaches workflows without adequate APIs, but adds perception errors, application waits, fragile state, and more expensive verification. There is no universal 1–3-second step latency. Measure capture, inference, action, page response, and recovery separately. Current model/tool contracts are documented in the computer-use lesson.

Pattern 3: generated code execution

Code execution lets a model write a program that performs a task inside a controlled runtime. Locality is a deployment choice; the same pattern can run on a workstation, server, or remote sandbox.

Architecture / visual model
flowchart LR T[Task and bounded input files] --> M[Model produces code artifact] M --> C[Validate artifact metadata<br/>and permitted execution scope] C --> S[Isolated executor<br/>time, memory, disk and network limits] S --> O[Exit state, bounded logs<br/>and output artifacts] O --> V[Independent outcome checks] V -->|Repair allowed and budget remains| M V -->|Accepted| R[Publish authorized result] V -->|Incomplete or unsafe| F[Stop with explicit status]
Read diagram source
flowchart LR
    T[Task and bounded input files] --> M[Model produces code artifact]
    M --> C[Validate artifact metadata<br/>and permitted execution scope]
    C --> S[Isolated executor<br/>time, memory, disk and network limits]
    S --> O[Exit state, bounded logs<br/>and output artifacts]
    O --> V[Independent outcome checks]
    V -->|Repair allowed and budget remains| M
    V -->|Accepted| R[Publish authorized result]
    V -->|Incomplete or unsafe| F[Stop with explicit status]

A CSV-analysis agent needs input files, a runtime, and an output directory. It usually does not need the employee's home directory, SSH keys, production database, or unrestricted internet access. Separate the agent controller's model connection from the generated program's network policy.

Execution contract

  1. Accept code as a structured tool argument or versioned artifact, not by blindly extracting the first Markdown fence.
  2. Choose the language/runtime from an allowlisted configuration, not an arbitrary executable path.
  3. Allocate a fresh scoped workspace with explicit read/write mounts.
  4. Apply CPU, wall-time, memory, process-count, disk, output-size, and egress limits.
  5. Terminate the process tree on timeout/cancellation and track cleanup separately.
  6. Capture exit status and artifact hashes; redact secrets from logs and returned context.
  7. Validate the requested result, then allow a bounded repair or return an explicit failure.

“Allow git; deny rm -rf” is inadequate isolation. Allowed programs, hooks, configuration, tests, and alternate command spellings can execute arbitrary code. User review helps decide intent; operating-system and service controls enforce the actual scope.

Tradeoff: code can efficiently transform large data without putting every row into model context, but it expands the executable attack surface and environment-maintenance burden. A repair loop may correct an error or weaken a test to hide it. Protect independent validation from the generated code.

Pattern 4: multi-agent orchestration

Multi-agent orchestration divides work among separate agent roles or runs and coordinates their results. A separate prompt or context window does not imply a separate process, filesystem, identity, or security domain. Decide those independently.

Structure Use when Main cost or failure
Router One specialist can handle the whole task Misrouting and lost context during handoff
Plan and execute Dependencies can be named and checked Bad plan propagates; replanning adds latency
Parallel specialists Independent subtasks have separate outputs Duplicate work, spend, inconsistent source versions
Hierarchical delegation Large tasks need bounded subtask ownership Depth, fan-out, and inherited authority grow
Peer collaboration Specialists need iterative shared problem solving Coordination loops and unclear ownership
Architecture / visual model
flowchart TD T[Task and shared budget] --> P[Coordinator builds dependency graph] P --> A[Policy research<br/>read-only source scope] P --> B[Order analysis<br/>authorized order scope] A --> J[Join results with source versions] B --> J J --> V[Validate a refund proposal] V --> C[One authorized commit path] C --> R[Reconcile and report]
Read diagram source
flowchart TD
    T[Task and shared budget] --> P[Coordinator builds dependency graph]
    P --> A[Policy research<br/>read-only source scope]
    P --> B[Order analysis<br/>authorized order scope]
    A --> J[Join results with source versions]
    B --> J
    J --> V[Validate a refund proposal]
    V --> C[One authorized commit path]
    C --> R[Reconcile and report]

Keep mutation ownership explicit. Two specialists should not independently commit the same refund. One shared budget service should atomically reserve attempts/spend across workers. Each child receives a narrower or equal scope, an output contract, a deadline, and cancellation behavior. A model's delegation request cannot grant additional authority.

When splitting saves money

Use hypothetical measured costs, not a universal “87%” or “90%” saving:

Component All-capable-model workflow Planner plus smaller workers
Planning/aggregation Included $0.20
Four work steps Included 4 × $0.08 = $0.32
Initial model cost $1.00 $0.52
25% of tasks need a full $1 repair — Expected $0.25
Expected model cost $1.00 $0.77

The expected saving is 23% before orchestration and review costs, assuming comparable accepted outcomes. If repairs become necessary on 60% of tasks, the split costs $1.12. Threshold: with a $1 repair, the repair fraction must be below 48% to beat $1 on model cost alone.

If four independent steps each take 5 seconds and coordination takes 2 seconds, ideal elapsed time is 7 rather than 22 seconds. This assumes simultaneous capacity and no dependencies; total worker compute is still 20 seconds plus coordination. See multi-agent design.

Isolation choices and their actual boundaries

Mechanism Boundary What to verify
OS process sandbox Configured filesystem, process, and network restrictions Platform support and every execution path
Docker container Namespaces, resource controls, capabilities, shared host kernel No privileged mode, host socket, excessive mounts, or unintended egress
gVisor Userspace application kernel mediating workload system calls Workload compatibility, configured resources, and runtime integrity
Firecracker microVM KVM-based virtual machine with a small device model Guest images, host hardening, resource sizing, and permitted I/O
WebAssembly runtime Isolated execution with explicitly imported host capabilities Every exposed host function, memory and execution limits
Managed sandbox Provider-operated execution environment Underlying guarantees, region, credentials, persistence, and billing

Docker security, gVisor architecture, Firecracker, and Wasmtime security describe different mechanisms. gVisor is not simply a syscall filter or a conventional VM. WebAssembly is not a full desktop replacement. A managed service such as E2B is a product/deployment layer, not a separate fundamental isolation mechanism.

Do not compare a published microVM boot time with a complete agent startup time. Image retrieval, dependency setup, credentials, browser launch, and application login may dominate readiness. Benchmark the actual cold and warm paths.

A supervised user machine can be an authorized execution environment, but watching the screen is not containment. Use least privilege and recoverable changes even for one user. Separate tenants and risk levels when code or connected content can be hostile.

State: conversation is only one record

State Example Authority and lifetime
Conversation/context Current task, selected tool schemas, recent observations Derived input to the model; can be compacted
Durable task state Step, attempt count, deadline, budget, cancellation Trusted task database across worker restarts
Workspace state Files, browser session, installed packages Execution environment; may expire independently
Operation ledger Refund intent, request hash, idempotency key, outcome Durable coordinator record, linked to source transaction
External business state Refund status and remaining balance Authoritative payment/order system
Long-term memory Prior permitted preferences or facts Retention and authorization rules across tasks

Do not write a file's requested content into a “current file” cache before confirming the write. Do not infer a refund from the model's conversational memory. After a crash, reconcile uncertain operations with the external system.

For money movement, store a durable intent before sending the request and an outcome after observing it. This does not make two independent systems atomic. The gap between external completion and local recording remains; an idempotency contract and status lookup close the operational recovery path.

Architecture / visual model
stateDiagram-v2 [*] --> Prepared Prepared --> Authorized: policy and exact approval valid Authorized --> InFlight: durable intent and attempt reservation InFlight --> Confirmed: authoritative success InFlight --> Rejected: authoritative rejection InFlight --> Unknown: response lost or worker crashed Unknown --> Confirmed: reconciliation finds success Unknown --> Rejected: reconciliation proves failure Unknown --> InFlight: safe same-operation retry contract Confirmed --> [*] Rejected --> [*]
Read diagram source
stateDiagram-v2
    [*] --> Prepared
    Prepared --> Authorized: policy and exact approval valid
    Authorized --> InFlight: durable intent and attempt reservation
    InFlight --> Confirmed: authoritative success
    InFlight --> Rejected: authoritative rejection
    InFlight --> Unknown: response lost or worker crashed
    Unknown --> Confirmed: reconciliation finds success
    Unknown --> Rejected: reconciliation proves failure
    Unknown --> InFlight: safe same-operation retry contract
    Confirmed --> [*]
    Rejected --> [*]

Unknown is a real state, not a transient error to hide. Cancellation stops further dispatch but cannot promise that an already-submitted transaction was undone. See durable execution.

Error handling: retry the operation contract, not the exception name

Observation Appropriate response
Invalid input before execution Return bounded corrective feedback; validate the new proposal again
Access denied Stop that operation; authorized identity repair may be a separate workflow
Rate limit or transient read failure Bounded backoff with jitter and provider retry hints
Write timeout or connection loss after dispatch Mark unknown; reconcile or use a valid same-operation idempotency contract
Conflict/stale version Fetch current state and replan; previous approval may no longer apply
Worker crash Fence the old worker, recover ledger state, then reconcile
Repeated irrelevant action or no progress Stop/replan within a separate repair budget

A 500/503 response does not universally prove a write had no effect. An idempotency key is useful only if the service enforces it for that operation, payload, identity, and retention window. A new key for every retry defeats deduplication. Error recovery and input correction are different: changing an amount creates a new proposal, not a retry of the same approved refund.

Executable retry-decision example

This pure Python function consumes normalized facts from trusted adapters, not arbitrary model output. dedupe_valid means the host has already verified the same operation/key/payload and the service's unexpired deduplication contract. transient denotes a retryable service condition; unknown denotes missing outcome evidence. It chooses the next class of action; it does not implement networking, reconciliation, backoff, or authorization.

import math


def retry_action(*, failure, operation, attempts, remaining_seconds,
                 dedupe_valid=False, max_attempts=3):
    if type(failure) is not str or failure not in {"input", "denied", "conflict", "transient", "unknown"}:
        raise ValueError("Unrecognized failure")
    if type(operation) is not str or operation not in {"read", "write"}:
        raise ValueError("Unrecognized operation")
    if type(attempts) is not int or attempts < 1:
        raise ValueError("attempts includes the initial dispatched attempt")
    if type(max_attempts) is not int or max_attempts < 1:
        raise ValueError("Invalid maximum")
    if type(dedupe_valid) is not bool:
        raise ValueError("Invalid deduplication evidence")
    if type(remaining_seconds) not in {int, float}:
        raise ValueError("Invalid deadline")
    try:
        finite = math.isfinite(remaining_seconds)
    except OverflowError:
        finite = False
    if not finite:
        raise ValueError("Invalid deadline")
    if failure == "denied":
        return "stop_denied"
    # Keep uncertainty visible even after the active-run budget expires.
    if operation == "write" and failure in {"transient", "unknown"}:
        if not dedupe_valid or attempts >= max_attempts or remaining_seconds <= 0:
            return "reconcile"
    if attempts >= max_attempts or remaining_seconds <= 0:
        return "stop_budget"
    if failure == "input":
        return "correct_and_reauthorize"
    if failure == "conflict":
        return "refresh_and_reauthorize"
    return "retry_with_backoff"

max_attempts=3 permits the initial attempt plus at most two further dispatches. Before sleeping/retrying, reserve capacity, honor the actual remaining deadline, and recheck cancellation and authorization. Reconciliation uses its own controlled queue/budget; this return value does not authorize unlimited background polling.

MCP integration: direct, multiple servers, or gateway

Topology Benefit Added responsibility
Direct client/server Few moving parts for a small integration Each host implements connection, identity, and policy correctly
Multiple server clients Independent service ownership and reuse Tool-name namespaces, separate credentials, per-source limits
Central gateway Shared routing, admission, auditing, and policy controls Availability dependency, token brokerage, data exposure, version translation
Architecture / visual model
flowchart LR H[Agent host and model adapter] --> G[Integration gateway<br/>identity, budgets and audit] G --> O[Order MCP server] G --> K[Knowledge MCP server] G --> C[Console task adapter] O --> D[Order service with object authorization] K --> S[Search with source ACLs] C --> W[Isolated GUI worker]
Read diagram source
flowchart LR
    H[Agent host and model adapter] --> G[Integration gateway<br/>identity, budgets and audit]
    G --> O[Order MCP server]
    G --> K[Knowledge MCP server]
    G --> C[Console task adapter]
    O --> D[Order service with object authorization]
    K --> S[Search with source ACLs]
    C --> W[Isolated GUI worker]

Current MCP uses dated protocol revisions. The July 28, 2026 revision changes initialization and request metadata, while individual servers may still support an older revision. The enterprise MCP design demonstrates why connector-by-connector compatibility matters. An adapter must not assume that every server upgrades at once.

MCP does define protocol errors and tool execution errors (isError), including current resultType handling. It does not supply a universal business-specific retry policy. The host maps the actual source contract into its own failure categories. See MCP tool errors.

OAuth/resource authorization and downstream service identity also need explicit configuration. A gateway must not forward a bearer token to an unintended resource or accept model-provided identity as authority. Global spend budgets, approval validity, and cross-system transaction recovery remain application responsibilities.

Small reviewed catalogs can stay static. For a large catalog, discover only trusted servers, validate and version their schemas, namespace tools, and present a relevant subset to the model. A server's tool description is untrusted integration content, not a grant of authority.

Evolve the design by finding flaws

Initial decision Flaw discovered Repair Cost/benefit
One agent with all credentials A retrieved page can influence a privileged action Scoped adapters and separate commit authority Integration work buys enforceable containment
Retry every timeout Refund can be duplicated Durable intent, deduplication, reconciliation More state and latency buy correct recovery
One shared browser Users and jobs interfere Session isolation and exclusive worker lease More capacity buys identity/state separation
Keep everything in messages Restart loses execution truth Durable task/operation records Database complexity buys recoverability
Spawn workers freely Vendor quota and cost explode Shared reservations and bounded fan-out Coordination overhead buys predictable limits
Trust test output Generated code can weaken its own checks Independent validation of exact artifacts Extra compute buys stronger evidence

A support assistant can use API reads, a code-based report generator, and a dedicated GUI fallback. It does not need multiple agents unless specialization or independent parallel work materially improves the result.

Capacity, latency, and operating metrics

Assume 6,000 tasks during an eight-hour day, a 10× peak-to-average ratio, and 20 seconds of mean active task time:

  • Average arrival rate: 6,000 / 28,800 = 0.2083 tasks/s.
  • Peak assumption: 2.083 tasks/s.
  • Mean active concurrency at that sustained peak: 2.083 × 20 ≈ 42.
  • At a 70% occupancy planning target: ceil(41.67 / 0.70) = 60 active-task slots.

These are planning estimates using Little's law and queueing, not proof of a tail-latency SLO. Approval waits should release expensive workers when possible. Model quotas, tool quotas, and GUI-session limits may cap concurrency before CPU does. At six tool attempts per task, peak demand is approximately 12.5 attempts/s before retries.

Track task success with valid evidence, unsafe/duplicate actions, unknown outcomes, time to reconcile, deadline misses, human interventions, and cost per accepted outcome. Keep sensitive payloads in controlled evidence storage; routine metrics need identifiers and counts rather than full customer records. An SLO error budget is the allowed amount of unreliability over a period; it is not the same as a per-task retry or token budget.

Interview questions and answer notes

  1. Why does typed tool calling not guarantee deterministic outcomes? The request format is structured, but external state, concurrency, model selection, and service failures still vary.
  2. Must an API-backed tool be exposed through MCP? No. Use MCP when interoperability or service separation justifies it; an in-process adapter may be sufficient.
  3. A payment timed out. Is exponential backoff enough? No. First establish idempotent same-operation retry or reconcile the unknown outcome.
  4. Why keep an operation ledger outside model memory? It records trusted execution state across restarts and supports deduplication and investigation.
  5. Can a read-only tool exfiltrate data? Yes. Its arguments or returned data can cross a trust boundary. Constrain scope, destinations, and allowed fields.
  6. Why not ask for approval before every tool call? Routine actions may already be authorized. Require a decision when policy or scope needs one, and bind material approval to the exact action.
  7. Does a VM prevent the agent from buying something? No. A valid authenticated browser or API credential can still perform that action.
  8. Why can more agents increase latency? Handoffs, duplicated context, contention, incompatible artifacts, and repair can outweigh parallelism.
  9. What happens when a user cancels a submitted refund? Stop new work, determine the source transaction state, and explain whether a separate authorized reversal is possible.
  10. What is the most useful success metric? A correctly completed intended operation with valid evidence, considered alongside side effects, review effort, latency, and total cost.

Final summary and closing remarks

Recall card Design consequence
Propose → authorize → execute → verify Keep model intent distinct from application authority and truth
Retry ≠ repair The same operation preserves identity; changed arguments need revalidation
Unknown ≠ failed Reconcile before risking a duplicate effect
Context ≠ durable state Persist the records needed for recovery outside the prompt
Parallel ≠ isolated Explicitly separate resources, identities, and mutation ownership
Cheap call ≠ cheap outcome Count failed work, repairs, review, and operations

Closing answer: begin with narrow tools and one bounded controller. Add durable state before background execution, isolate generated code and browser sessions, and keep irreversible actions behind explicit business policy. Use measured quality, recovery, and cost to justify extra agents or more complex infrastructure.

Previous: Tool-use landscape. Next: OpenClaw architecture.

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← Tool-use agents: choose the execution model before the product
NEXT LESSONOpenClaw: designing a persistent assistant around a trusted gateway →

Explore the diagram