A tool-use agent is a system in which a model selects actions, software executes allowed actions, and the resulting observations guide subsequent decisions. A tool might search a database, edit a file, run a test, or operate a browser. The model's request is not itself an executed action. The surrounding application determines access, execution, state, and completion.
For interview preparation, compare architectures before comparing brands. “Local,” “cloud,” “IDE,” and “computer use” describe different dimensions; they are not mutually exclusive product categories. Product details below were checked against official documentation on September 24, 2026.
Ecosystem overview
Start with the components that any implementation must place somewhere:
Read diagram source
flowchart LR
U[User task and authenticated identity] --> H[Agent host<br/>state and orchestration]
H <-->|context and proposed actions| M[Model endpoint<br/>local or hosted]
H --> P[Policy and authorization]
P --> X[Tool execution<br/>local, remote or isolated]
X --> S[Files, APIs, browser or desktop]
S --> O[Observation with provenance]
O --> H
H --> V[Completion checks and evidence]
A hosted model can control a local process. A locally installed editor can launch a remote worker. An isolated worker can still possess a token that modifies production. Always identify where data is sent and what authority crosses each boundary.
The related tool-calling lesson covers schema design; execution architecture develops these boundaries into a service.
Category taxonomy
| Dimension | Choices | Question to ask |
|---|---|---|
| Interaction surface | Chat, terminal, editor, web dashboard, schedule | How does a user start and supervise work? |
| Model location | On-device, self-hosted endpoint, provider API | Where do prompts, images, and tool results go? |
| Execution location | User machine, controlled server, managed worker | Which filesystem and network can actions reach? |
| Action interface | Typed API, shell/code, browser structure, screen coordinates | How precisely can we express and verify an action? |
| Isolation | OS sandbox, container, VM, dedicated host | What damage can an allowed process cause? |
| Task lifetime | Interactive turn, background task, durable workflow | What survives a disconnect, restart, or retry? |
| Integration | Built-in tool, MCP server, skill, custom adapter | Who supplies instructions, code, and credentials? |
Self-hosted does not mean offline. You must account for model calls, telemetry, plugins, package downloads, and connected services. Open source does not mean unrestricted data ownership: licenses, provider terms, customer agreements, and deployment choices remain separate.
Four action interfaces
Read diagram source
flowchart TD
R[Required operation] --> A{Supported authorized API?}
A -->|Yes| B[Typed tool with domain validation]
A -->|No| C{Controlled code or CLI interface?}
C -->|Yes| D[Restricted command or code executor]
C -->|No| E{Usable browser or accessibility structure?}
E -->|Yes| F[Semantic element interaction]
E -->|No| G[Visual computer use]
B --> V[Check actual application outcome]
D --> V
F --> V
G --> V
This is a starting preference, not a universal ranking. A broken API may be less useful than a supported UI. A GUI test must exercise the GUI even when a backend API exists. In either case, choose a supported interface and establish an observable success condition.
| Interface | Concrete operation | Typical failure | Repair |
|---|---|---|---|
| Typed API | Retrieve the current user's order status | Valid JSON names someone else's order | Authorize the record on the server |
| Shell/code | Run a repository's test suite | Tests execute untrusted setup code | Isolate the process and remove unnecessary secrets |
| Browser structure | Click a labeled “Save draft” button | Label matches several controls | Scope the locator and verify the saved draft |
| Screen coordinates | Select an icon in a desktop application | Screen changed after observation | Reobserve, relocate, then verify the effect |
Product profiles: what they help you build
The profiles are examples of implementation choices, not endorsements or a complete market census. Avoid universal “safe/unsafe” scores: configuration and granted capabilities determine the exposure.
OpenClaw: a persistent gateway for assistants
OpenClaw connects messaging channels and other entry points to agent sessions, tools, and models. Its gateway owns routing, connections, configuration, and credentials. It is MIT licensed and supports personal and team setups. Installing it locally does not make calls to hosted models local. See the project documentation.
Current documentation separates the trusted gateway from configurable execution environments. Sandboxing is off by default; supported isolation backends include Docker/Podman, SSH, and OpenShell. An explicit sandbox request without a configured runtime fails instead of silently executing on the host. This is more precise than labeling the whole product either “unsandboxed” or “secure by default.” See the trust boundary.
State is now described as database-first, with global and per-agent SQLite stores and version-aware migrations. Workspace instructions and memory artifacts are not substitutes for the runtime database. See state and upgrades.
Interview relevance: separate message identity, session isolation, durable state, and execution authority. A pleasant assistant personality does not enforce access control. Continue with the OpenClaw deep dive.
OpenHands: a software-agent platform with separate components
The current OpenHands architecture distinguishes Agent Canvas, the browser interface, from the Software Agent SDK and Agent Server, which provide agent execution and service access. Managed Cloud and licensed Enterprise offerings are separate deployment choices. SandboxServer and AutomationServer address environment management and scheduled/event-driven work. The legacy local GUI is deprecated; it should not be taught as the current architecture. Each repository has its own license; “everything is MIT” is an unsafe assumption. See the current introduction.
Read diagram source
flowchart LR
C[Agent Canvas or custom client] <-->|conversation API and events| A[Agent Server<br/>Software Agent SDK]
A --> W[Configured workspace and tools]
S[Environment management] --> W
E[Schedule or verified event] --> O[Automation service]
O --> A
W --> D[Diff, logs and test evidence]
Interview relevance: UI, orchestration, execution, and scheduling need independent contracts. A per-task workspace is useful, but its isolation depends on mounts, privileges, network access, and credentials.
Open Interpreter: inspect the current project, not an old tutorial
The current Open Interpreter repository describes an Apache-2.0 Rust coding agent forked from Codex, with configurable harnesses, MCP/ACP integration, and native command sandboxing. It is substantially different from the older Python conversational interpreter and “Computer API” descriptions. Current commands include i and interpreter; harness emulation is not proof of identical behavior to the emulated product. See the current repository.
Interview relevance: a model, harness, execution policy, and editor integration can change independently. Verify the actual release, interfaces, and license you plan to deploy. Do not infer that a local installation must execute without a sandbox or require confirmation before every action.
Claude computer use: a model tool contract, not a hosted desktop
Computer use lets an application supply screen observations and execute model-requested input actions. The application still provides the environment, action implementation, policy, and verification. Current Anthropic documentation uses computer_toolset_20260801 for newer supported models; compatibility varies by model and provider platform. Do not mix a historical tool schema with a current model without checking support. See the computer-use documentation.
Read diagram source
sequenceDiagram
participant H as Agent host
participant M as Vision model
participant G as Policy gate
participant D as Isolated desktop
H->>D: Capture current observation
D-->>H: Screenshot and environment identity
H->>M: Goal and observation
M-->>H: Proposed input action
H->>G: Scope, policy and state checks
G-->>H: Allow, deny or require review
H->>D: Execute allowed action
D-->>H: New observation
H->>H: Check application postcondition
The loop can operate many graphical applications, but not literally every application or task. Rendering, accessibility, authentication, prohibited actions, and visual ambiguity create limits. A click acknowledgment does not prove that a payment or submission succeeded. The computer-use chapter explains state, coordinates, and recovery.
Claude Code and Codex: coding workflows across several surfaces
Claude Code supports terminal, IDE, desktop, and web workflows. It inspects repositories, edits files, runs commands, and uses configured tools. Project guidance, skills, hooks, and memory support the workflow; they are not operating-system isolation. See Claude Code's overview and the detailed lesson.
Codex also has local terminal/editor/desktop surfaces and separate cloud execution. Codex cloud tasks use isolated managed environments and can produce reviewable changes without occupying the local machine. Its documented setup phase and agent phase have different network/secret behavior. Local command sandbox settings do not automatically govern browser or MCP connections. See Codex cloud and security boundaries.
Read diagram source
flowchart LR
T[Task and exact starting revision] --> L[Local controlled workspace]
T --> R[Remote isolated workspace]
L --> P[Patch and verification evidence]
R --> P
P --> C[Independent checks on resulting revision]
C --> M[Repository merge and release policy]
Interview relevance: source isolation with a Git worktree is not a process sandbox. A passing command is not necessarily a valid test of the requirement. Evaluate the final diff, tests, affected behavior, and integration with current main.
Editor and background agents
| Product or surface | Current documented capability | Architectural question |
|---|---|---|
| Cursor Cloud Agents | Isolated VM development environments, background work, computer/browser use, MCP, and multi-repository work | Which source-control identity and secrets enter each worker? |
| Cascade in Devin Desktop | Code, Plan, and Ask modes; current Windsurf documentation redirects here | Is a mode a workflow preference or an enforced capability limit? |
| Cline | Editor and terminal applications, SDK, task board, and several model-access paths | Which actions are preauthorized in this deployment? |
| GitHub Copilot cloud agent | Background research/planning/changes in an ephemeral GitHub Actions environment; distinct from IDE agent mode | Can untrusted repository code access privileged CI credentials? |
| Google Jules | GitHub-connected tasks in a VM, with repository setup and plan review | How is the chosen branch and environment reproduced? |
| Google Antigravity 2.0 | Standalone desktop application for coordinating coding and knowledge-work agents | How are parallel tasks, permissions, and shared resources isolated? |
Do not rank one as “more autonomous” solely from marketing language. Compare the configured actions it can take without another decision and the quality of its recovery. A Plan mode name alone does not prove a read-only security boundary: Cascade's documentation lists all tools for Plan, while Ask has search tools only.
Comparison matrix: choose by required outcome
| Required outcome | Candidate starting point | Essential acceptance evidence |
|---|---|---|
| Personal messaging assistant | OpenClaw with scoped channel/session configuration | No cross-user disclosure; intended recipients; durable task recovery |
| Custom software-agent application | OpenHands SDK/server or another supported agent SDK | Reproducible worker, complete events, cancellation and artifact ownership |
| Interactive repository change | Claude Code, Codex, Cline, Cursor, or Devin Desktop | Focused diff, meaningful tests, protected existing work |
| Background repository maintenance | Managed coding worker or controlled self-hosted service | Exact starting revision, bounded privileges, integrated verification |
| Legacy desktop workflow | Computer-use model plus controlled execution environment | Correct application outcome, duplicate prevention, recoverable state |
| Internal search/operations assistant | Typed domain tools and selected MCP integrations | Per-user authorization, source evidence, bounded writes and expenditure |
This table narrows an evaluation, not selects a winner. Product licensing affects permitted use and distribution; it does not make a runtime secure. A container with a mounted production credential can perform production actions.
Memory, skills, MCP, and ACP are different concepts
| Concept | Definition | What it does not guarantee |
|---|---|---|
| Session state | Records needed to continue the current conversation/task | Correct long-term facts or isolation from other users |
| Long-term memory | Persisted information retrieved across sessions | Truth, current authorization, or appropriate retention |
| Workspace instructions | Project/agent guidance supplied as context | Enforced filesystem, network, or business permissions |
| Skill | Packaged instructions and optionally supporting code/resources | Trustworthiness merely because it is installed |
| MCP | Protocol for exposing tools and context to AI applications | A universal identity, business authorization, or transaction layer |
| ACP | Protocol connecting an agent to compatible clients such as editors | A replacement for the agent's tool-execution policy |
Use memory and state, MCP, and safety for the corresponding design details. Static tool catalogs remain reasonable for a small, reviewed tool set. Dynamic discovery helps larger catalogs but adds availability, provenance, versioning, and policy checks.
What has changed—and how to evaluate it
The cited product documentation shows three practical developments: products span multiple interfaces; background execution is increasingly exposed as a first-class workflow; and tools, skills, and agent-client integrations allow more extensibility. These are observations about supported features, not measured market-share or adoption claims.
Use an evaluation matrix instead of stars, revenue estimates, or leaderboard screenshots:
- Representative tasks: include easy work, missing information, dependency changes, and ambiguous requests.
- Outcome quality: accepted changes or completed business operations, including regressions and incorrect side effects.
- Human effort: review, correction, approval, and escalation time.
- Runtime behavior: completion latency, tool calls, retries, cancellations, and restart recovery.
- Exposure: files, network destinations, third-party content, credentials, and retained artifacts.
- Economics: total cost per accepted outcome under the intended deployment and support model.
A benchmark score describes a model/harness/task-set combination. Benchmark interpretation explains why it is not the probability that your workflow succeeds.
Worked cost comparison
Assume 1,000 independent pilot tasks, with all figures hypothetical and excluding initial development:
| Measured input | Interactive workflow | Background workflow |
|---|---|---|
| Accepted outcomes | 850 | 900 |
| Model/runtime cost per attempted task | $0.30 | $0.80 |
| Average human review/repair per attempted task | 6 minutes | 3 minutes |
| Human time valued at | $60/hour | $60/hour |
| Total cost | $6,300 | $3,800 |
| Cost per accepted outcome | $7.41 | $4.22 |
For the interactive workflow: 1,000 × $0.30 + 1,000 × 6/60 × $60 = $6,300. Background execution is cheaper in this example despite higher model cost because review time falls. If it instead takes 8 minutes of review, its cost becomes $8,800 / 900 = $9.78 per accepted outcome. Include failed attempts in the numerator; add integration, operations, and security costs before purchasing or building.
Interview exercise: select an internal engineering assistant
Prompt: “Support 200 engineers with repository questions, small fixes, and occasional legacy-console work.” Do not immediately answer with a product name.
Functional requirements
- Answer repository questions with references to the relevant revision.
- Propose code changes in a separate, identifiable workspace.
- Run agreed validation and produce a diff with evidence.
- Access approved internal tools under the requesting user's authority.
- Operate the legacy console only for explicitly supported tasks.
- Let users inspect progress, cancel work, and recover an interrupted task.
Non-functional requirements
- Isolate users, workspaces, credentials, and task artifacts.
- Set task deadlines, tool-call budgets, and concurrency limits.
- Record enough evidence to investigate an incorrect change without retaining secrets.
- Preserve uncommitted work and prevent unauthorized release actions.
- Measure accepted outcomes, regressions, review time, and complete cost.
Initial design: interactive coding agent, scoped repository access, isolated execution, and a narrow catalog of internal tools. Return proposals through existing review/CI. Add a remote task queue only when background duration or concurrency justifies it. Keep a dedicated GUI worker for the legacy console so its session and credentials do not spread to ordinary repository tasks.
Likely flaw: a successful test report can refer to an earlier revision. Repair: bind validation and approval to the exact resulting artifact; rerun affected checks after a material change. Cost: additional compute and latency. Benefit: a reviewable assurance about the change actually being merged.
Interview questions and answer notes
- A tool is self-hosted. Can customer code still leave the network? Yes: model inference, telemetry, connected tools, and package/network access may be external. Trace each path.
- An agent runs in a VM. Can it delete production records? Yes, if a reachable service accepts credentials available to that VM. Isolation and authorization are separate controls.
- Which product should a Slack assistant use? First establish identity, task scope, state retention, and tool needs. A messaging gateway is a candidate, not a complete authorization design.
- Why prefer an API to screen clicks for an order lookup? It can provide typed identifiers and explicit errors. You still need record authorization and freshness checks.
- Is a skill equivalent to a tool? No. A skill guides a workflow; its instructions or bundled code may invoke tools. Installation is a supply-chain decision.
- Does MCP remove custom integration work? It standardizes an interface. You still map identities, permissions, data contracts, versions, failure behavior, and operational budgets.
- Is an agent with fewer approval prompts better? Only when its standing permissions match the intended work and remain enforceable. Prompt count alone is not a quality metric.
- What does background execution add? Independent task lifetime and scheduling. Durable state, cancellation, credentials, cleanup, and duplicate handling become explicit requirements.
- What is wrong with “GUI agents take 1–3 seconds per step”? It is an unsupported universal latency claim. Measure observation, inference, action, application wait, and retries for the chosen environment.
- What would make you reject a high-scoring coding product? Unacceptable data handling, unsupported isolation, excessive review burden, poor recovery, or integration costs on the target workload.
Final summary and practice notes
| Remember | Say in an interview |
|---|---|
| Surface ≠ execution location | “I will separately place the UI, model, controller, and tools.” |
| Schema ≠ authorization | “The server checks who may act on this specific resource.” |
| Sandbox ≠ transaction | “An uncertain external mutation must be reconciled before retry.” |
| Product feature ≠ outcome | “We measure accepted work and its full cost.” |
| Current name ≠ stable contract | “We pin the release and verify the interfaces we depend on.” |
Closing answer: choose the least complex architecture that can execute the required operations with observable outcomes. Explain where authority resides, what fails, how recovery works, and which measurements would justify a different product or execution model.