Agent safety means preventing or limiting harmful outcomes from an agent's decisions and actions. Security protects the system and its data against unauthorized access, manipulation, and disruption. Governance establishes who owns those risks, which controls are required, and how their effectiveness is reviewed. These concerns overlap, but none is satisfied by a model saying that an action is safe.
A support assistant may read the correct account and still issue the wrong refund. It may propose the correct refund but execute it twice after a timeout. It may obey an injected instruction inside a ticket. Each failure needs a different control.
This chapter develops those controls from a concrete threat model through a complete interview design. Workloads and budgets are hypothetical design assumptions. Sources and dated requirements were checked on September 24, 2026.
Start with assets, authority, and failure outcomes
A text-only application can cause serious harm through disclosure or misleading advice. Tools add the ability to affect external systems, often through credentials more powerful than the person reading the response. Memory and background work can extend that influence beyond one conversation.
| Term | Meaning | Support-workflow example |
|---|---|---|
| Asset | Something requiring protection | Customer identity, refund balance, credentials |
| Threat actor | A party that may deliberately cause harm | Customer inserting hostile instructions in a ticket |
| Trust boundary | A transition between differently trusted components or data | Ticket text entering the model; proposal entering the payment adapter |
| Vulnerability | A weakness that can enable harm | Adapter trusts a model-supplied tenant ID |
| Failure mode | How behavior can depart from requirements | A retry creates a second refund |
| Risk | Potential harm assessed using its likelihood and consequences | Unauthorized refunds at the deployment's scale |
| Control | A measure intended to reduce risk | Server-side account binding and atomic operation identity |
| Residual risk | Risk remaining after controls | Provider outages leave some results unresolved |
Do not begin with a universal incident percentage or a fixed list of security products. Identify the workload, sensitive data, permitted effects, adversary access, and recovery obligations. Then choose measurable controls. See prompt injection and agent isolation.
The basic authority boundary
Defense in depth uses multiple complementary controls so that one failure does not remove every protection. Controls should address different failure mechanisms; several classifiers sharing the same blind spot are not fully independent defenses.
Read diagram source
flowchart LR
U[Authenticated request and trusted scope] --> W[Workflow controller]
D[External documents, messages and tool data] --> M[Model proposes next step]
W --> M
M --> P[Policy and business validation]
W --> P
P --> A[Scoped execution adapter]
A --> S[Authoritative service]
S --> R[Recorded result and reconciliation]
R --> W
The model may propose an action; trusted code determines whether that exact action is permitted. An authenticated tool result can still contain attacker-authored text. Authentic transport establishes its source, not the authority of every sentence inside it.
Use a risk taxonomy to find omissions
The OWASP Top 10 for Agentic Applications 2026, published in December 2025, is a useful cross-check. It is distinct from the August 2026 LLM Top 10. The following short descriptions map the agentic IDs to engineering review questions; they are paraphrases, not replacement category names. See the official agentic taxonomy.
| ID | Risk to inspect | Concrete review question |
|---|---|---|
| ASI01 | Redirected objectives | Can external content change the task's authority? |
| ASI02 | Misapplied tools | Can a valid tool create an invalid business effect? |
| ASI03 | Abused identity or privilege | Whose permissions actually authorize execution? |
| ASI04 | Compromised dependencies | Can an extension update acquire new access unnoticed? |
| ASI05 | Unintended code execution | What can generated code reach? |
| ASI06 | Corrupted persistent context | Can an untrusted observation become a durable instruction? |
| ASI07 | Unsafe agent messages | Are sender, scope, freshness, and content checked? |
| ASI08 | Propagating failures | Can one bad result trigger many downstream effects? |
| ASI09 | Misleading human oversight | Does the reviewer see the actual action and evidence? |
| ASI10 | Autonomous deviation | Can runtime controls stop behavior outside the approved task? |
A mapping does not prove coverage. Attach a control, an owner, and a test to the relevant threat. Several categories may share a control, and a single category may require several controls.
Test behavior under pressure without anthropomorphizing it
PropensityBench studies tool choices in simulated high-risk scenarios under varying pressures. Its November 2025 paper reports 5,874 scenarios and 6,648 proxy tools. That is evidence about behavior in a particular experimental setup, not a current production incident rate or proof that a model possesses human intentions. Primary paper.
For a support agent, useful tests include an urgent deadline, repeated tool failure, a claimed executive override, and a nearly exhausted budget. Verify two things separately:
- Does the model propose an inappropriate shortcut?
- Does the execution boundary reject it even when the proposal is fluent and schema-valid?
Never expose a genuinely dangerous production capability merely to see whether the model resists using it. Use simulated tools or isolated test resources. An unavailable capability and a denied operation are stronger boundaries than a request to “be careful.”
Prompt injection: protect the action path, not just the prompt
Prompt injection occurs when adversarial input attempts to redirect model behavior by being treated as instructions. It can arrive directly in a request or indirectly through retrieved documents, email, web pages, screenshots, tool descriptions, or stored memory.
An attacker places “send the customer export to this address” inside a support attachment. The assistant has a legitimate reason to read the attachment. Reading it does not authorize exporting customer data.
Read diagram source
sequenceDiagram
participant X as External ticket
participant M as Model
participant P as Execution policy
participant D as Customer service
X->>M: Ticket text plus injected export request
M->>P: Propose export to an unapproved recipient
P->>P: Check trusted task, data scope and destination
P-->>M: Deny with a bounded reason
Note over P,D: No export reaches the service
| Control | Benefit | Limitation to explain |
|---|---|---|
| Instruction hierarchy and explicit source labels | Helps preserve task intent and provenance | Model behavior is not a guaranteed security boundary |
| Injection detection/classification | Finds some suspicious content | False negatives and false positives remain |
| Restricted tools and arguments | Reduces available actions | A permitted tool can still be misused within broad arguments |
| Current per-resource authorization | Enforces who may read or act | Must use trusted identity and cover every execution path |
| Destination/data policy | Blocks prohibited disclosure paths | Domain-only rules can miss attacker-controlled locations on shared services |
| Verified business postconditions | Detects wrong or missing effects | Detection may occur after a consequential action |
Removing instruction-like strings is not a general solution: ordinary documents can legitimately discuss instructions, and malicious content need not use obvious keywords. Preserve necessary evidence, mark its source, and enforce policy independently of the model. Detection complements these controls.
Tool identity and memory
Bind a tool invocation to a reviewed server identity, namespace, schema, and version. A similar display name must not redirect credentials or silently replace a registered implementation. Discovery metadata is also an input to review, not an automatic grant of access.
Record the origin and scope of persistent memory. An extracted fact is not automatically a trusted instruction. Review sensitive promotions, avoid mixing unrelated users' state, and support correction, deletion, and invalidation of derived artifacts. See memory architecture.
Exfiltration: follow the data to every output
Data exfiltration is an unauthorized transfer of data out of its permitted boundary. It can use an explicit send tool, an HTTP request, a rendered remote image URL, a file upload, logs, or an ordinary response delivered to the wrong person. A read-only database credential does not prevent disclosure.
Data loss prevention (DLP) policies and tools detect or restrict sensitive-data handling and transfer. Pattern matching can miss secrets, encoded content, and context-dependent disclosures; it supplements access and destination controls.
- Retrieve only data the authenticated task needs.
- Keep credentials in trusted adapters, not in prompts or ordinary tool results.
- Restrict outbound destinations, methods, paths, and audiences according to the operation.
- Validate redirects and network destinations at the component making the connection; prevent access to prohibited internal/metadata services.
- Apply appropriate DLP checks and output size limits.
- Monitor unusual destinations and aggregate transfers, with explicit alert ownership.
A proxy cannot inspect encrypted application content merely because traffic passes through it. Decide where inspection occurs and how that component sees permitted plaintext. Avoid recording the sensitive material being detected.
Read/write separation: a read worker can prepare a narrow, typed proposal for a send service. This reduces credential exposure. It does not help if the send service blindly forwards any text, file, or URL supplied by the reader. Validate the handoff against audience and data policy.
Tradeoff: narrow adapters take more development effort than unrestricted browsing or shell access, but make permissions and auditing tractable. For public web research, broad internet retrieval may be required; place it in a separate environment without private credentials and control what crosses back.
Wrong tools and cascading failures
A tool call can be valid JSON and still target the wrong customer, currency, record version, or operation. A downstream agent can repeat a false fact in equally valid JSON. Schema validation checks structure, not truth or authorization.
| Failure | Useful control | Why another common control is insufficient |
|---|---|---|
| Wrong tool chosen | Clear contracts, narrowed registry, operation policy | A familiar tool name does not establish intent |
| Wrong account | Trusted identity and current resource authorization | A model-generated account ID is only a proposal |
| Plausible false value | Source references, domain invariants, reconciliation | Type/range checks may accept it |
| Repeated effect | Scoped operation key and authoritative deduplication | A message queue can redeliver |
| Multi-agent amplification | Typed scoped handoffs, fan-out limits, independent verification | A signature proves origin, not correctness |
| Failing dependency | Bounded retries, circuit breaker, fallback/queue | Repeating a non-idempotent write may make the problem worse |
A circuit breaker temporarily stops calls to a dependency according to a failure policy. It is not a complete incident response system. A dead-letter queue isolates messages for inspection; it needs capacity, ownership, access control, and a replay procedure that preserves operation identity.
Choose isolation from the threat model
A sandbox restricts what executed code can access or affect. A browser profile, Python virtual environment, and operating-system isolation boundary solve different problems. There is no universal ordering that makes every WebAssembly deployment safer than every VM.
| Mechanism | Main boundary | Suitable workload | Remaining design work |
|---|---|---|---|
| Hardened ordinary container | Namespaces, resource controls, privileges; shared host kernel | Managed workloads where the assessed boundary is acceptable | Patch host, restrict capabilities/mounts/network; assess hostile-code exposure |
| gVisor | Interposed userspace system-call implementation | Compatible Linux workloads needing reduced host-kernel exposure | Test compatibility/performance; configure network and host resource limits |
| WebAssembly runtime | Module memory/control flow and explicitly supplied host interfaces | Supported languages/modules with bounded host capabilities | Audit host functions, configure filesystem/network capabilities and execution limits |
| MicroVM | Virtualization boundary with guest kernel | Hostile code needing a compatible guest OS environment | Secure host/VMM, constrain devices and credentials, patch and limit resources |
| Dedicated host/account | Additional infrastructure and administrative separation | Sensitive workloads needing stronger tenant separation | Higher cost and operations; application authorization still applies |
gVisor's security model explicitly distinguishes kernel-exposure reduction from network, resource, and hardware concerns. Wasmtime exposes outside capabilities through host interfaces; WebAssembly is not limited to pure computation when suitable interfaces are deliberately granted. Firecracker's production guidance requires host hardening and jailer-equivalent process constraints.
Measure startup, memory, and syscall-heavy behavior on the actual image and workload. A demonstration's boot time is not a service-wide cold-start SLO, and VMM overhead is not the guest's total memory.
Interview choice: for a multi-tenant service executing arbitrary native code, begin with a hardened microVM or appropriately evaluated gVisor deployment, scoped credentials, and controlled egress. Explain the compatibility and cost tradeoff. Do not treat a product name or a default container as proof of isolation.
Authorization: identity, operation, resource, conditions
Least privilege grants only the access necessary for a task. Default deny rejects operations not explicitly permitted. Capability-based security conveys authority through possession of appropriately protected, restricted capabilities; it is not simply a synonym for a tool allowlist or any role check.
A permission decision should bind:
- Subject: authenticated user/service and tenant.
- Operation: the actual action, not only a broad tool name.
- Resource: account, record, workspace, and destination.
- Conditions: amount, scope, version, expiry, approval, and current policy.
Use fixed business operations or safely parameterized queries where practical. A read-only SQL user can still disclose too many rows or run expensive queries. Apply resource authorization, query controls, result limits, and timeouts. Database row-level security also depends on connection role and policy configuration; an administrator or bypass-capable role defeats the intended boundary.
Risk tiers are a product policy, not a universal table
| Example operation | Possible starting policy | Condition that changes the answer |
|---|---|---|
| Public product lookup | Automatic | Result contains private customer data |
| Own-account status | Automatic after authorization | Identity/session is uncertain |
| Support reply draft | Automatic preparation | Sending introduces audience and content obligations |
| Small eligible refund | Automatic only within explicitly approved business policy | Account, amount, eligibility, or duplicate state changes |
| Exceptional refund | Named reviewer, sometimes separation of duties | High amount or regulatory/business constraint |
| Credential/permission change | Dedicated identity workflow | A conversation cannot substitute for required verification |
A read may be highly sensitive; a routine write may be safely automated. Blanket “all reads allowed, all writes approved by a human” creates both security gaps and unnecessary review. Some operations should be unavailable to the agent altogether.
Make approval specific, current, and reviewable
Human-in-the-loop approval places a human decision at a defined workflow point. Approval authorizes a particular proposal; it does not replace identity checks, business constraints, or execution verification.
A useful review card shows:
- Action and affected account/record, using independently resolved identity.
- Exact amount, currency, destination, recipients, or diff.
- Source evidence and any uncertainty or missing data.
- Expected effect, reversibility, alternatives, and deadline.
- Policy reason requiring review.
- The version being approved and an explicit approve/reject action.
A model explanation can help, if labeled as such and accompanied by inspectable evidence. Do not use it as the sole proof of safety. A 100% approval rate is not automatically a failed gate; sample decision accuracy, missed errors, reviewer workload, and time spent inspecting relevant evidence.
Approval and execution contract
| Field | Purpose |
|---|---|
| Proposal ID and canonical payload hash | Binds review to the exact normalized action |
| Tenant, subject and resource | Prevents approval reuse across identities/records |
| Resource/policy version | Detects changed eligibility or permissions |
| Reviewer and decision | Records authorized oversight |
| Expiry | Prevents indefinite reuse |
| Consumption/operation ID | Binds one approval to the logical effect |
At commit, revalidate current authority and preconditions and atomically reserve/consume the approval with the operation when they share a transaction boundary. Hash a stable canonical representation; an arbitrary JSON serialization is not a universal canonical format. If anything material changes, prepare a new proposal and obtain the required review.
Remote gap: a local transaction cannot atomically commit a remote refund by itself. Use the provider's documented idempotency contract, retain unknown outcomes, and reconcile. If the remote resource can change after local validation, use supported conditional writes or enforce the invariant in the authoritative service.
A review deadline may expire the proposal or escalate it according to policy. Expiry must never silently become approval. Do not impose an arbitrary 24-hour deadline on every workflow.
Bound work and spending across the whole task
Rate limits restrict activity over time. Concurrency limits restrict simultaneous work. Quotas/budgets restrict aggregate consumption over a specified scope and period. They are related but not interchangeable.
| Limit | Example assumption | Enforcement location |
|---|---|---|
| Per-tenant tool rate | 30 calls/minute with a bounded burst | Shared admission service |
| Task attempt budget | At most 12 tool attempts, including retries | Durable workflow record |
| In-flight work | At most four active tool calls/task | Shared task reservation |
| Model token budget | Separate input/output and context limits | Model gateway and task ledger |
| Spend budget | At most $0.50 in reservable task charges | Cost reservation and settlement |
| Query/output bound | 100 rows or 256 KiB per result | Adapter before materialization/return |
| Wall-clock deadline | Ten minutes from accepted task | Controller and execution adapter |
| External effects | Refund amount/count per account and period | Authoritative business service |
These numbers are examples, not recommended defaults for all deployments. Child agents and retries share the parent budget; creating a new worker must not reset it.
Overshoot example: a task has $0.30 remaining. Four workers each read that balance and independently start a call costing up to $0.10. They can incur $0.40. Atomic reservations admit only three such calls. Settle actual charges and release unused reservations; handle expired workers and delayed billing without incorrectly returning spent capacity.
A hard monetary bound requires an enforceable upper bound on each admitted operation. Unknown vendor charges need a conservative reservation or a clearly described soft budget. Cancellation does not refund work already billed.
Validate structure, meaning, authority, and effect separately
- Structure: expected types, required fields, lengths, additional fields, and formats.
- Meaning: currency, totals, supported states, resource identity, and domain invariants.
- Authority: current actor, resource permissions, policy, and approval.
- Output handling: safe rendering, sensitive-data restrictions, destination control.
- Effect: authoritative result ID, version, or independently checked postcondition.
Structured generation reduces some malformed outputs; it does not establish truth or permission. A second model can assist classification but can share the first model's errors. Deterministic checks should enforce rules that have deterministic definitions. See building tools.
Render untrusted HTML/Markdown using appropriate sanitization, safe links, and controlled remote resources. Content sanitization for browser execution is a distinct problem from natural-language prompt injection. Do not evaluate a returned string as shell, SQL, JavaScript, or a template merely because it passed a text filter.
Keep useful evidence without creating a second sensitive-data store
An audit trail records attributable events needed to reconstruct relevant decisions and effects. It does not require access to a model's hidden chain of thought. A generated rationale is not a faithful substitute for execution evidence.
| Event | Useful fields | Data to minimize or protect |
|---|---|---|
| Request accepted | Task, actor, tenant, purpose, timestamp | Full private conversation |
| Proposal prepared | Operation, normalized payload reference/hash, resource version | Credentials, unrestricted document text |
| Policy decision | Decision, rule/version, reason code, resource scope | Sensitive values unnecessary to explain the rule |
| Review | Reviewer, exact proposal, decision, time | Unrelated personal details |
| Execution | Operation key, attempt, adapter/version, remote receipt/state | Bearer tokens and unredacted response bodies |
| Reconciliation | Observed effect, conflict, repair decision | Excessive copies of source records |
| Budget/incident | Reserved/settled units, stop event, unresolved scope | Secret-bearing stack traces |
Protect evidence with restricted access, encryption, reliable collection, integrity controls, and an explicit retention/deletion policy. Append-only storage and separate administration can make tampering harder to conceal. They do not mean every employee must be technically unable to delete any record forever.
A transactional event/outbox can bind a business action to its audit event. An asynchronous debug buffer alone can lose the final event on crash. Define which actions must stop when required audit persistence is unavailable; keep a bounded degraded policy for other telemetry.
Retention correction: there is no universal “finance seven years, healthcare six years” rule for all agent logs. HHS distinguishes required Security Rule documentation retention from medical-record retention. Determine the record category, entity, jurisdiction, contractual duties, legal holds, and data-minimization obligations. HHS Security Rule summary, HHS medical-record retention FAQ.
Sizing: 100,000 tasks/day × 12 events/task × 1 KiB/event is about 1.23 GB/day in decimal units, or 36.9 GB over 30 days, before indexes, replication, attachments, and backups. Screenshots or full documents can dominate this budget; retain only what the evidence policy requires.
Stop authority, then account for work already in flight
A kill switch is an independently enforceable mechanism to stop specified activity. A chat message saying “stop” is not enough when workers have queued actions or hold credentials.
| Scope | Mechanism | Important limitation |
|---|---|---|
| Task | Durable cancellation state and adapter checks | An already submitted remote action may complete |
| Tool/agent/tenant | Deny new admissions and invalidate permits | Disconnected workers may hold stale state |
| Runtime | Terminate or isolate worker environments | Process termination is not transaction rollback |
| Credential/network | Revoke credentials, restrict egress, disable adapter | Revocation propagation and token caching vary |
| Fleet | Independent operator control with incident runbook | Recovery and service continuity still need owners |
For urgent containment, revoke admission first and choose whether existing work should drain or stop immediately based on the threat. Requiring two people for every emergency halt can delay containment; separation of duties may be more appropriate for restoring privileges or high-impact configuration changes.
Read diagram source
sequenceDiagram
participant O as Operator control
participant C as Admission and policy service
participant W as Worker
participant R as Remote service
O->>C: Disable operation scope and advance policy generation
C-->>W: Cancel or invalidate lease
W->>C: Check authorization before next effect
C-->>W: Denied
Note over W,R: Earlier submitted request may still finish
R-->>C: Receipt or later reconciled outcome
C->>C: Record affected operations for recovery
An independent adapter should enforce current cancellation/policy, including for a worker that ignores a notification. Fence stale writers at the authoritative service where supported. A network partition may prevent an immediate cancellation signal; short-lived permits and a defined fail-closed policy bound further access, but do not reverse completed effects.
Set and test containment targets for the actual control path. Distinguish “no new effects admitted,” “worker stopped,” and “all remote outcomes reconciled.” Do not promise universal one-second shutdown and complete reversal.
Governance is ownership plus repeatable evidence
NIST AI RMF organizes work around Govern, Map, Measure, and Manage. It is voluntary guidance, not a certification or an ordered four-step checklist. As of this review, NIST says RMF 1.0 is being revised. NIST framework, RMF Core.
| Governance question | Concrete record or practice |
|---|---|
| Who owns this deployment? | Business owner, technical owner, incident contact |
| What is it allowed to do? | Purpose, users, data classes, tools, effects, exclusions |
| Which dependencies matter? | Model, framework, extension, adapter and policy versions |
| What evidence permits launch? | Quality/security gates, residual-risk decision, recovery drill |
| What requires another review? | New privilege, data source, vendor, model, material incident |
| How is it monitored? | Verified outcomes, incidents, false positives, human workload |
| How is it retired? | Disable authority, migrate users, apply retention/deletion rules |
Keep an agent registry, permission inventory, dependency review, and named incident ownership. Vendor frameworks can help map controls to a platform; their number of listed risks or controls does not prove that your deployment is covered.
The OWASP project's Agent Control Standard, added to its resources in September 2026, describes middleware hooks for portable runtime policies. Evaluate implementation coverage and bypass paths in the framework you actually use; the existence of an interface is not enforcement evidence. ACS resource.
A practical maturity progression is documented ownership → enforced boundaries → measured operation → tested change and recovery. Treat this as a planning aid, not a claim about the percentage of organizations at each level. Policy changes themselves need accountable review; a model must not quietly expand its own permissions because a metric improved.
Test the boundary and the recovery path
| Test | Inject | Required observation |
|---|---|---|
| Indirect injection | Hostile text in a normal ticket/attachment | Prohibited effect denied even if proposed |
| Tenant isolation | Valid record ID from another tenant | No data or existence disclosure beyond policy |
| Stale approval | Change amount, account, policy, or source version | Commit rejected or reviewed again |
| Duplicate execution | Redeliver the same approved logical operation | One effect or controlled reconciliation |
| Unknown outcome | Timeout after the remote service commits | No blind new-key retry |
| Memory poisoning | Save an adversarial extracted “rule” | No promotion to execution policy |
| Tool substitution | Change server identity/schema/version | Review/deny according to registry policy |
| Cascading falsehood | Well-formed but wrong upstream value | Domain evidence catches it or routes uncertainty |
| Budget race | Several workers reserve remaining budget | Aggregate bound maintained |
| Emergency halt | Stop while queued, executing, and disconnected | Admission denied and unresolved effects enumerated |
| Audit outage | Fail evidence persistence before an effect | Required action policy enforced without silent loss |
| Recovery | Restore state and replay pending work | No permission widening or duplicate effects |
Use representative legitimate tasks too: an aggressive filter that blocks most valid work may produce excellent attack-blocking numbers and an unusable product. Track false positives, quality, cost, reviewer minutes, and failure severity. Repeat trials for variable model behavior and separate model proposal rate from actual unauthorized-effect rate.
Statistical reminder: zero failures in 300 independent representative trials gives a one-sided 95% upper bound of about 1%, not proof of zero risk. Adaptive attackers and correlated test cases violate the simple independent-trial interpretation. Use the bound as one piece of evidence, alongside structural control tests and adversarial work.
Current legal scope: classify the use, not the word “agent”
Agent autonomy does not automatically make a system high-risk under the EU AI Act. Assess the intended use, relevant category, actor obligations, exceptions, and other applicable laws. Disclosure duties and high-risk duties are distinct; disclosure is not a universal right to the model's hidden reasoning.
The Commission's current FAQ gives December 2, 2027 for Annex III high-risk rules and August 2, 2028 for high-risk AI embedded in regulated products. Transparency duties generally applied from August 2, 2026, with a specific December 2, 2026 transition for Article 50(2) marking/detection for systems already marketed before August 2. This is not a blanket delay of interaction disclosure. Commission enforcement FAQ.
For an actual launch, record applicability, approved data use, required oversight, notices, incident obligations, and evidence retention with qualified owners. SOC 2, HIPAA, PCI DSS, privacy law, and the AI Act have different scopes; citing all of them does not establish compliance. Continue in AI governance and compliance.
Interview design: support actions across many customer accounts
Assumptions: 40,000 support tasks/month. The first release answers authorized order questions, prepares replies, and issues eligible refunds through a payment service. Some refunds require human review. We design the controls around the actual business policy, not around a blanket ban on useful writes.
Functional requirements
- Authenticate the user/operator and bind each task to its permitted customer resources.
- Retrieve scoped order, ticket, and payment evidence.
- Prepare a versioned reply or refund proposal.
- Enforce eligibility, amount, destination, approval, and duplicate rules.
- Execute through scoped adapters and record authoritative outcomes.
- Route exceptions and reconcile unknown results.
- Support operator containment, audit retrieval, correction, and controlled recovery.
Non-functional requirements
- Isolation: prevent cross-customer disclosure or effects.
- Correctness: enforce business invariants even after injection or model error.
- Durability: preserve approvals, operation identity, and unresolved states across crashes.
- Responsiveness: set separate targets for automatic responses and review-dependent work.
- Containment: measure how quickly each adapter stops new admissions after revocation.
- Economics: bound resource use and account for human review and recovery.
- Privacy: minimize data passed to models and retained in evidence stores.
Basic design and first flaw
Chat UI → model → CRM/payment tools. A system prompt asks the model to verify identity and avoid duplicate refunds.
Flaw: the same probabilistic component interprets a hostile ticket and decides whether it is authorized. A timeout can also lead it to submit a second refund.
Repair: move identity, eligibility, approval binding, and operation identity into trusted services. The model prepares proposals and explains results; it does not supply its own trusted role or decide whether a prior effect exists.
Detailed design
Read diagram source
flowchart TD
U[Authenticated support interface] --> C[Durable task controller and trusted scope]
C --> M[Model with bounded context and tool registry]
M <--> R[Scoped evidence adapters]
R <--> D[CRM, orders and payment records]
M --> P[Normalized versioned proposal]
P --> V[Policy, eligibility and current resource checks]
V -->|Review required| H[Reviewer sees exact action and evidence]
H --> K[Commit coordinator]
V -->|Explicit automatic policy| K
K --> L[Approval, operation and audit transaction]
L --> A[Scoped payment or messaging adapter]
A --> E[External service with supported operation identity]
E --> O[Receipt, rejection or unknown outcome]
O --> Q[Reconciliation and exception queue]
Q --> C
S[Independent containment control] --> V
S --> A
L --> T[Protected event store and monitoring]
Q --> T
The commit coordinator atomically records the local operation and necessary evidence, then dispatches through a controlled adapter. Remote atomicity depends on the service contract. Worker termination after dispatch may leave unknown, which the reconciliation queue must resolve.
Flaws found during review
| New flaw | Repair | Cost/benefit |
|---|---|---|
| Reviewer approves one amount, proposal changes | Immutable normalized proposal plus current preconditions | Extra versioning; prevents changed-action approval reuse |
| Two workers retry after a timeout | Shared operation ledger and supported remote idempotency | Storage and reconciliation work; prevents common duplicate effects |
| Malicious attachment targets shared cloud URL | Data/audience policy and controlled retrieval/rendering | Integration effort; stronger than domain-only filtering |
| Stop message sits behind work | Independent admission control and adapter revocation | Extra control path; effective even when the model ignores stop |
| Reviewer queue becomes the bottleneck | Narrow eligible automatic policy and measured exception routing | Requires evidence; avoids unmanageable blanket approval |
| Model/provider update changes behavior | Versioned evaluation, canary, rollback and re-review triggers | Release overhead; detects regressions before broad rollout |
Capacity and full cost
Assume 40,000 tasks across 22 eight-hour workdays, a 5× sustained peak, and a mean of 18 seconds of active runtime per task across the workload, excluding human-review waiting. Peak arrival is about 0.316 tasks/second, giving 5.7 active tasks on average at that peak. At 65% planned occupancy, start with nine slots, then load-test variability and dependency quotas. Persist human-review waits instead of holding an active model worker for hours.
If 12% of tasks require three review minutes, that is 240 reviewer hours/month. At 120 productive hours/person/month, it needs two people of capacity before absence and peak coverage. At 30% review, it needs 600 hours, or five people on the same assumption.
| Monthly cost assumption | 12% review | 30% review |
|---|---|---|
| Model/tools at $0.07/task | $2,800 | $2,800 |
| Runtime, storage, monitoring | $900 | $900 |
| 20 maintenance hours at $100 | $2,000 | $2,000 |
| Review at $45/hour | $10,800 | $27,000 |
| Total | $16,500 | $32,700 |
| Cost per attempted task | $0.41 | $0.82 |
At 95% verified completion, the corresponding costs per completion are about $0.43 and $0.86. Initial development, exceptional incident work, and any separate service fees are excluded and must be added when applicable. Do not invent a percentage reduction in security incidents to justify a control; measure the relevant quality and failure outcomes.
Closing remarks
“I would launch with scoped reads and a narrow set of policy-approved support actions. Every effect passes current authorization and domain checks; review binds to exact proposals; remote uncertainty is reconciled. The security boundary remains effective when the model makes a bad proposal. I would expand autonomy only after testing tenant isolation, duplicate prevention, containment, reviewer capacity, and full cost on representative work.”
Use incident evidence carefully
A production incident, a researcher demonstration, and a simulated benchmark answer different questions. Do not present an unattributed story as a measured deployment result or infer a root cause merely from a social-media description.
- Replit's database-separation announcement supports separating development and production data/authority; it does not prove a universal cause or scope for every reported incident.
- Cisco's discussion of third-party skills includes research involving a deliberately malicious extension. Treat that as evidence of an attack mechanism, not a claim that every extension installation was compromised.
- Memory-poisoning research motivates testing persistent write channels. Check its systems and experimental conditions before transferring any measured rate to your own design.
An actionable postmortem records the verified timeline, affected operations/data, contributing control gaps, containment evidence, recovery, and tests that would detect recurrence. A proposed fix should address the demonstrated mechanism. For example, adding schema validation cannot by itself fix a plausible but false value that already passes the schema.
Interview questions and answer notes
- Can a schema-valid call be unsafe? Yes: identity, policy, business meaning, and current state are separate checks.
- Does read-only mean low risk? No. Private reads can cause disclosure and resource exhaustion.
- Does separating reader and writer agents stop exfiltration? Only if their handoff and the writer's authority are actually constrained.
- Does prompt sanitization prevent every injection? No. Use detection as one layer and enforce action policy independently.
- What is wrong with a broad database tool plus an allowlist of tool names? The allowed tool may still reach unauthorized records or expensive operations.
- Does a microVM protect a production API credential deliberately placed inside it? It constrains host access, not the legitimate powers of that credential.
- What makes an approval stale? Changed payload, resource state, policy, permissions, or expiry.
- Can an independent kill switch guarantee an already submitted refund is canceled? No. Track and reconcile the remote outcome.
- Why reserve budget atomically? Concurrent workers otherwise each spend the same remaining balance.
- Should all model reasoning be stored for compliance? No. Retain required attributable evidence with appropriate minimization and retention.
- Does zero observed failure prove safety? No. State sampling assumptions, uncertainty, and structural-control evidence.
- Does using an agent automatically trigger EU high-risk classification? No. Intended use and applicable legal criteria determine obligations.
- What makes governance operational? Named ownership, enforced scope, release evidence, monitoring, and tested response/recovery.
- What is the strongest interview demonstration? Trace a concrete malicious or mistaken proposal through the control that prevents or contains its effect.
Final notes
Remember scope → evidence → authorization → execution → verification → recovery. Use this as a recall aid, not a new security standard. Explain each control through the failure it addresses and its remaining limits. Useful autonomy comes from a bounded task and trustworthy execution, supported by people who can inspect, stop, and repair the system.
Next: Real-time voice agents.