This is a hypothetical interview scenario. Volumes, timings, error rates and costs are planning assumptions, not measured production results. Model and platform facts have primary-source links.
Interview focus: automate supported legacy workflows while keeping business authority, exact approvals, uncertain outcomes and human capacity explicit.
60-second interview answer
I would first separate stable API or scripted steps from the parts that genuinely require interpreting a screen. The agent would work in a fresh restricted session, read the approved task and propose bounded business actions. A trusted controller would verify the company, report, fields and current state before any change. Required approval would cover the exact proposed action and expire after relevant changes. The executor would record intent, perform the action and verify the remote outcome. Unknown outcomes would require reconciliation before retry. I would measure complete workflow quality, exception staffing and full cost, not just clicks or model success.
Remember: Observe → Propose → Authorize → Record → Act → Verify or reconcile.
Interview problem and scope
A 4,000-person company processes 14,000 expense reports weekly using a corporate-card portal without a usable API, an expense tool with unreliable CSV import and Workday for cost-center lookup. Each report involves four to seven fields. The business estimates nine manual minutes per report, a 2.3% correction rate and $80,000 in quarterly cleanup.
Three contractors assist today, but they cannot represent the full manual workload: 14,000 × nine minutes is 2,100 hours/week. At 40 hours each, three people could handle about 800 reports. Clarify the share already automated and work performed by other employees before claiming staff savings.
For this scenario, company policy requires separate authorized approval of payments over $2,500. This is an internal threshold, not a universal SOX rule. Bank-detail changes, record deletion and policy administration are outside the agent's authority. Start with receipt reconciliation and draft preparation; extend submission only where the target system and controller can enforce the necessary boundary.
Functional requirements
- Accept a task with confirmed company, employee, report identity, receipt evidence and permitted operations.
- Read required data from the three systems using approved credentials and task-scoped sessions.
- Propose field changes and attachments with evidence; detect mismatched totals, currencies or cost centers.
- Obtain applicable approval for the exact change or submission and reject stale approvals.
- Execute permitted actions, verify target state and reconcile uncertain outcomes before retrying.
- Support pause, cancellation, human takeover and a durable audit of proposals, approvals and results.
Nonfunctional requirements
- Authority: every write passes a trusted enforcement boundary; model wording cannot grant permission.
- Isolation: separate browser profile, filesystem and credentials between tasks; apply destination and resource restrictions.
- Completion: target 84% of eligible reports prepared without technical repair. Required human approvals still count separately.
- Latency: target machine p95 under 12 minutes, with a separate example four-hour high-risk approval queue target.
- Cost: target a $0.30 average machine allowance per report and explicitly budget exceptions, retries and people.
- Recovery and audit: persist required action intent before execution, record verified outcomes and resolve gaps without duplicate business effects.
Workload, workers and human capacity
Assume all 14,000 weekly reports are candidates for the scoped workflow and six minutes of active machine time per report. Arrivals are concentrated into 40 working hours.
| Quantity | Calculation | Implication |
|---|---|---|
| Mean worker occupancy | 14,000 × 6 / 60 / 40 = 35 workers | Waiting for human approval should release expensive active capacity |
| Fivefold machine burst | 35 × 5 = 175 workers | With 25% headroom, plan about 219 worker slots and validate tail behavior |
| Technical exceptions | 14,000 × 16% = 2,240 reports/week | Five minutes each requires 186.7 review hours |
| Routine confirmation | 14,000 × 30 seconds | 116.7 hours/week, even when drafting succeeds |
| Additional high-risk approval | 10% × 14,000 × two minutes | Another 46.7 hours/week |
| Total modeled human work | 186.7 + 116.7 + 46.7 ≈ 350 hours/week | At 30 productive review hours/person, at least 12 reviewer equivalents |
These categories deliberately count routine confirmation and additional specialist work separately; measure overlap before final staffing. At 15% quarterly growth, 14,000 × 1.15⁴ is approximately 24,486 reports/week after four quarters. Capacity planning must include growth, end-of-month bursts and vendor throttling.
Baseline and implementation choices
Start with supported APIs and deterministic scripts for stable lookups. Use a computer-use model only for interface interpretation where ordinary integration is unavailable or demonstrably costly to maintain. Keep final payment authorization in the established human workflow until semantic action controls are validated.
| Layer | Candidate | What it must actually provide |
|---|---|---|
| Workspace | Firecracker microVM or qualified managed sandbox | Isolation, lifecycle, resource limits, cleanup and restricted credentials |
| Browser/controller | Playwright/Chromium or a desktop adapter | Fresh observation, bounded actions and verified target state |
| Model | Evaluated Sonnet 5 or another supported computer-use model | Proposals under the selected tool/schema version |
| Authorization | Application policy, credential broker and trusted executor | Company, record, operation and approval enforcement |
| Orchestration | Durable jobs, leases and action records | Recovery without concurrent duplicate execution |
| Evidence | Restricted event and object storage | Traceable intent/result, retention and access controls |
Firecracker is a KVM-based microVM technology. Its advertised minimal startup characteristics are not a ready, authenticated browser latency promise. Account for image preparation, network setup, browser startup, login and vendor response time. A managed service can reduce operational work but still needs equivalent isolation and recovery checks. Firecracker.
Current integrations differ. OpenAI documents code-execution and structured computer-action paths; the former is recommended for its general integration. Claude's current computer-use toolset and earlier tool versions have model/platform compatibility differences, so pin the supported combination. A broad code-execution recommendation is not permission to bypass this business workflow's controls. OpenAI computer use, Claude computer-use tooling.
Detailed architecture
Read diagram source
flowchart TB
USER[Authorized expense task] --> INTAKE[Validate identity evidence and task scope]
INTAKE --> QUEUE[(Durable task queue)]
QUEUE --> ORCH[Scheduler and worker lease]
ORCH --> VM[Fresh restricted browser workspace]
VM --> OBS[Observe report and current account]
OBS --> MODEL[Model proposes business action]
MODEL --> POLICY[Trusted scope and state validation]
POLICY --> DEC{Permitted and sufficiently verified?}
DEC -->|No| HANDOFF[Stop or hand off with evidence]
DEC -->|Read or standing authorization| INTENT[(Durable action intent)]
DEC -->|Approval required| REVIEW[Review exact diff or submission]
REVIEW -->|Approved current revision| INTENT
INTENT --> EXEC[Trusted executor rechecks and dispatches]
EXEC --> RESULT[Read authoritative result and save receipt]
RESULT --> KNOWN{Outcome known?}
KNOWN -->|Success and work remains| OBS
KNOWN -->|Complete| CLOSE[Revoke session and clean workspace]
KNOWN -->|Unknown| RECON[Reconcile target state before any retry]
RECON --> HANDOFF
“Trusted executor” means the model cannot use a second route around it. If the model has arbitrary shell/CDP access to a privileged browser, a narrow submit_report tool does not constrain its capabilities. Keep unrestricted code execution out of that session, restrict available operations and prefer server-enforced application roles. Domain allowlists alone do not distinguish tenants sharing a hostname.
Some legacy UIs cannot provide a reliable, independently checked business-action boundary. In those cases the design remains draft-only and a person submits in the destination application. Do not claim a screenshot classifier can guarantee every generic click is harmless.
APIs and records
| Operation | Contract |
|---|---|
POST /expense-tasks |
Validated company/report identity, scope and evidence references; deduplicate submission retries |
POST /tasks/{id}/proposals |
Controller builds canonical business payload and digest from validated values |
POST /proposals/{id}/approvals |
Authorized reviewer approves the digest, policy revision and expiry |
POST /actions/{id}/dispatch |
Atomically reserve the approved action under the current task lease |
POST /tasks/{id}/takeover |
Fence model actions and revoke/stop the worker session before transferring control |
| Record | Fields and purpose |
|---|---|
| Task | Company, report, employee, allowed operations, evidence and budget |
| Observation | Account/report identity, remote version or verified state fingerprint, time and source locations |
| Proposal | Typed operation, amount in minor units/currency, recipient, field diff, attachments and digest |
| Approval | Reviewer/role, proposal digest, policy revision, expiry and consumption state |
| Action intent | Stable action ID, task lease epoch, approved bytes, dispatch state and attempt history |
| Outcome | Remote receipt/status, observed values, known/unknown state and reconciliation evidence |
A state fingerprint is useful only if it includes the fields relevant to the action and comes from a reliable observation. If the UI lacks a trustworthy revision token, re-read and compare the important values immediately before dispatch. This narrows the race; it does not create atomicity that the remote application does not offer.
Isolation and identity
A disposable workspace prevents one task inheriting another's cookies, downloads or clipboard. Restrict outbound destinations, downloads, uploads, CPU, memory, runtime and session duration. Revoke credentials and remove temporary data after completion or cancellation; a paused VM may retain its filesystem and memory.
Isolation does not prevent authorized network effects. A sandbox with a valid payment session can still send money. Use read-only or draft-only application roles where possible, and separate the submission authority from the model-controlled workspace.
The broker obtains credentials the destination actually accepts. Minting an internal JWT does not make Workday or a card portal understand it. OAuth resource indicators can help obtain tokens for a particular resource when supported, but audience binding is neither tenant authorization nor replay prevention at that resource. RFC 8707.
Store long-lived secrets outside model context and task logs. Where a legacy application requires a browser session, scope its account privileges and isolate its cookies. Expiry and logout behavior must be tested against that application. A worker lease cannot revoke a payment already committed remotely.
Approval belongs to a specific action
| Action class | Policy in this scenario | Reviewer sees |
|---|---|---|
| Permitted read/navigation | Standing task authorization | Recorded account/report scope |
| Draft fields and receipt attachment | Bundled report-level confirmation | Before/after values, currency, receipt and supporting evidence |
| Permitted payment/report submission | Required role approval; separate control above $2,500 | Exact company, report, recipient, amount, operation and current state |
| Bank-detail changes, deletion or policy administration | Outside agent scope | Handoff to established procedure |
Bundling the four-to-seven field edits of one report reduces repeated confirmations while preserving the full diff. A changed recipient, amount, report, policy or source evidence invalidates the relevant approval. Do not teach a four-second universal approval time or assume people will reliably examine dozens of nearly identical dialogs.
This pure gate illustrates checks over validated controller records. The digest is computed outside the model over the complete canonical action and relevant scope. Approval and dispatch still need atomic reservation in the orchestration database.
def action_gate(proposal, observed, grant, now):
for field in ("tenant_id", "report_id", "state_version", "policy_version", "lease_epoch"):
if proposal[field] != observed[field]:
return "STALE_OR_WRONG_SCOPE"
if not observed["policy_permits_action"]:
return "DENIED"
if proposal["kind"] == "read":
return "READY_FOR_DISPATCH"
if not grant or grant["consumed"] or now >= grant["expires_at"]:
return "APPROVAL_REQUIRED"
if grant["action_digest"] != proposal["action_digest"]:
return "APPROVAL_REQUIRED"
return "READY_FOR_DISPATCH"
The policy engine determines whether an operation is a read; the model cannot relabel “click Submit” as read-only. The grant comes from an authenticated reviewer with the required current role. The executor verifies these facts again when reserving dispatch and consumes the grant once. Its local action ID links records; it provides remote deduplication only if the destination supports that contract.
Prompt injection and allowed operations
A receipt may contain instructions to change a bank account or ignore the stated total. Indirect prompt injection occurs when such untrusted material influences the agent as instructions. Keep the task and policy in trusted controller state; receipt text supplies evidence only.
Use independent destination, account, field and operation restrictions. Detectors, captions and trust labels can aid review but do not enforce authority. CaMeL studies a separation of control/data and capability-based enforcement; adding a warning tag or a second vision model does not implement that system. CaMeL research.
Prefer an allowlist of business operations such as read_report, prepare_field_diff and attach_approved_receipt. A list of generic click/type/scroll tools is not itself a business allowlist: those tools can express almost any action available to the logged-in user. Downloads and uploads need object-level restrictions as well as allowed hosts.
Action state and uncertain outcomes
Read diagram source
stateDiagram-v2
[*] --> Proposed
Proposed --> Rejected: Scope or policy fails
Proposed --> Ready: Permitted read under standing authorization
Proposed --> AwaitingApproval: Permitted change
AwaitingApproval --> Ready: Exact current action approved
AwaitingApproval --> Expired: Approval deadline passes
Ready --> Proposed: Relevant state changes
Ready --> IntentRecorded: Atomic dispatch reservation
IntentRecorded --> Executing: Current executor dispatches
Executing --> Succeeded: Target outcome confirmed
Executing --> Failed: Confirmed not applied
Executing --> Unknown: Timeout or connection loss
Unknown --> Reconcile: Inspect remote state
Reconcile --> Succeeded: Matching result found
Reconcile --> Failed: Conclusively not applied
Reconcile --> HumanReview: Outcome cannot be established
Failed --> Proposed: New bounded attempt
Succeeded --> [*]
After a submission timeout, look up the payment or report using the available remote identifiers and recorded business details. If the target cannot establish whether it accepted the operation, stop for manual reconciliation. A local task retry, screenshot replay or click on the same coordinates cannot make that action exactly once.
Cancellation stops new work but cannot undo a committed payment. Human takeover must fence dispatch, stop/revoke the old worker's session and reconcile any in-flight action before another executor continues. Losing heartbeats alone does not prove the old worker stopped acting.
Audit integrity and completeness
Persist required intent before making a change, then record the observed outcome. A hash chain can make alteration detectable only relative to a protected trusted anchor. An attacker able to rewrite an entire local chain can recompute it; store signed checkpoints or roots separately and protect the signing keys.
An intact chain does not prove completeness. A missing remote action can leave the remaining records internally consistent. Reconcile intent, dispatch records, application receipts and expected sequence transitions. Screenshots support review but are not a financial ledger or a promise of byte-for-byte reproduction of a live website.
Capture evidence deliberately. Recording one 200 KB frame per second for six minutes is 72 MB/report, or about 1 TB/week at this volume. Seven years of that hypothetical stream is roughly 367 TB before redundant storage. Event-based frames, redaction and scoped retention can materially change the cost; do not hide it in a token estimate.
Failure modes and mitigations
F1: The interface changes
Run canaries against approved test records, verify labels and resulting values, and pause the affected workflow after a failed check. Accessibility labels help but can change too. A vision model adapting successfully once is not proof of general UI stability.
F2: The agent repeats an ineffective action
Track (report state, proposed action, failure class) alongside step/time/token budgets. An unchanged screen may hide successful server processing, while animation may change pixels without progress. Stop repeated failed transitions rather than relying only on screenshot similarity.
| Observation | Correct progress judgment |
|---|---|
| Upload dialog unchanged but a receipt ID now exists | Real progress; verify attachment processing within a deadline |
| Submit disabled with missing cost-center error | Resolve the missing field; repeating the click is not progress |
| Submission request timed out | Reconcile remote outcome; do not repeat Submit |
F3: A receipt contains prompt injection
Treat the text as evidence, enforce permitted operations outside the model and verify every attachment/destination. Freeze the task for review if the agent proposes an out-of-scope operation. The detection score does not replace the action boundary.
F4: The browser enters the wrong company
Verify account and report identity before sensitive reads and every write. Separate profiles and application roles. A shared vendor hostname and a valid audience token cannot distinguish Company A from Company B.
F5: Audit storage fails
Pause actions requiring a durable record. After recovery, reconcile incomplete action intents with the target before retry. Acknowledging intent storage does not mean the outcome was recorded or that the remote write succeeded.
F6: A task exhausts its budget
Reserve expected spend before each model/tool call and account for unknown billed outcomes. Bound steps, wall time, retries and aggregate tenant spend. An expensive fallback requires a budget decision before dispatch, not after the invoice arrives.
F7: Reviewers rubber-stamp approvals
Show concise meaningful diffs with receipt evidence, reduce redundant confirmations and staff queues realistically. Audit sampled approved reports and use separate safe training exercises. Do not insert fake financial actions into live work to test attention.
F8: Receipt extraction loses a number or currency
Compare total, tax, currency, date and merchant against the source image and relevant record. Missing or contradictory evidence goes to clarification. OCR confidence values are not interchangeable across models or a guarantee of financial correctness.
F9: A model or tool version retires
Pin the tested model/tool combination and monitor official lifecycle notices. Qualify an alternative on the same workflow, action-boundary and recovery suite before switching. A schema-compatible replacement is not automatically behavior-compatible.
F10: The browser or worker crashes
Use heartbeats, expiring leases and an independent cleanup service. Fence the old worker at the trusted executor, terminate/revoke its session and reconcile dispatched actions. Requeue only after duplicate-execution risk is resolved; cap retries and preserve evidence for human takeover.
Operational Considerations
Quality and deployment
Evaluate complete reports in a test tenant with adversarial receipts, ambiguous currencies, UI changes, wrong-account states, duplicate submissions and lost responses. Measure prepared-without-repair rate, post-review correctness, executed unauthorized actions, approval age, unknown-outcome resolution and full cost. A rejected unsafe proposal differs from an executed unauthorized action; do not combine the counts.
OSWorld 2.0 contains 108 longer workflows with a median human completion time of about 1.6 hours and distinguishes binary completion from partial progress. Its tasks and protocol differ from earlier OSWorld results. Use such benchmarks as research context, then qualify the actual expense workflow; a benchmark percentage cannot justify this company's payment permissions. OSWorld 2.0.
Replay recordings in a test environment with controlled fixtures. Replaying production browser actions can repeat real writes. Shadow evaluation must be read-only or isolated from production effects. Roll out by workflow and tenant cohort, retain the previous model/controller bundle and stop changes when a critical boundary fails.
Cost model
Assume 40,000 billed input tokens and 8,000 output tokens across a complete task, including screenshot representation, tool definitions and repeated context. At Sonnet 5's $2/$10 per million-token rates, model cost is $0.16/task. Claude pricing.
| Component | Per-report allowance |
|---|---|
| Normal model workload | $0.16 |
| Sandbox and browser | $0.04 |
| Scoped audit evidence | $0.04 |
| Evaluation and sampling | $0.03 |
| Partial machine cost | $0.27 |
At 14,000 × 52/12 reports per month, this totals $16,380/month. The audit allowance assumes minimized, measured evidence retention, not continuous seven-year video. Include idle capacity, vendor licenses, support and human review in the operating budget.
The same token counts on GPT-6 Astra at $10/$50 cost $0.80 in model calls alone. If 5% of tasks need an additional full Astra attempt after the normal attempt, average partial machine cost rises to $0.31, already above the $0.30 target. This is an economic gate to evaluate, not a reason to silently exceed the budget. OpenAI pricing.
Using the 350 weekly human hours above and an assumed $40/hour gives about $60,667/month in review labor. Combined with the normal machine allowance, the partial total is approximately $77,047/month, before the other excluded costs. Do not compare it only with three contractors whose workload covers a small fraction of all reports.
On-call response
| Trigger | Immediate action | Evidence needed before resuming |
|---|---|---|
| Executed out-of-scope action | Stop writes and affected sessions | Scope, approval and remote-effect investigation |
| UI canary failure | Pause that workflow or use manual handling | Corrected mapping and complete test report |
| Approval backlog | Reduce admissions and assign review capacity | Queue age and measured handling time |
| Rising costs or no-progress loops | Stop excess work within budget policy | Repeated-state analysis and bounded recovery |
| Unknown submission outcome | Hold the task for reconciliation | Remote receipt or conclusive non-application |
Lowering model temperature is not a security incident response. Fix the failed authority boundary and validate recovery before re-enabling writes.
Interview follow-ups
1. Why not use scripts for everything?
Use them where workflows and selectors are stable. A model can help interpret variable layouts, but adds inference cost and error modes. Compare maintenance, task quality and recovery against APIs and deterministic automation.
2. Does a microVM make payment automation safe?
It isolates local state and resources. It does not restrict the authority of a live payment session by itself. Application roles, trusted action enforcement and exact approvals are separate requirements.
3. How do you prevent a stale approval from paying the wrong amount?
Bind approval to the canonical action, account/report identity, relevant state and policy version. Revalidate at atomic dispatch reservation. If the target cannot support the required state check, keep submission with the human.
4. Can the agent retry after a browser timeout?
Not until it establishes what happened. Reconcile the remote payment/report and use supported deduplication where available. Unknown outcomes require human resolution when the legacy system cannot provide a conclusive check.
5. What proves the audit is complete?
No hash chain alone does. Compare durable intents, trusted dispatch records and remote receipts, investigate missing transitions and protect anchors. Screenshots are supporting evidence with privacy and retention costs.
6. Does 84% machine success mean one reviewer can handle exceptions?
No. At this volume, 16% technical exceptions already require about 187 hours/week at five minutes each. Add routine confirmation and high-risk approval to get a realistic staffing model.
7. How do you choose the next model?
Check its supported tool schema and lifecycle, then run the same representative and adversarial workflows with complete-cost measurement. Use an isolated shadow and bounded rollout; never promote from a generic benchmark rank alone.
Closing notes
The useful automation boundary is a verified business action under real authority, not a sequence of plausible clicks. Start with reliable draft preparation, add submission only where controls can be enforced, and close with the limits of legacy state checks, human review capacity and uncertain-outcome recovery.
Related: Tool-use landscape, Agent fundamentals, LLM security.