A useful tool-agent use case has a concrete task, a permitted action surface, a verifiable result, and an acceptable recovery path. “Add an agent to operations” is not a requirement. “Prepare a replenishment request from current stock and open orders, then create the approved request without duplication” is specific enough to design and evaluate.
This chapter provides original hypothetical interview scenarios, not claims about customer deployments or the author's employment history. Every workload, cost, and target below is an explicit design assumption. Published research and product evidence are identified separately and were checked on September 24, 2026.
The aim is to explain where model judgment helps, where ordinary software should remain in control, and what evidence would justify expanding the deployment.
A five-question selection test
| Question | Promising answer | Warning sign |
|---|---|---|
| What exact outcome is useful? | Saved record, validated draft, tested change, or sourced answer | “More autonomous activity” |
| Why does this need a model? | Variable language, incomplete structure, or context-dependent interpretation | A fixed rule already solves the task cheaply |
| Can we establish correctness? | Observable postconditions and independent checks | Only the agent's confident final message |
| What authority is necessary? | Narrow, attributable, revocable access | Broad administrator access for convenience |
| Can we recover economically? | Bounded retry, reconciliation, or human handling | Irreversible effects with no usable evidence |
Start with the existing process as the baseline. A deterministic workflow may be the right product, with a model used only for extraction, explanation, or exception handling. A fixed sequence of five steps does not require five agents.
Developer productivity
Useful tasks include code navigation, scoped implementation, refactoring, test preparation, and build-failure investigation. An agent can inspect a repository, propose a diff, run checks, and package review evidence. Compilation and tests provide feedback, but passing them does not establish complete correctness, security, or maintainability.
| Workflow | Useful output | Measure | Failure to include in evaluation |
|---|---|---|---|
| Feature work | Reviewed change meeting acceptance criteria | Time to accepted change, escaped defects | Plausible code that misses the requirement |
| Refactoring | Behavior-preserving transformation | Regression results and review effort | Changed edge-case semantics |
| Test generation | Tests that detect relevant faults | Mutation/bug detection and maintenance | Tests that simply mirror the implementation |
| Code review | Actionable findings with locations/evidence | Precision, missed important defects | Large volumes of low-value comments |
| CI triage | Reproducible diagnosis and scoped repair | Recovery time and recurrence | Treating a flaky pass as a fixed root cause |
Giving an agent shell access changes the attack surface even when CI already uses containers. Repository scripts, dependencies, model context, network access, and credentials become part of the design. Keep release authority separate from the ability to prepare a change. See the autonomous coding case.
Evidence, with limits: METR's early-2025 study measured a 19% slowdown for its participating experienced open-source developers. Its February 2026 update said later results were difficult to interpret because of selection and measurement effects. Neither finding establishes a universal current speedup or slowdown. Use task- and team-specific evaluation rather than a promised “2–10× productivity” multiplier. METR update.
Business process automation
Document extraction, reconciliation, data entry, and recurring reports often have observable outputs. High volume can justify integration work, but neither high volume nor a familiar form makes the data correct.
Read diagram source
flowchart LR
I[Documents and source records] --> E[Extract values with source evidence]
E --> V[Validate types, totals, identity and policy]
V -->|Complete and permitted| D[Prepare downstream record]
V -->|Missing or conflicting| H[Exception review]
H --> D
D --> C[Authorized commit]
C --> R[Reconcile saved record and source]
Use parser/OCR output as evidence to validate, not as unquestioned truth. Two models can agree on the same incorrect amount. Reconcile the record identity, currency, line items, totals, and permitted destination. Keep accounting entry and payment execution distinct.
For a legacy interface, use supported APIs or semantic controls where available, and computer use for the remaining gap. Budget for changing layouts, login recovery, review, and unknown submissions. Measure cost per verified completion, including spending on failed attempts.
Customer support, sales, and onboarding
The meaningful distinction is the operation's authority and consequence, not a universal support-tier number.
| Task | Reasonable starting scope | Important check |
|---|---|---|
| Order status | Read authorized order and explain current state | Customer/order binding and fresh source |
| Account recovery | Guide the supported identity process | Never replace identity verification with conversational confidence |
| Billing complaint | Gather evidence and prepare a proposed resolution | Policy, amount, approval, and transaction outcome |
| Sales research | Summarize permitted public/CRM information | Provenance, freshness, and access terms |
| Outreach | Prepare accurate, scoped drafts | Recipient and sending authorization |
| Onboarding | Explain setup and prepare allowed configuration | Identity, resource limits, and current product behavior |
| Identity/KYC workflow | Orchestrate approved checks and review | Applicable process, specialist decisions, and protected data |
An agent can route routine, well-understood cases while handing exceptions to people. “Escalate on any uncertainty” can overwhelm reviewers and discourage useful uncertainty reporting. Define actionable reasons to escalate, priority, context, and a fallback owner. Measure wrong resolutions and human rescue, not only conversation containment. See customer support automation.
IT operations
Start with scoped evidence gathering: correlate an alert with recent releases, relevant logs, dependencies, and a current runbook. Root cause is a hypothesis until supported by evidence; timestamp correlation alone is insufficient.
| Action | Why context matters | Appropriate control |
|---|---|---|
| Query logs | Read access may expose sensitive data | Scoped queries, redaction, bounded results |
| Suppress an alert | Can hide an active incident | Reason, affected scope, expiry, and reversal |
| Restart a workload | May lose in-memory work or reduce quorum | Service-specific policy, health preconditions, blast-radius limit |
| Roll back a release | May conflict with a migrated schema | Tested compatibility and verified target revision |
| Scale down | Can remove capacity or stateful replicas | Workload/SLO constraints and authorized change |
| Publish an incident update | Communicates external commitments | Evidence-backed wording and destination authority |
A restart is not universally low risk, and read-only access should not mean “can observe everything.” Enforce task-specific permissions and bounded changes. A dry run is useful evidence but may omit effects that occur only during real execution. See reliable agent operations.
Research and analysis
Research agents gather sources, extract comparable facts, perform calculations, and distinguish observations from interpretation. Citation presence alone is insufficient: the cited passage must support the claim, match the relevant time period, and remain accessible to the intended reader.
- Define the decision and the source scope.
- Retrieve dated primary evidence where available.
- Normalize units, periods, definitions, and entities before comparison.
- Record contradictory and missing evidence.
- Separate supported facts, calculations, assumptions, and inferences.
- Review the resulting recommendation at a level appropriate to its impact.
Correctness is not unique to research. Coding tests can miss bugs; reconciled totals can hide wrong accounts; a compliance form can be syntactically valid but substantively wrong. Verification must fit the domain. See financial analysis and enterprise knowledge.
Interview scenario 1: a distribution operations assistant
Assumptions: 120 trusted operators across eight distribution sites ask about stock, shipments, and replenishment. There are 1,200 requests per working day. Current warehouse and purchasing systems remain authoritative. The first release reads status and prepares replenishment requests; it does not choose new suppliers, change bank details, or authorize purchases.
Functional requirements
- Authenticate the operator and identify permitted sites and records.
- Answer stock/shipment questions using current source values and timestamps.
- Reconcile on-hand stock, reservations, and outstanding replenishment before drafting.
- Prepare a replenishment request with item, quantity, site, source evidence, and version.
- Create the exact approved request under a deduplication contract.
- Show pending/unknown outcomes and route exceptions to an accountable person.
- Produce a daily report with explicitly authorized recipients.
Non-functional requirements
- Scope: prevent cross-site disclosure or action beyond the operator's role.
- Freshness: include source timestamps; do not present cached stock as current availability.
- Responsiveness: aim for p95 status answers within eight seconds under the assumed load.
- Correctness: never report a purchase request as saved without its authoritative ID.
- Recovery: preserve operation state across gateway/worker failure.
- Economics: measure full cost, review minutes, and avoided repeated work.
Basic design and first flaw
A chat bot forwards an operator's request to a model with warehouse and purchasing tools. This can answer a simple question, but “order 40 units” may ignore 30 units already on order or a site the operator cannot manage.
Repair: keep identity and site scope in trusted request context. Retrieve stock, reservations, and open orders before creating a versioned proposal. The business service validates permitted quantities and authorization. This adds source calls but prevents the model from treating incomplete inventory context as permission.
Detailed architecture
Read diagram source
flowchart TD
U[Authenticated operator or permitted team channel] --> G[Trusted gateway and session routing]
G --> P[Model proposes status answer or replenishment draft]
P <--> R[Scoped warehouse and shipment reads]
R <--> W[Warehouse and carrier systems]
P --> D[Deterministic replenishment service]
D <--> V[Current stock, reservations and open orders]
D --> I[Versioned intent and operation ledger]
I --> H[Authorized reviewer]
H --> C[Commit with current policy and preconditions]
C --> E[Purchasing system]
E --> O[Saved request ID or unresolved outcome]
O --> I
I --> G
OpenClaw can supply the trusted gateway, sessions, and tool loop. Its workspace/persona files do not implement warehouse authorization. One shared gateway suits a trusted team; unrelated customers need separate trust boundaries. See the OpenClaw design chapter and its official security model.
Data and failure contracts
| Record | Minimum evidence | Why it exists |
|---|---|---|
| Request | Operator, permitted scope, source message/event ID | Attribution and redelivery handling |
| Inventory observation | Item/site, values, source revision/time | Grounds the proposed quantity |
| Replenishment intent | Normalized payload, evidence references, expiry | Exact reviewable proposal |
| Operation | Scoped key, payload hash, state, remote request ID | Correct retry/reconciliation |
| Delivery | Intended audience, report revision, send outcome | Prevents blind notification repeats |
Second flaw: two operators approve overlapping drafts for the same item/site. Serializing chat messages is insufficient because stock and orders can change through other applications. Revalidate in the authoritative transaction where possible; otherwise use the purchasing system's concurrency controls and reconcile outcomes. A business service may reject a stale intent and require a new review.
Third flaw: supplier text says to ignore stock rules. Keep external content as untrusted evidence, with no ability to change tool policy or reviewer identity. A scanner can add detection, but the deterministic service must still reject an unauthorized action.
Capacity and economics
At 1,200 requests in eight hours, a 4× sustained peak is 0.167 requests/second. With a 12-second mean active task across the mixed workload, expect about two active tasks; at 65% planned occupancy, start with four slots and load-test tail behavior and downstream quotas. Measure the simpler status-answer path separately against its eight-second p95 target.
For 22 workdays, assume 26,400 requests/month. Model/tool usage at $0.03/request costs $792; infrastructure/monitoring is $400; 12 maintenance hours at $100 cost $1,200. If 10% need two review minutes at $45/hour, review costs $3,960. Total: $6,352/month, before initial development.
If the previous process took three human minutes/request at $45/hour, its allocated labor value was $59,400/month. That is a workload estimate, not an automatic cash saving: retained staff may use the freed time elsewhere. Validate task mix and actual displaced work before claiming a return.
Closing tradeoff: the gateway makes a familiar interface available quickly; trustworthy stock evidence and purchasing controls determine whether it can act correctly. Expand from read-only status to narrowly authorized requests only after measuring freshness, duplicate effects, and review effort.
Interview scenario 2: validate a legacy policy-system migration
Assumptions: a policy-administration platform has COBOL code, batch jobs, legacy files, and terminal workflows. The team wants to migrate one bounded policy-servicing capability first. Coding and computer-use agents assist analysis and testing; domain engineers own accepted behavior. There is no assumed percentage reduction in project duration.
Functional requirements
- Inventory programs, copybooks, batch dependencies, file layouts, interfaces, and operational schedules.
- Link candidate business rules to executable paths, source locations, and observed examples.
- Implement a scoped replacement behind a controlled routing boundary.
- Run repeatable differential tests on authorized, isolated data.
- Compare outputs and side effects, including boundary/error behavior.
- Reconcile migrated records and manage a controlled cutover and rollback decision.
- Preserve a human-reviewed record of intentional behavior changes.
Non-functional requirements
- Semantic fidelity: preserve approved arithmetic, dates, ordering, and error behavior.
- Isolation: test runs must not send real payments, notices, or external business changes.
- Repeatability: pin input snapshot, compiler/runtime configuration, fixtures, and generated revision.
- Data integrity: reconcile identity, counts, checksums where applicable, balances, and exceptions.
- Availability: agree a migration window and recovery objective for the selected capability.
- Auditability: retain rule-to-test-to-change evidence with appropriate access control.
Basic design and first flaw
Translate a program, compile it, then compare a few screens. This proves that some code builds and some views look familiar. It does not prove that batch posting, file encoding, numeric behavior, or rare policy transitions remain correct.
Repair: create a differential test harness around the domain capability. Run old and new implementations from the same input snapshot with controlled time, dependencies, and random inputs. Capture semantic results and intended side effects in isolated sinks.
Read diagram source
flowchart LR
S[Authorized input snapshot and versioned fixtures] --> L[Legacy program and isolated dependencies]
S --> N[Replacement at exact revision]
L --> A[Canonical results and captured side effects]
N --> B[Canonical results and captured side effects]
A --> D[Differential checks and domain invariants]
B --> D
D --> E[Investigate mismatches with domain expert]
E --> R[Approved behavior and regression fixtures]
R --> G[Controlled traffic and data migration gate]
What to compare
| Area | Important cases | Why a generic translator misses them |
|---|---|---|
| Decimal arithmetic | Scale, rounding point, negative values, overflow | A binary float or a later rounding step can change results |
| Record formats | Packed/zoned fields, encoding, padding, signed values | Display text is not the complete storage contract |
| Control flow | Loop boundaries, fall-through, exceptional paths | Similar-looking code can have different evaluation behavior |
| Batch processing | Ordering, restart points, duplicate input, partial completion | Interactive tests do not exercise overnight workflows |
| Dates and schedules | Cutoffs, leap dates, business calendars, timezone | A valid timestamp may mean a different business period |
| Side effects | Posting, notification, downstream export | Matching screens can hide different external effects |
For a Java replacement, BigDecimal supports decimal arithmetic with explicit scale and rounding control. It does not automatically reproduce an existing program's evaluation and rounding rules; those must be specified and tested. See Java decimal arithmetic.
Worked rounding example: if the approved rule is round decimal 2.675 to two places using half-up, the answer is 2.68. A binary floating-point representation can behave differently. This is an example of a declared rule, not a claim that every COBOL calculation uses half-up.
Where agents help
- A coding agent maps dependencies, proposes translations, and adds targeted tests.
- A retrieval tool locates source/rule evidence rather than relying on a generated summary alone.
- A computer-use worker exercises a terminal or GUI flow when no suitable test interface exists.
- Deterministic comparison and domain review decide whether a difference is acceptable.
Use terminal protocol/screen-field APIs where they exist; visual control is a fallback. Pin the font, geometry, locale, and environment for repeatability while also testing expected variants. A matching screenshot is one observation, not a complete migration acceptance test.
Cutover and recovery
- Complete the bounded capability's acceptance criteria and identify unresolved behavior.
- Establish a consistent data snapshot and transformation mapping.
- Use supported change capture/replay or an agreed write freeze to close the migration gap.
- Reconcile migrated data by record identity and domain invariants, not only row count.
- Route a controlled cohort or capability to the replacement.
- Monitor correctness and operations; apply the preplanned rollback or forward-repair path.
Second flaw: both systems write to the same downstream service during a parallel run. Redirect test/shadow effects into isolated sinks and keep one authoritative writer. This costs environment work but prevents duplicate business actions.
Third flaw: a new system accepts writes that the old schema cannot represent. “Switch traffic back” is now insufficient. Plan backward-compatible data changes or a tested reverse/forward transformation before cutover.
Capacity and investment example
A suite of 12,000 two-system fixtures averaging three seconds of worker occupancy needs 36,000 worker-seconds. Twenty workers at 75% effective utilization give a planning duration of 40 minutes, excluding startup, long-tail cases, and investigation. Run critical small suites more often and the full suite at release gates.
Suppose the scoped migration pilot requires 600 engineering hours at $120/hour and $4,000 in tools/test infrastructure: $76,000. If measured ongoing benefit is $8,000/month and additional operations cost $2,000/month, simple payback is about 12.7 months. If net benefit falls to $3,000/month, payback is about 25.3 months. Do not claim a seven-month transformation without a measured scope and plan.
Closing tradeoff: agents can accelerate understanding and implementation, but behavior preservation and data cutover remain engineering obligations. Migrate a bounded capability with explicit acceptance evidence, then use what the pilot teaches to estimate the rest.
Interview scenario 3: regulatory reporting evidence and preparation
Assumptions: an operations team must prepare a particular reporting obligation from transaction data. Qualified owners define the applicable jurisdiction, product scope, rules, submission destination, and deadlines. The assistant gathers evidence, explains exceptions, and prepares a package. It does not invent legal applicability or assume all obligations go to one regulator.
As one concrete distinction, ESMA's EMIR materials describe reporting to trade repositories and include validation, reconciliation, and response mechanisms. That differs from treating every derivatives report as a generic SEC/FINRA web form. Consult the actual applicable rules and interfaces. ESMA reporting guidance.
Functional requirements
- Ingest transaction events and revisions with stable identity.
- Apply an approved, effective-dated ruleset to the defined reporting scope.
- Assemble required fields and source evidence into the correct package version.
- Detect missing values, conflicts, late source corrections, and unsupported cases.
- Present the exact package and exceptions for the required review.
- Submit through the approved service and track acceptance, rejection, correction, and cancellation.
- Reconcile expected reports against submitted and acknowledged records.
Non-functional requirements
- Completeness: reconcile the in-scope event population, not just successful submissions.
- Version fidelity: bind each result to source, rule, schema, and package revisions.
- Authorization: separate preparation, review, and submission authority where required.
- Timeliness: measure deadline risk and exception age against the applicable obligation.
- Recovery: retain durable submission state and avoid uncontrolled duplicates.
- Privacy/audit: protect sensitive records and preserve required decision evidence.
Basic design and first flaw
Four model agents read trades, classify rules, generate forms, and submit them. A well-formed message between agents cannot prove that the right obligation was selected. More model roles also do not create independent evidence.
Repair: use deterministic event processing, approved rule evaluation, and schema validation for prescribed operations. Use a model for source interpretation assistance, exception explanations, and review preparation. A workflow controller owns state and progression.
Read diagram source
flowchart TD
T[Transaction events and corrections] --> I[Validated event ledger]
I --> R[Approved effective-dated rule service]
R --> P[Versioned package builder and field validation]
P --> X[Exception work queue]
X --> A[Agent-assisted evidence and explanation]
A --> H[Qualified reviewer]
P --> H
H --> S[Authorized submission service]
S --> E[Applicable reporting endpoint]
E --> K[Acceptance, rejection and correction state]
K --> L[Reconciliation ledger]
I --> L
L --> X
A specialized-agent arrangement may help with different documents, tools, or permissions. It remains optional. Keep typed handoffs, provenance, shared budgets, and deterministic release gates regardless of how many agents you use.
Records and freshness
| Record | Version/evidence to retain | Failure it reveals |
|---|---|---|
| Source event | Event ID, revision, business time, ingestion time | Correction or late arrival |
| Rule decision | Applicable rule/version, effective interval, inputs | Stale rule or wrong applicability |
| Package | Exact bytes/hash, schema version, source references | Changed report after approval |
| Review | Reviewer, reviewed package hash, exceptions, decision | Approval of a different package |
| Submission | Logical report ID, attempt, endpoint receipt, status | Accepted transport but rejected report |
| Reconciliation | Expected/submitted/accepted counts and record differences | Missing reports masked by success averages |
Second flaw: a ruleset changes after a package was prepared. Check whether the change affects the package's applicable period and obligation. Recompute and obtain review when needed; do not blindly replace historical rules with the newest version for every past event.
Third flaw: the submission HTTP request succeeds but the endpoint later rejects the report. Keep submitted and accepted distinct and route the rejection with its due time. A transport acknowledgment is not regulatory acceptance.
Reviewer capacity and economics
Assume 8,000 packages/month. If every package receives four minutes of required review and 10% need twelve additional minutes for exceptions:
8,000 × 4/60 + 800 × 12/60 = 693.3 reviewer hours/month.
At 120 productive review hours/person/month, this is about 5.8 people of capacity, before absence and peak coverage. Adding an agent does not reduce a mandatory review step to zero.
If measured preparation time falls from 18 to seven minutes/package, gross released preparation time is 1,466.7 hours/month. Whether those hours create savings or more capacity depends on work allocation and whether the measured baseline includes the same review obligations. Do not double-count the same minutes in preparation savings and reviewer savings.
Closing tradeoff: specialize model work where it helps people interpret evidence, while rules, versioned packages, submission state, and reconciliation remain explicit. Judge the design on completeness and accepted outcomes, not the number of agents or generated reports.
Calculate value with units and an honest baseline
Net benefit is a currency amount. Return on investment (ROI) is a ratio over a stated period. Payback is the time to recover an initial investment under an assumed cash-flow model. Do not use these labels interchangeably.
| Quantity | Simple definition | Caveat |
|---|---|---|
| Gross benefit | Realized labor/capacity value + measured error reduction + other non-overlapping benefit | Avoid counting faster work and released labor twice |
| Operating cost | Model, tools, runtime, storage, review, maintenance, support, and incidents | Include failed/abandoned work |
| Monthly net benefit | Monthly gross benefit − monthly operating cost | State how benefits are monetized |
| Simple payback | Initial investment / positive monthly net benefit | No finite payback when net benefit is nonpositive |
| Period ROI | (Period benefits − period costs) / period costs | Include initial investment and recurring costs for that period |
| Cost per verified completion | All relevant operating cost / verified completions | State the denominator and observation window |
Worked sensitivity table
Suppose the old process used 2,000 hours/month. Automation releases 800 hours, and the relevant labor value is $50/hour. Only a fraction may become realized savings or valuable redeployed capacity. Assume operating cost of $12,000/month and initial investment of $60,000.
| Realized fraction of released-hour value | Monthly benefit | Monthly net benefit | Simple payback |
|---|---|---|---|
| 25% | $10,000 | −$2,000 | No positive payback under these assumptions |
| 50% | $20,000 | $8,000 | 7.5 months |
| 100% | $40,000 | $28,000 | About 2.1 months |
At 50% realization over twelve months, benefits are $240,000 and costs are $60,000 + 12 × $12,000 = $204,000; period ROI is about 17.6%. This simplified example ignores discounting and financing.
Keep the metric denominators visible
- Straight-through processing: verified completions with no human intervention divided by eligible tasks, with eligibility defined in advance.
- False-completion rate: tasks marked successful whose required outcome is wrong or absent, divided by all tasks marked successful.
- Escalation rate: tasks requiring a specified human action divided by attempted tasks.
- End-to-end time: intake to verified resolution, including queues and review.
- Operational workload: human minutes, unresolved outcomes, and oldest exception.
- Resource use: tool calls, tokens, retries, and environment occupancy per task/outcome.
Attempt counts, token usage, and uptime are useful diagnostic and capacity metrics. They simply do not establish business value on their own. Compare against a measured baseline, retain absolute correctness requirements, and report uncertainty; a worse-than-human failure rate is not the only possible release blocker.
Published evidence and practical failure lessons
| Evidence | What it supports | What it does not prove |
|---|---|---|
| Replit's July 21, 2025 database-separation announcement | Development and production data need distinct operating boundaries | A prompt or rollback button prevents every harmful change |
| Cisco's March 2026 discussion of malicious third-party skills | Extensions can introduce credential theft and injected instructions | A scanner catches every attack or every installation has the same exposure |
| June 2026 memory-poisoning research | Persistent writes can carry adversarial influence into later tasks | Every memory implementation has the same measured exploitability |
Apply the lessons as concrete controls:
- Keep development credentials and data separate from production authority; exercise restoration against real recovery objectives.
- Review and pin extension versions, restrict execution and network access, and monitor unexpected effects.
- Record memory origin and scope, review sensitive promotions, and support correction/deletion of contaminated state.
- Keep critical authorization policy in an enforced service; a system prompt alone is not a policy engine.
- Evaluate reviewers on sampled decision quality, not on a required correction quota. A 100% approval rate may be appropriate for a narrow, correct workload; investigate it with evidence rather than assuming rubber-stamping.
Additional exercise: invoice entry at 5,000 invoices/month
Functional requirements
- Ingest invoices with source and supplier identity.
- Extract required fields with page/region evidence where available.
- Validate totals, currency, purchase-order matching, receipt evidence, and duplicate identity.
- Route policy exceptions and required approvals.
- Create an accounting entry in the authorized ERP with a stable operation key.
- Reconcile saved entries and unresolved attempts.
Non-functional requirements
- Prevent unauthorized supplier/account changes and duplicate entries.
- Protect documents and bank details in storage, logs, and model requests.
- Meet an agreed processing deadline with measured reviewer capacity.
- Preserve a trace from input document to reviewed data and final ERP record.
- Bound retries, costs, and unresolved-outcome age.
Scope exclusion: this design prepares/posts approved accounting entries; payment release uses its separate authorized process.
Basic design: OCR → model extraction → ERP write. Flaw: a visually plausible invoice can contain changed bank details or an already-paid invoice reference. Repair: match authoritative supplier/PO/receipt data, flag changes, and use business-level duplicate checks. Two model reads are helpful evidence but do not establish independent truth.
Read diagram source
flowchart TD
D[Invoice and source identity] --> E[Extract fields with evidence]
E --> V[Deterministic totals and schema checks]
V --> M[Supplier, PO, receipt and duplicate checks]
M --> P[Versioned proposed accounting entry]
P --> H[Required approval and exception review]
H --> C[Scoped idempotent ERP adapter]
C --> R[Saved entry or unknown outcome]
R --> Q[Reconciliation and exception queue]
Q --> H
Serializing every ERP call does not prevent duplicates caused by retries or another application. Use the ERP's supported concurrency/idempotency contract and a scoped operation ledger. Reconcile individual entries; matching a bank-statement total cannot prove that each invoice went to the correct account.
Economic check: at 5,000 invoices/month, 15% requiring six review minutes means 75 review hours/month. At $40/hour, that is $3,000, before model, ERP, runtime, maintenance, and other approval work. A different exception mix changes the answer more than a small token-price improvement.
Closing answer: “I would automate evidence extraction and routine preparation, retain supplier and payment authority in the existing business controls, and verify each saved ERP entry. The launch decision depends on false entries, duplicate prevention, exception workload, and full unit cost, measured on representative invoices and failure cases.”
Interview questions and answer notes
- How do you decide whether a task needs an agent? Identify where variable interpretation adds value beyond a fixed workflow and whether the outcome can be verified.
- Do passing tests prove coding-agent correctness? No. Test coverage, requirements, security, and operational behavior still matter.
- Is a status lookup automatically safe? No. The result may be private, stale, or for the wrong account.
- Why does a stock assistant need open-order data? On-hand quantity alone can cause unnecessary replenishment.
- Why are matching screens insufficient for migration? They can miss arithmetic, batch behavior, storage semantics, and side effects.
- Can old and new systems both perform live effects during a shadow run? That risks duplicates; isolate shadow effects and define one authoritative writer.
- Why is the newest rule not always correct for an old transaction? Applicability depends on the relevant effective interval and obligation.
- Does submission success mean a report was accepted? No. Track transport, validation, acknowledgment, correction, and rejection separately.
- Does an approval rate of 100% prove poor review? No. Measure sampled accuracy and reviewer engagement in context.
- Are hours saved the same as cash saved? No. State whether time reduces spending, avoids hiring, or creates measured redeployed capacity.
- Can global write serialization solve invoice duplicates? No. Repeated operations and other writers still require identity, concurrency, and idempotency controls.
- What makes a compelling interview close? A bounded launch scope, verifiable outcome, recoverable failures, measured economics, and a clear expansion criterion.
Final notes
Remember task → evidence → authority → action → recovery → value. Treat forecasts as forecasts, examples as examples, and measured outcomes as measured outcomes. A credible design earns its automation level through observable quality and practical recovery.
Next: Safety and governance.