Numerical examples are illustrative unless explicitly sourced.
Human-in-the-loop (HITL) systems include human input, judgment or authorization at defined points in an automated process. The person may supply missing information, review a proposal or resolve an exception. Human involvement is a system component whose effectiveness must be measured.
Remember: Right action. Right reviewer. Right evidence. Right time.
Understand why a human is involved
Human review can serve different purposes. A domain expert may supply missing judgment. An account owner may grant permission. An operator may resolve an uncertain external outcome. Those are different jobs, and one “Approve” button does not automatically solve all three.
Consider an agent proposing a $400 refund. The reviewer needs the relevant order, policy evidence, proposed amount, destination, and reason the case was escalated. A one-line summary can help orient them, but it must link to the evidence and show consequential details. If the display hides the destination, the reviewer cannot meaningfully approve where the money goes.
Human-in-command means a person directs decisions throughout the task. Approval before execution allows the system to prepare a proposal but pauses before a consequential action. Escalation on uncertainty or policy conditions sends selected cases to a person. After-the-fact audit checks completed work. Audit can improve future behavior but cannot prevent the already-completed action. Choose the pattern according to the consequence and reversibility.
Follow an approval through its full lifecycle
The agent creates a structured proposal with an ID and version. The application validates basic eligibility, stores the proposal and evidence references, and moves the job into an awaiting-review state. A durable wait means a worker restart does not discard the pending decision.
The lifecycle has explicit transitions:
- Validate and persist the exact proposed action and evidence.
- Determine whether existing authorization covers it or policy requires review.
- Assign a qualified reviewer and an expiry when review is required.
- Record approval, rejection or requested changes against the proposal version.
- Revalidate current permissions and business preconditions at execution.
- Execute under a durable operation identity and record the verified outcome.
If the amount changes, an approval for the previous exact amount does not authorize the new one. Re-evaluate the policy and any explicitly bounded standing authorization. Handle duplicate approval messages so one decision cannot start multiple independent actions. Avoid repeatedly asking for permission already granted within its valid scope.
An interrupt pauses execution at a defined point. A resume supplies the decision and lets the workflow continue under its framework's semantics. Saving application state does not preserve a model's private internal thoughts, and resuming from an earlier state does not undo an external action. If a human edits the plan after a payment, the system must account for the payment that actually occurred.
Work together on an explicit plan
A person can help form the plan before approving an action. This is sometimes called co-reasoning or a shared scratchpad. The useful shared artifact is a reviewable plan, relevant facts, assumptions, open questions, and evidence—not a claim to expose the model's private internal thoughts.
For a travel task, the plan might say: “Book the Tuesday flight to Paris, France; arrive before 14:00; budget $800; no purchase yet.” The user corrects the destination to Paris, Texas. Store plan version 2, record the corrected fact and its source, invalidate the old search results, and recompute the itinerary. A reviewer can now inspect how that correction changes the proposal. Merely replacing the destination label on the old itinerary would preserve the wrong airports and prices.
Distinguish editing from authorization. Editing the plan does not approve a purchase. Any eventual approval binds to the exact itinerary, amount, traveler, and proposal version. If a ticket was already bought, show that completed effect and use a separate cancellation or exchange process; editing the scratchpad cannot undo it.
Practice question: When is collaborative planning better than a final approval button? Answer: When the person has missing context or can catch a wrong assumption before expensive work begins. Keep final-action checks as well, because agreeing on the goal does not authorize every means of achieving it.
Treat the review queue as a service with capacity
If 1,000 tasks per hour arrive and 20% need review, the queue receives 200 reviews per hour. At three minutes per review, that is 600 reviewer-minutes, or ten hours of review work each hour before breaks, variation, and headroom. A design with three reviewers cannot sustain that workload. These are illustrative assumptions; measure actual arrivals and handling times.
| Capacity calculation | Result |
|---|---|
| Reviews arriving per hour | 1,000 × 20% = 200 |
| Work per hour | 200 × 3 = 600 reviewer-minutes |
| Fully occupied reviewers | 600 ÷ 60 = 10 |
| Illustrative staffing at 80% planned utilization | 10 ÷ 0.8 = 12.5, rounded up to 13 |
This average-capacity calculation does not prove a queue-delay SLA. Bursts, case mix, shift coverage and skill-specific routing require additional modeling and measurement.
Reduce unnecessary escalation by improving the product and the quality of evidence, not by asking reviewers to approve faster without understanding. Route to appropriate expertise, prioritize by consequence and deadline, support explicit deferral, and monitor queue age and reviewer disagreement. Self-reported model confidence or token probability is not automatically a calibrated probability that an action is safe.
Over-reliance occurs when people accept machine proposals too readily. Clear evidence, manageable workload, training, independent sampling, and feedback can help. Practice drills belong in a disclosed training environment; quietly inserting harmful errors into real customer decisions is not an acceptable attention test.
Approval is a state machine
Imagine a refund above a business-defined limit. The agent may prepare the proposal, but it cannot authorize itself.
Read diagram source
flowchart LR
A[Draft exact action] --> B[Persist pending review]
B --> C{Authorized reviewer decides}
B -->|Expiry reached| E[Stop or hand off]
C -->|Approve| D[Revalidate and execute safely]
C -->|Edit| A
C -->|Reject| E
The approval record contains tenant, job, action arguments, proposal version or hash, reviewer identity, decision time, and expiry. A changed recipient or amount must satisfy the applicable policy and authorization for the revised proposal. At execution, recheck current permissions and business preconditions. A permission held yesterday may have been revoked today.
Persist the wait so deployments and worker restarts do not lose it. Define timeout behavior: expire, route to another qualified reviewer or return a safe partial result; silence is not approval. The runtime should release compute while waiting when supported. LangGraph interrupts documents resume behavior and code re-execution around an interrupt.
Select the intervention point
| Pattern | Useful when | Limitation |
|---|---|---|
| Human directs each step | Expert judgment dominates | Slow and expensive |
| Approval before external action | Consequences are hard to reverse | Queue delays and reviewer mistakes |
| Review final draft before publication | Content can be inspected before exposure | Long drafts may be poorly reviewed |
| Escalation for exceptions | Routine low-risk work is well measured | Missed exceptions bypass the reviewer |
| Post-action audit | Effects are reversible and risk is acceptable | Cannot prevent the original harm |
Choose thresholds from impact and evidence. Token probability measures likelihood of text, not the probability that a refund is correct. A model saying “95% confident” is not a calibrated risk estimate. Useful escalation inputs include missing required evidence, conflicting sources, policy exceptions, out-of-distribution cases, and a risk model validated on held-out cases.
Make review possible
A good review screen answers: What will change? For whom? Why is it allowed? What evidence supports it? What remains uncertain? Can it be undone? It gives approve, reject, and edit options without making approval the thoughtless default.
Show a concise action summary and source evidence, not a claim to expose the model's true internal reasoning. Sensitive data should be visible only to qualified reviewers. A long generated rationale can persuade a tired person without establishing correctness.
Do not force cosmetic word edits as proof of attention. Use explicit review criteria, training, quality sampling, and controlled exercises. Secretly injecting errors into live consequential workflows can itself create harm and corrupt measurements.
Measure review latency, missed harms, unnecessary escalations, reviewer agreement, overturns and outcomes after approval. Reducing escalation rate is not success if harmful actions escape review.
Manager follow-ups
“Can approval be reused?” Only for the policy-defined operation and unchanged proposal within its validity window; recheck authorization.
“How do you avoid fatigue?” Remove low-value approvals only after measuring risk, improve evidence presentation, and staff the queue. Do not solve fatigue by hiding uncertainty.
“Can I time-travel to before a mistake?” You can restore a workflow checkpoint; external effects require reconciliation or compensation.
Close the page and explain how an approval survives a restart, expires safely, and cannot authorize a changed amount. Then compare with durable execution.
Interview questions with developed answers
Q1: How do you design human review without fatiguing operators?
Sample answer: I make each review necessary and reviewable. Escalation rules reflect consequence, uncertainty, and policy, and the interface shows the proposed change, authoritative evidence, and reason for escalation. I size staffing from arrival rate and handling time, monitor queue age, and route cases to the right expertise. I reduce repeated low-value approvals through better bounded automation where evidence supports it. A reviewer still needs access to details; compressing every case to one sentence can remove the information needed for a sound decision.
Follow-up: What if review arrivals exceed capacity? Restrict automation scope, defer eligible work, or add capacity rather than silently dropping review.
Q2: What is over-reliance in human-in-the-loop systems?
Sample answer: It is accepting the system's recommendation without sufficient independent judgment. It can arise from fluent explanations, past success, time pressure, or poor interface design. I show verifiable evidence and uncertainties, make rejection and correction practical, keep workload manageable, and audit outcomes and disagreements. I use disclosed training exercises to build skill. Requiring a meaningless edit before approval does not prove understanding, and undisclosed error traps in production can cause harm. Human review must be evaluated as a fallible part of the system.
Follow-up: Which metric helps? Reviewer error and disagreement rates alongside handling time, queue pressure, and downstream outcomes.
Q3: What must be stored while waiting for approval?
Sample answer: The exact proposal and version, relevant evidence references, request identity, required reviewer authority, expiry, and current workflow state. The approval event references that proposal. On resume, I verify that authorization and business prerequisites still hold. This prevents a restart or changed proposal from turning an old approval into permission for a new action. I also deduplicate events and preserve an audit trail. The stored state is the application's decision record, not a frozen copy of inaccessible model reasoning.
Follow-up: What if the reviewer changes the amount? Treat it as a new proposal requiring the appropriate validation and approval.
Q4: When should an agent escalate because it is uncertain?
Sample answer: I use signals that are validated for the task, such as missing evidence, conflicting records, failed checks, unfamiliar cases, or a calibrated classifier. I combine uncertainty with impact: a small drafting ambiguity and an uncertain payment need different treatment. Model statements such as “I am 95% confident” are not enough. I measure how escalation affects missed failures, false escalations, reviewer load, and user outcomes, then adjust the policy. A safe abstention should explain what information or authority is needed next.
Follow-up: Can a confident model still need approval? Yes; approval can be a business or authority requirement independent of confidence.
Q5: Does editing an old checkpoint undo a mistake?
Sample answer: It can change the application's future execution path, but it does not reverse outside effects already performed. If the agent sent an email or refunded money, that fact remains after restoring an older state. I inspect the external record, preserve operation IDs, and define any correction as a separate authorized action. State editing is useful for diagnosis and controlled recovery, but it needs safeguards against repeating completed work or concealing history. The reviewer must know what actually happened before choosing the next step.
Follow-up: What should the UI show? Completed effects and uncertain outcomes as well as the editable proposal.
60-second interview answer
Human review is a designed control, not a guarantee of safety. I choose which actions need advance approval based on impact, reversibility, and uncertainty. The reviewer sees the exact proposed change, supporting evidence, and important missing information. Approval is tied to that proposal and expires; the application rechecks permission before execution. I persist the wait, prevent duplicate execution, and define what happens if nobody responds. I measure reviewer errors and queue capacity as well as agent performance, because an overloaded approval queue becomes a rubber stamp.