Learnastra AI SYSTEM DESIGNAnup Rai

Concept · Understand the mechanism

Loop Engineering

By Anup Rai10 min readReviewed September 2026

An agent loop repeatedly assembles context, selects an action, executes permitted work, observes the result and decides whether to continue. The surrounding application code is often called the agent harness. “Loop engineering” describes designing and improving that control system; it is not a universally standardized set of maturity levels.

This chapter brings together planning, recovery, memory, evaluation and durable execution. A stronger model and a better harness can both improve results. Neither universally compensates for every weakness in the other.

Distinguish the things being improved

Concern Main question Example change
Prompt What instruction does one call receive? Clarify the required output
Context What evidence and state does it receive? Retrieve the relevant schema revision
Run control What happens between calls? Bound retries and enforce permissions
Invocation What starts a run? Authenticated request or deduplicated event
Improvement process How is a new version evaluated and released? Compare a tool-schema change on held-out cases

These concerns interact, but do not have to form nested autonomous loops. Scheduled execution does not imply self-improvement, and an automatic critic does not automatically produce a better release.

The core research ideas also differ. ReAct interleaves reasoning and actions with observations. Reflexion uses verbal feedback retained across attempts rather than ordinary weight updates. Planning and parallel function calling can change when observations are needed. These are related techniques, not a proven single evolutionary chain from one system to all later agents.

Design a bounded task before choosing a loop

Example task: update documentation for a changed API schema and prepare a reviewable patch.

Functional requirements

  1. Identify changed endpoints against a specified schema revision.
  2. Update relevant examples and descriptions.
  3. Validate links, example syntax and agreement with the schema.
  4. Produce a diff with evidence and unresolved questions.
  5. Stop without publishing unless publication is authorized.

Non-functional requirements

  1. Preserve unrelated edits and keep work in the permitted scope.
  2. Bound execution time, model/tool usage and parallel work.
  3. Retain enough progress and artifact identity to resume safely.
  4. Check completion against the current artifacts, not an earlier version.
  5. Record failed, blocked, cancelled and partial outcomes accurately.

A fixed script may handle deterministic schema substitutions. Add adaptive model decisions for ambiguous explanations or exceptions. The point of a loop is to use intermediate evidence when the next useful action is not known in advance.

Implement the control boundary

Architecture / visual model
flowchart TD T[Authorized trigger and goal] --> L[Load current task state] L --> B{Deadline and budget permit work?} B -->|No| Z[Record bounded partial outcome] B -->|Yes| C[Assemble relevant context] C --> M[Model proposes next action or completion] M -->|Action proposal| P{Validate and authorize proposal} P -->|Denied or invalid| E[Classify error and remaining recovery] P -->|Permitted action| X[Execute with timeout and operation identity] X --> O[Record observation and artifact revision] M -->|Completion claim| V[Check required outcomes and conduct] O --> V V -->|Complete| D[Record verified completion] V -->|Incomplete with useful next work| N[Check progress and update plan] N --> B E -->|Recoverable within scope| N E -->|Blocked| H[Report evidence and needed decision]
Read diagram source
flowchart TD
    T[Authorized trigger and goal] --> L[Load current task state]
    L --> B{Deadline and budget permit work?}
    B -->|No| Z[Record bounded partial outcome]
    B -->|Yes| C[Assemble relevant context]
    C --> M[Model proposes next action or completion]
    M -->|Action proposal| P{Validate and authorize proposal}
    P -->|Denied or invalid| E[Classify error and remaining recovery]
    P -->|Permitted action| X[Execute with timeout and operation identity]
    X --> O[Record observation and artifact revision]
    M -->|Completion claim| V[Check required outcomes and conduct]
    O --> V
    V -->|Complete| D[Record verified completion]
    V -->|Incomplete with useful next work| N[Check progress and update plan]
    N --> B
    E -->|Recoverable within scope| N
    E -->|Blocked| H[Report evidence and needed decision]

The model's completion claim is input to the controller, not the sole authority. Conversely, the controller may stop a timed-out or cancelled run even if the model keeps requesting tools. A final natural-language response is not a necessary prerequisite for a valid stop.

Component Contract
Trigger Identity, scope, deduplication key and reason to run
Task state Goal, constraints, versions, results and pending work
Context builder Selected evidence within a known budget
Tool gateway Validated arguments, current authorization and limits
Result recorder Completed, failed or unknown outcomes with provenance
Verifier Observable acceptance criteria and calibrated judgment where needed
Controller Stop states, retry rules, deadlines and no-progress handling

A structured tool response can distinguish success, business rejection, incomplete result and unknown write outcome. Reducing everything to SUCCESS or FAILED loses the ambiguity that matters most after a timeout.

Compare useful loop patterns

Pattern Mechanism Benefit to test Main risk/cost
ReAct-style interaction Choose next action after an observation Adapt to unexpected evidence Repeated context and serial latency
Reflexion-style retry Use feedback from an earlier attempt Avoid repeating a diagnosed mistake Incorrect lessons and extra attempts
Plan and execute Execute a proposed dependency structure Reduce repeated planning Stale assumptions
ReWOO Plan with references to observations that are obtained later Reduce repeated model prompting Plan may need repair when observations differ
LLMCompiler Plan function calls and dispatch ready dependencies Parallelize independent calls Incorrect dependencies or expensive fan-out
Generator and evaluator Produce, grade and revise an artifact Improve a measurable quality criterion Correlated judge errors and over-optimization
Coordinator and workers Delegate bounded independent investigations Separate contexts and reduce elapsed time Coordination, duplicate work and integration failures

ReWOO separates planning from external observations, while LLMCompiler organizes function execution around dependencies and parallelism. Their reported benchmark savings do not establish fixed cost or latency rankings for every workload. A tool may require no model call; one model response may propose several tools.

A fresh-context approach starts a new model interaction from durable task records and selected evidence. It can reduce irrelevant history, but a poor handoff can lose exceptions, duplicate completed work or repeat an old mistake. Preserve source references and completed operation IDs, and test continuity before treating context resets as an improvement.

Enforce budgets before dispatch

  1. Set a deadline and explicit maximum allocations for the task.
  2. Reserve capacity before starting a call or child task.
  3. Apply provider/tool limits to the dispatched operation.
  4. Settle actual usage and release unused reservations.
  5. Propagate cancellation and reconcile any in-flight effects.

Enforcement must live in trusted runtime or service code that the untrusted workload cannot modify or bypass. It does not require a physically separate service for every deployment. A prompt asking the model to spend less is not an enforced cap.

Suppose a task has 60,000 tokens allocated. Planning uses 6,000; three workers each receive a 12,000-token reservation; final integration reserves 10,000. That leaves 8,000 tokens for recovery or other allowed work. These token counts are an illustrative accounting unit, not a dollar budget: input, output, cached input and different model rates must be priced separately.

If every worker independently checks the same remaining balance and then spends it, concurrent calls can overshoot. Use atomic reservation or equivalent coordination. A strict ceiling also requires a bound on each call's maximum charge, accounting for in-flight work and billing granularity. A monitor that reports usage after completion is useful telemetry but may not enforce a hard cap.

Limit Purpose Avoid
Total spend/tokens Bound the whole task Ignoring child tasks and failed calls
Per-call output/time Limit one expensive operation Waiting for an unbounded call to finish
Concurrent calls Contain bursts and downstream pressure Equating parallelism with free capacity
Retry count/budget Bound recovery amplification Nested invisible SDK retries
Deadline Bound elapsed time Restarting the clock after each retry
Progress condition Stop repeated ineffective work Treating all repeated polling as failure

High token throughput is not by itself evidence of a runaway loop; a healthy parallel job can be expensive. Alert on spending relative to the task allocation and useful progress. There is no universal 4,000-token-per-minute or thirty-turn failure threshold.

Define honest stop states

State Required evidence User-facing meaning
Succeeded Required outcome and conduct checks passed Completed within the specified scope
Blocked Missing input, permission or dependency Cannot proceed until the named condition changes
Failed A known unsuccessful outcome with no permitted recovery The requested work did not complete
Budget exhausted Limit reached with recorded progress Partial work is available; completion is unproven
Cancelled Cancellation accepted; outstanding effects accounted for Further work stopped where possible
Needs reconciliation An external effect remains uncertain Do not claim success or repeat blindly

Success criteria can combine deterministic checks and human judgment. A design review or writing task can have a rubric, a bounded number of revisions and a human acceptance step; it does not need to be rejected merely because aesthetic quality lacks an exact Boolean formula.

For the documentation patch, “all tests pass” is necessary only for the selected relevant tests and is not sufficient by itself. The patch could remove examples or disable a test. Check required content, scope and the actual diff. A coverage percentage is a proxy, not proof that the right behavior is covered.

Detect lack of progress without inventing magic thresholds

Record normalized actions, artifact versions, error categories and completed subgoals. Look for repeated failures with unchanged causes, cycling between states or repeated plans without new evidence.

For example, three status polls can be normal while an asynchronous job runs. Three edits that alternate a field between the same two values without resolving validation may be an oscillation. The response depends on the task: wait under a deadline, repair the cause, revise the plan or stop.

Plan-text similarity alone is weak evidence. A correct plan may remain unchanged while useful work completes. A substantially reworded plan may repeat the same failing strategy. Tune progress rules on labeled traces and inspect both false stops and missed loops.

Preserve context without treating summaries as truth

Technique Helps with Failure to check
Explicit task state Recover goal, versions and completed work State is stale or unauthenticated
Selective retrieval Fit relevant evidence into context Missing crucial exceptions
Summary/compaction Reduce repeated history Lost uncertainty, scope or identifiers
External artifact references Avoid copying large outputs every turn Broken references or access changes
Separate worker contexts Isolate independent investigations Lost shared dependencies
Stable eligible prefixes Improve computational cache reuse Preserving obsolete instructions for a cache hit

“Context rot” is an informal description of quality degradation with poorly managed context. Measure accuracy against length, distraction and evidence placement rather than assuming every long context fails. A larger window can help some tasks; selecting and preserving the right information still matters. Research such as Lost in the Middle shows why evidence placement deserves evaluation.

Follow the chosen provider's documented message-continuation contract for any opaque state. Do not invent a universal requirement to expose or preserve private reasoning text. Persist observable task state and results separately from model-specific continuation data.

Verify with evidence appropriate to the task

Verifier Strength Limit
Tests, schema checks and recomputation Reproducible checks of explicit properties Tests can be incomplete or expensive
Authoritative service lookup Establishes external state under its contract Read freshness and availability matter
Calibrated model grader Scales semantic judgment Can be biased, injected or wrong
Qualified human review Handles context and ambiguous tradeoffs Capacity, disagreement and fatigue

A separate model is not necessarily an independent verifier. It may share the generator's mistaken assumptions. For a numerical total, recomputation is more appropriate than asking another model whether it looks right. For an external write, inspect the receiver's record.

The 2023/2024 study Large Language Models Cannot Self-Correct Reasoning Yet found limitations in intrinsic self-correction under its tested settings. It is evidence to evaluate feedback quality, not a timeless claim that no model can ever revise its own answer. Useful feedback and outcome checks matter more than the number of critic personas.

Improve the harness through controlled releases

  1. Classify failures from representative, authorized traces.
  2. Identify the likely cause: model capability, context, tools, permissions, state or control flow.
  3. Make a targeted change and state the expected improvement.
  4. Compare with the baseline using repeated tasks and appropriate uncertainty estimates.
  5. Check held-out cases, regressions, cost and severe failure gates.
  6. Roll out within a bounded scope and retain rollback capability.

Do not automatically rewrite and publish the harness after every failed task. A reflection can be a useful hypothesis, but promoting it requires evaluation. Replacing a stable model-mediated step with ordinary code can reduce cost and variance once the contract is understood.

Track acceptable task success, repeated-trial consistency, full cost per acceptable outcome, tail latency, unknown writes, no-progress stops and human review demand. The suitable number of evaluation trials depends on the uncertainty and effect size; three to six runs is not a universal statistical guarantee.

Interview practice

Q1: What is the harness responsible for?

Context assembly, permitted execution, state recording, budgets, progress checks and termination. The model proposes actions; trusted code enforces the applicable boundaries. Model quality still affects whether the proposed plan is useful.

Q2: When is a loop unnecessary?

When a fixed, known transformation or workflow meets the requirements without adaptive choices. A single model call may be enough for one-shot drafting. Add a loop when intermediate observations materially affect the next action.

Q3: How do you stop a tool from being retried hundreds of times?

Classify its failure, enforce one shared retry/deadline budget and detect unchanged causes. Fix ambiguous tool results. Preserve unknown write outcomes instead of retrying them as ordinary read failures.

Q4: Must a different agent verify every result?

No. Use the strongest relevant evidence: deterministic checks, authoritative state or calibrated human/model judgment. A second agent can help but is not automatically independent or correct.

Q5: Can a new context solve every long-run failure?

No. It may remove distracting history, but can also lose requirements and repeat work. Build the new context from trustworthy task records and evidence, then test continuity and outcome quality.

Q6: How do you cap a parallel run's spending?

Reserve allocations atomically before dispatch, bound individual calls and include all descendants and retries. Settle actual use and account for in-flight work during cancellation. Monitoring total usage after the fact alone is not a hard limit.

Q7: What is the danger of a perfect-looking success metric?

The agent may satisfy the proxy while violating the goal, such as deleting a failing test. Evaluate intended outcomes, scope and prohibited conduct separately, and protect the grader from the agent's modifications.

Q8: How would you close this design?

Explain the loop's purpose, trusted execution boundary, durable state, stop conditions and verification evidence. Compare quality, cost and latency with the simpler baseline, then identify which failure category the next change is intended to improve.

Final notes

Recall card: Observe → choose → authorize → act → record → verify → continue or stop. Reliable operation comes from explicit contracts and measured behavior, not from adding an unlimited number of iterations.

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← Durable Execution for Long-Running Agents
NEXT LESSONMemory Architectures →

Explore the diagram