Learnastra AI SYSTEM DESIGNAnup Rai

Complete design interview

Design an Agent That Produces a Reviewable Code Change

By Anup Rai11 min readReviewed September 2026

Interview problem: turn a scoped repository task into a complete proposed patch, demonstrate what was checked, and publish it only through the authorized review workflow.

A coding agent repeatedly observes repository state, chooses an allowed action, examines the result and decides what to do next. It differs from an inline code assistant because a task may span files, run tools and survive worker failures. “Autonomous” describes execution within a boundary; it does not grant unlimited repository or production access.

This is a hypothetical Learnastra interview scenario. Costs and workload targets are illustrative. The central engineering problem is preserving the relationship between the user's request, the exact changed artifact, its validation evidence and authorization to publish.

1. Clarify the contract

Use a concrete task: “Add rate limiting to the API.” Before editing, clarify the limit, identity key, shared versus per-instance enforcement, burst allowance, exemptions, response behavior and operation when the limiter store fails. Otherwise a syntactically correct patch can implement the wrong product behavior.

Functional requirements

  1. Accept a task with repository, base revision, permitted paths/tools and acceptance criteria.
  2. Inspect a codebase with 1,000 or more files without assuming the whole repository fits in context.
  3. Produce a plan and a cumulative multi-file diff in an isolated workspace.
  4. Run relevant checks, retain their evidence, and attempt bounded repairs.
  5. Return useful partial work when time, attempts, spend or permissions stop the task.
  6. Support cancellation, recovery and review comments without losing task history.
  7. Require human approval of the exact proposed change before committing/publishing in this scenario; merging remains a separately authorized action.

Nonfunctional requirements

  1. Keep unrelated developer changes intact and enforce repository/tenant isolation.
  2. Set attempt, token, compute, wall-clock and dollar limits outside the model.
  3. Target model/tool expenditure below $0.50 per admitted small task; measure cost per accepted change separately.
  4. Propose p95 completion below ten minutes for the scoped small-task class. Larger migrations need a different queue and contract.
  5. Persist enough state to recover after a worker crash without replaying uncertain external side effects blindly.
  6. Report failed, skipped, unavailable and passed checks distinctly; “tests pass” does not prove absence of regressions.
  7. Keep repository data and secrets within the approved processing/retention boundary.

Out of scope initially: automatic production rollout, arbitrary credential use, unlimited dependency upgrades, and rewriting a repository merely because a broad task could be interpreted that way.

2. Estimate the worker problem

Suppose the product receives 10,000 admitted tasks/day, with a six-minute mean active runtime and a fivefold peak over the 24-hour mean.

Quantity Calculation Planning result
Mean task arrivals 10,000 / 86,400 0.116/s
Peak task arrivals 5 × 0.116 0.579/s
Active tasks at peak 0.579 × 360 seconds About 208
Allocated memory at 4GB/task 208 × 4GB About 832GB
Sandbox work/day 10,000 × 6 minutes 1,000 sandbox-hours

The memory estimate is an allocation scenario, not a recommendation to reserve that much for every task. Provision from measured language/build workloads, task residence, concurrency and headroom. A provider's token quota can become the bottleneck before CPU does. Admission control must consider both.

3. Baseline: one model and a controlled edit–test loop

Start with a single task worker, an isolated checkout, a model that can read/edit through bounded tools, and the repository's validation commands. Give the model actionable failures and a limited chance to repair them. Return a full diff and evidence for review.

Separate planning, editing and reviewing as stages. They need not be separate models or permanent agents. Add parallel specialists only when measured gains justify the coordination and conflicting-edit risks. See agent design patterns.

Baseline flaw Improvement Benefit Cost or limitation
Broad instruction produces the wrong behavior Explicit acceptance contract and plan Testable scope Clarification effort
Model rereads the entire repository Layered retrieval and versioned context Less irrelevant context Index freshness and missed dependencies
Repair loop repeats the same failure Normalized failure signature and bounded attempts Stops unproductive work May stop a repairable task early
Worker crash loses the task Durable states and isolated checkpoints Recovery without starting over More state/lease management
“Passed” belongs to an earlier patch Artifact-bound validation evidence Honest review status Rechecking after edits
Agent makes PR twice after a timeout Stable publication intent and reconciliation Avoids duplicate external actions Host-specific integration

4. Detailed task lifecycle

Architecture / visual model
flowchart TD U[Task and scoped authorization] --> ADMIT[Validate request and reserve budget] ADMIT --> JOB[(Durable task, base revision and state version)] JOB --> LEASE[Worker lease and attempt fence] LEASE --> WS[Isolated snapshot of authorized repository state] WS --> PLAN[Acceptance contract and change plan] PLAN --> CONTEXT[Retrieve current symbols and source] CONTEXT --> EDIT[Bounded edit and tool execution] EDIT --> CHECK[Required checks tied to artifact hash] CHECK --> RESULT{Checks and task criteria satisfied?} RESULT -->|No, budget remains| REPAIR[Specific failure feedback] REPAIR --> EDIT RESULT -->|No, stop condition| PARTIAL[Partial diff, limitations and recovery record] RESULT -->|Yes| REVIEW[Full diff and validation for human review] REVIEW -->|Changes requested| FOLLOW[Scoped follow-up against current artifact] FOLLOW --> CONTEXT REVIEW -->|Approved artifact| PUBLISH[Restricted publisher with stable intent] PUBLISH --> HOST[Commit and PR after base/version check] HOST --> CI[Repository CI and separate merge policy]
Read diagram source
flowchart TD
    U[Task and scoped authorization] --> ADMIT[Validate request and reserve budget]
    ADMIT --> JOB[(Durable task, base revision and state version)]
    JOB --> LEASE[Worker lease and attempt fence]
    LEASE --> WS[Isolated snapshot of authorized repository state]
    WS --> PLAN[Acceptance contract and change plan]
    PLAN --> CONTEXT[Retrieve current symbols and source]
    CONTEXT --> EDIT[Bounded edit and tool execution]
    EDIT --> CHECK[Required checks tied to artifact hash]
    CHECK --> RESULT{Checks and task criteria satisfied?}
    RESULT -->|No, budget remains| REPAIR[Specific failure feedback]
    REPAIR --> EDIT
    RESULT -->|No, stop condition| PARTIAL[Partial diff, limitations and recovery record]
    RESULT -->|Yes| REVIEW[Full diff and validation for human review]
    REVIEW -->|Changes requested| FOLLOW[Scoped follow-up against current artifact]
    FOLLOW --> CONTEXT
    REVIEW -->|Approved artifact| PUBLISH[Restricted publisher with stable intent]
    PUBLISH --> HOST[Commit and PR after base/version check]
    HOST --> CI[Repository CI and separate merge policy]

The orchestration service owns state transitions and budgets. The sandbox executes repository-controlled code. The publisher holds narrowly scoped credentials outside the sandbox. The model proposes actions; it does not mint authorization or mark itself approved.

API and records

POST /coding-tasks
{request_id, repository_id, base_revision, task, acceptance_criteria,
 path_scope, limits} → task_id

GET /coding-tasks/{id}
→ state, progress, artifact_hash?, complete_diff?, checks[], limitations[], usage

POST /coding-tasks/{id}/approval
{artifact_hash, base_revision, authorized_action} → approval_id

POST /coding-tasks/{id}/cancel
→ cancellation_requested
Record Essential fields
Task Request key, trusted owner/repository scope, base snapshot, criteria, budget and state version
Attempt Worker lease/fence, plan, tools, start/end, spend and failure signature
Artifact Cumulative patch, complete changed-file manifest, content hashes and base revision
Check Command/configuration, tested artifact hash, status, exit/result metadata and restricted logs
Approval Approver, exact artifact/base, permitted publication action, expiry/revocation
Publication intent Stable task/artifact key, branch/commit/PR identifiers and confirmed or uncertain outcome

Use optimistic state transitions or equivalent serialization so two workers cannot independently publish the same task. Lease expiry alone does not stop an old worker: tools and shared writes need a current attempt fence, or access must be revoked when the worker is replaced.

5. Retrieve code by evidence, not a fixed top-ten rule

Start with paths, exact symbols, errors and configuration references. Semantic summaries help with conceptual queries, but a language server/compiler index often resolves definitions and references more precisely. Tree-sitter provides incremental parsing and concrete syntax trees; extracting a symbol graph requires additional language-specific logic. A compiler may also build an abstract syntax tree (AST) and resolve types. Syntax structure alone cannot fully resolve dynamic dispatch or runtime dependency injection.

Architecture / visual model
flowchart LR Q[Task, symbol or failing test] --> FIND[Exact search, paths and optional semantic summaries] FIND --> CAND[Candidate files at a known revision] CAND --> GRAPH[Definitions, imports, references and dependency context] GRAPH --> READ[Read actual source, tests and configuration] READ --> FIT[Tokenizer-aware context with output reserve] FIT --> MODEL[Plan or edit] MODEL -->|Need more evidence| FIND
Read diagram source
flowchart LR
    Q[Task, symbol or failing test] --> FIND[Exact search, paths and optional semantic summaries]
    FIND --> CAND[Candidate files at a known revision]
    CAND --> GRAPH[Definitions, imports, references and dependency context]
    GRAPH --> READ[Read actual source, tests and configuration]
    READ --> FIT[Tokenizer-aware context with output reserve]
    FIT --> MODEL[Plan or edit]
    MODEL -->|Need more evidence| FIND
  1. Exclude unauthorized/generated/vendor paths according to the task policy.
  2. Pin retrieval to the working snapshot, including authorized uncommitted inputs.
  3. Find likely files, then inspect their real contents before editing.
  4. Follow relevant callers, tests, shared middleware and configuration.
  5. Invalidate or update summaries/symbols after edits; a stale summary is a hint, not authority.
  6. Fit a bounded context and request more evidence when needed instead of fabricating missing interfaces.

For rate limiting, inspect route registration, authentication, deployment replicas, existing caches and tests. A local in-memory counter in each API replica does not implement a shared global limit. The agent should discover that requirement before generating code.

6. Sandboxing and authority

Use a disposable environment with restricted filesystem mounts, outbound network, process resources and lifetime. E2B is one current managed sandbox option; its documentation describes Linux VMs and supports persistence. Do not assume a sandbox is always a temporary container or automatically erased when an application disconnects. Configure cleanup, storage and resume behavior explicitly.

Boundary Enforced behavior Why it matters
Repository snapshot Only task-authorized inputs; path/symlink controls Protects other repositories and local work
Network/dependencies Approved registries/destinations and policy checks Install scripts and tests can execute arbitrary code
Secrets No production credentials; broker narrowly scoped actions Isolation cannot protect a secret handed to hostile code
Tool outputs Bounded bytes/time and safe rendering Huge logs and control sequences can harm the controller/UI
Publication Separate trusted service validates approval and artifact Repository comments cannot authorize push/merge
Cleanup Kill processes and enforce storage lifetime Canceling a model call does not terminate its tools

Treat code comments, build output, issue text and fetched documents as untrusted task data. Instructions inside them cannot expand scope. See agentic security and sandboxing.

7. Bounded repairs and durable recovery

A repair should use the failing test/diagnostic, the attempted change and the original acceptance criteria. Keep failures structured rather than recursively embedding the entire previous prompt. Stop on exhausted budget, deadline, repeated nonprogress or attempt limit.

This executable orchestration example uses application adapters. It returns all accumulated task edits and rolls back an interrupted attempt only inside the task's isolated workspace:

import asyncio

async def execute_with_retry(task, runtime, workspace, budget,
                             max_attempts=3, timeout_seconds=120):
    baseline = await workspace.checkpoint()
    previous_error = None
    feedback = None
    reason = "attempt_limit"
    try:
        async with asyncio.timeout(timeout_seconds):
            for attempt in range(max_attempts):
                if not budget.can_start_attempt():
                    reason = "budget_limit"
                    break
                checkpoint = await workspace.checkpoint()
                try:
                    checks = await runtime.edit_and_check(
                        task=task, feedback=feedback, budget=budget
                    )
                except BaseException:
                    await runtime.stop_tools()
                    await workspace.restore(checkpoint)
                    raise
                if checks.passed:
                    reason = "ready_for_human_review"
                    break
                if checks.failure_signature == previous_error:
                    reason = "repeated_failure"
                    break
                previous_error = checks.failure_signature
                feedback = checks.failures
    except TimeoutError:
        reason = "time_limit"
    return {
        "status": reason,
        "diff": await workspace.diff_since(baseline),
        "checks": await workspace.latest_check_results(),
        "spent_usd": budget.spent_usd,
    }

The 120-second limit demonstrates an active repair-loop cap; it is not the earlier end-to-end ten-minute SLO. Admission, queue waiting, setup, checkpoints and cleanup need their own bounded deadlines. Python cancellation is cooperative: a process/sandbox supervisor must enforce hard termination and cleanup if a tool or coroutine fails to stop. Unexpected infrastructure errors propagate to the durable controller, which records the failure and recoverable artifact.

The adapters must reserve maximum allowed spend before each model/tool operation, settle actual usage, retain charges for unsuccessful work, and bind every check to its artifact. After rollback, old check results remain historical evidence and cannot certify the restored artifact. Empty/missing results do not count as a pass.

A normalized failure signature should include failing checks and meaningful error classes while excluding volatile paths/timestamps. A repeated signature suggests nonprogress; different signatures do not prove progress. Also compare artifact changes and fulfilled criteria.

Crash versus uncertain side effect

Interruption Recovery
Worker dies before a sandbox edit finishes Revoke its lease/access; recover the last known artifact or discard that isolated attempt
Test process hangs Supervisor terminates it; mark timeout, not failed application behavior or passed
Model times out after billing may have occurred Record uncertain usage and reconcile; do not release reserved budget as if the call were free
PR creation times out Look up the stable publication intent/branch/head before retrying
Base branch moves after approval Detect the mismatch, merge/rebase under policy, rerun affected checks and obtain approval for the new artifact/base

See error handling and recovery for durable state and uncertain outcomes.

8. Verification and review must test the requested behavior

Capture the baseline test state so pre-existing failures are visible. Do not simply excuse every later failure as “already broken”: compare which tests failed and whether the patch worsened them.

For the rate-limiting task, exercise:

  1. Requests below and above the configured limit, including the window boundary.
  2. Correct identity separation and resistance to spoofed forwarding headers.
  3. Shared enforcement across two API replicas and concurrent increments.
  4. The specified datastore-outage policy and permitted exemptions.
  5. Expected response codes/headers and unaffected authentication behavior.
  6. Relevant regression, type, dependency and security checks on the final patch.

A scanner such as Semgrep detects covered patterns; it does not prove business correctness or complete security. Test edits also need review: deleting an assertion can improve the pass rate while weakening the product.

For a 20-file task, record dependencies as a directed acyclic graph and process prerequisites first. Add shared middleware before wiring its consumers. A repaired prerequisite can invalidate earlier downstream results; rerun affected checks and the required final suite. Parallel edits need isolated branches/workspaces and a conflict-aware integration stage.

9. Model and cost decisions

Claude Sonnet 5 is a current candidate for repeated edit-and-test turns; compare alternatives on accepted patches in representative repositories. More capable candidates such as Claude Opus 5.5 may justify their cost for hard tasks, but escalation must fit a separate allowance or require an expanded task budget. Do not infer a model's quality as a fixed fraction of a benchmark leader. Current Claude models.

For an illustrative Sonnet 5 task at $2 input/$10 output per million tokens, uncached standard processing:

Phase Input / billed output tokens Model cost or allowance
Planning 8,000 / 2,000 $0.036
Retrieval embedding Allocated example budget $0.001
Coding per attempt 15,000 / 3,000 $0.060
Sandbox, tests, storage Hypothetical per-task allowance $0.030
At 1.5 mean attempts $0.036 + $0.001 + 1.5 × $0.060 + $0.030 $0.157
At three attempts Same inputs, three coding turns $0.247

The published model rates support the token calculation. Infrastructure and embedding allowances require actual workload pricing; testing is not free. Reasoning, growing tool transcripts and repeated context belong in billed tokens. The table is an average scenario, not a worst-case reservation.

If 60% of admitted tasks become accepted changes, the $0.157 mean cost yields $0.157 / 0.60 ≈ $0.262 per accepted change before human review and omitted overhead. A ten-minute review at an assumed $100/hour costs $16.67—much more than the model calls. Evaluate saved developer time and defects rather than optimizing only token price.

10. Evaluation, rollout and operational ownership

Evaluate the entire model/tool/harness configuration on repository tasks with known acceptance checks. Pin task snapshots, dependency environments and evaluation versions. Keep private holdouts and check for benchmark contamination; a public leaderboard cannot establish product readiness.

Measure accepted and retained patches, requested behavior, regressions/security findings, reviewer time, partial-result usefulness, p50/p95 completion, token/tool spend, repeated failures and cancellation cleanup. Report task mix and denominator; selected successful demos are not a success rate.

Pilot read-only explanations, then proposed edits, then reviewed publication. Operations owns queue and sandbox capacity; the platform team owns leases, budgets and credential boundaries; repository owners define meaningful checks and merge policy. Roll back model/prompt/tool configurations while retaining exact artifacts and audit records for already completed tasks.

Interview follow-ups

1. Why not start with many specialist agents? Stages give checkpoints without multi-agent coordination cost. Add specialists after measuring a real gap, and isolate parallel edits so agents cannot silently overwrite one another.

2. How do you preserve unrelated work? Capture authorized input changes in the base snapshot, work in an isolated checkout, and return a cumulative patch tied to that base. Restore only task-owned checkpoints; application to the user's workspace requires a version/conflict check.

3. What prevents an infinite repair loop? Runtime-enforced attempts, time, tokens and spend, plus repeated-failure/nonprogress detection. The model's promise to stop is not enforcement.

4. Is passing CI enough to commit? The scenario also requires exact-artifact human approval. CI provides evidence within its coverage; the acceptance criteria, changed tests and unresolved limitations still need review.

5. What if the process dies after creating a PR? The durable publication record may have an uncertain outcome. Reconcile the branch/head or provider identifier before retrying, using a stable intent to avoid duplicate PRs.

6. How would you defend the $0.50 budget? Reserve worst-case operation costs, enforce output/tool limits and stop before the next action would exceed the allowance. Report actual cost per admitted task and per accepted change; do not claim a favorable average enforces every task's cap.

60-second interview answer

I would run a bounded edit-and-test workflow in an isolated repository snapshot. The task starts with acceptance criteria, scoped tools and a reserved budget. Retrieval follows actual symbols and dependencies, while durable checkpoints preserve a cumulative patch across repairs and failures. Checks and human approval refer to the exact artifact, and a separate publisher reconciles uncertain external actions. I would measure accepted changes, defects, review time and total cost, then expand autonomy only where the evidence supports it.

Remember: Scope → Snapshot → Edit → Verify → Approve → Publish.

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← Design a Fresh Market-Intelligence Search Service
NEXT LESSONDesign a Multi-Tenant Contract-Analysis Platform →

Explore the diagram