Interview problem: turn a scoped repository task into a complete proposed patch, demonstrate what was checked, and publish it only through the authorized review workflow.
A coding agent repeatedly observes repository state, chooses an allowed action, examines the result and decides what to do next. It differs from an inline code assistant because a task may span files, run tools and survive worker failures. “Autonomous” describes execution within a boundary; it does not grant unlimited repository or production access.
This is a hypothetical Learnastra interview scenario. Costs and workload targets are illustrative. The central engineering problem is preserving the relationship between the user's request, the exact changed artifact, its validation evidence and authorization to publish.
1. Clarify the contract
Use a concrete task: “Add rate limiting to the API.” Before editing, clarify the limit, identity key, shared versus per-instance enforcement, burst allowance, exemptions, response behavior and operation when the limiter store fails. Otherwise a syntactically correct patch can implement the wrong product behavior.
Functional requirements
- Accept a task with repository, base revision, permitted paths/tools and acceptance criteria.
- Inspect a codebase with 1,000 or more files without assuming the whole repository fits in context.
- Produce a plan and a cumulative multi-file diff in an isolated workspace.
- Run relevant checks, retain their evidence, and attempt bounded repairs.
- Return useful partial work when time, attempts, spend or permissions stop the task.
- Support cancellation, recovery and review comments without losing task history.
- Require human approval of the exact proposed change before committing/publishing in this scenario; merging remains a separately authorized action.
Nonfunctional requirements
- Keep unrelated developer changes intact and enforce repository/tenant isolation.
- Set attempt, token, compute, wall-clock and dollar limits outside the model.
- Target model/tool expenditure below $0.50 per admitted small task; measure cost per accepted change separately.
- Propose p95 completion below ten minutes for the scoped small-task class. Larger migrations need a different queue and contract.
- Persist enough state to recover after a worker crash without replaying uncertain external side effects blindly.
- Report failed, skipped, unavailable and passed checks distinctly; “tests pass” does not prove absence of regressions.
- Keep repository data and secrets within the approved processing/retention boundary.
Out of scope initially: automatic production rollout, arbitrary credential use, unlimited dependency upgrades, and rewriting a repository merely because a broad task could be interpreted that way.
2. Estimate the worker problem
Suppose the product receives 10,000 admitted tasks/day, with a six-minute mean active runtime and a fivefold peak over the 24-hour mean.
| Quantity | Calculation | Planning result |
|---|---|---|
| Mean task arrivals | 10,000 / 86,400 | 0.116/s |
| Peak task arrivals | 5 × 0.116 | 0.579/s |
| Active tasks at peak | 0.579 × 360 seconds | About 208 |
| Allocated memory at 4GB/task | 208 × 4GB | About 832GB |
| Sandbox work/day | 10,000 × 6 minutes | 1,000 sandbox-hours |
The memory estimate is an allocation scenario, not a recommendation to reserve that much for every task. Provision from measured language/build workloads, task residence, concurrency and headroom. A provider's token quota can become the bottleneck before CPU does. Admission control must consider both.
3. Baseline: one model and a controlled edit–test loop
Start with a single task worker, an isolated checkout, a model that can read/edit through bounded tools, and the repository's validation commands. Give the model actionable failures and a limited chance to repair them. Return a full diff and evidence for review.
Separate planning, editing and reviewing as stages. They need not be separate models or permanent agents. Add parallel specialists only when measured gains justify the coordination and conflicting-edit risks. See agent design patterns.
| Baseline flaw | Improvement | Benefit | Cost or limitation |
|---|---|---|---|
| Broad instruction produces the wrong behavior | Explicit acceptance contract and plan | Testable scope | Clarification effort |
| Model rereads the entire repository | Layered retrieval and versioned context | Less irrelevant context | Index freshness and missed dependencies |
| Repair loop repeats the same failure | Normalized failure signature and bounded attempts | Stops unproductive work | May stop a repairable task early |
| Worker crash loses the task | Durable states and isolated checkpoints | Recovery without starting over | More state/lease management |
| “Passed” belongs to an earlier patch | Artifact-bound validation evidence | Honest review status | Rechecking after edits |
| Agent makes PR twice after a timeout | Stable publication intent and reconciliation | Avoids duplicate external actions | Host-specific integration |
4. Detailed task lifecycle
Read diagram source
flowchart TD
U[Task and scoped authorization] --> ADMIT[Validate request and reserve budget]
ADMIT --> JOB[(Durable task, base revision and state version)]
JOB --> LEASE[Worker lease and attempt fence]
LEASE --> WS[Isolated snapshot of authorized repository state]
WS --> PLAN[Acceptance contract and change plan]
PLAN --> CONTEXT[Retrieve current symbols and source]
CONTEXT --> EDIT[Bounded edit and tool execution]
EDIT --> CHECK[Required checks tied to artifact hash]
CHECK --> RESULT{Checks and task criteria satisfied?}
RESULT -->|No, budget remains| REPAIR[Specific failure feedback]
REPAIR --> EDIT
RESULT -->|No, stop condition| PARTIAL[Partial diff, limitations and recovery record]
RESULT -->|Yes| REVIEW[Full diff and validation for human review]
REVIEW -->|Changes requested| FOLLOW[Scoped follow-up against current artifact]
FOLLOW --> CONTEXT
REVIEW -->|Approved artifact| PUBLISH[Restricted publisher with stable intent]
PUBLISH --> HOST[Commit and PR after base/version check]
HOST --> CI[Repository CI and separate merge policy]
The orchestration service owns state transitions and budgets. The sandbox executes repository-controlled code. The publisher holds narrowly scoped credentials outside the sandbox. The model proposes actions; it does not mint authorization or mark itself approved.
API and records
POST /coding-tasks
{request_id, repository_id, base_revision, task, acceptance_criteria,
path_scope, limits} → task_id
GET /coding-tasks/{id}
→ state, progress, artifact_hash?, complete_diff?, checks[], limitations[], usage
POST /coding-tasks/{id}/approval
{artifact_hash, base_revision, authorized_action} → approval_id
POST /coding-tasks/{id}/cancel
→ cancellation_requested
| Record | Essential fields |
|---|---|
| Task | Request key, trusted owner/repository scope, base snapshot, criteria, budget and state version |
| Attempt | Worker lease/fence, plan, tools, start/end, spend and failure signature |
| Artifact | Cumulative patch, complete changed-file manifest, content hashes and base revision |
| Check | Command/configuration, tested artifact hash, status, exit/result metadata and restricted logs |
| Approval | Approver, exact artifact/base, permitted publication action, expiry/revocation |
| Publication intent | Stable task/artifact key, branch/commit/PR identifiers and confirmed or uncertain outcome |
Use optimistic state transitions or equivalent serialization so two workers cannot independently publish the same task. Lease expiry alone does not stop an old worker: tools and shared writes need a current attempt fence, or access must be revoked when the worker is replaced.
5. Retrieve code by evidence, not a fixed top-ten rule
Start with paths, exact symbols, errors and configuration references. Semantic summaries help with conceptual queries, but a language server/compiler index often resolves definitions and references more precisely. Tree-sitter provides incremental parsing and concrete syntax trees; extracting a symbol graph requires additional language-specific logic. A compiler may also build an abstract syntax tree (AST) and resolve types. Syntax structure alone cannot fully resolve dynamic dispatch or runtime dependency injection.
Read diagram source
flowchart LR
Q[Task, symbol or failing test] --> FIND[Exact search, paths and optional semantic summaries]
FIND --> CAND[Candidate files at a known revision]
CAND --> GRAPH[Definitions, imports, references and dependency context]
GRAPH --> READ[Read actual source, tests and configuration]
READ --> FIT[Tokenizer-aware context with output reserve]
FIT --> MODEL[Plan or edit]
MODEL -->|Need more evidence| FIND
- Exclude unauthorized/generated/vendor paths according to the task policy.
- Pin retrieval to the working snapshot, including authorized uncommitted inputs.
- Find likely files, then inspect their real contents before editing.
- Follow relevant callers, tests, shared middleware and configuration.
- Invalidate or update summaries/symbols after edits; a stale summary is a hint, not authority.
- Fit a bounded context and request more evidence when needed instead of fabricating missing interfaces.
For rate limiting, inspect route registration, authentication, deployment replicas, existing caches and tests. A local in-memory counter in each API replica does not implement a shared global limit. The agent should discover that requirement before generating code.
6. Sandboxing and authority
Use a disposable environment with restricted filesystem mounts, outbound network, process resources and lifetime. E2B is one current managed sandbox option; its documentation describes Linux VMs and supports persistence. Do not assume a sandbox is always a temporary container or automatically erased when an application disconnects. Configure cleanup, storage and resume behavior explicitly.
| Boundary | Enforced behavior | Why it matters |
|---|---|---|
| Repository snapshot | Only task-authorized inputs; path/symlink controls | Protects other repositories and local work |
| Network/dependencies | Approved registries/destinations and policy checks | Install scripts and tests can execute arbitrary code |
| Secrets | No production credentials; broker narrowly scoped actions | Isolation cannot protect a secret handed to hostile code |
| Tool outputs | Bounded bytes/time and safe rendering | Huge logs and control sequences can harm the controller/UI |
| Publication | Separate trusted service validates approval and artifact | Repository comments cannot authorize push/merge |
| Cleanup | Kill processes and enforce storage lifetime | Canceling a model call does not terminate its tools |
Treat code comments, build output, issue text and fetched documents as untrusted task data. Instructions inside them cannot expand scope. See agentic security and sandboxing.
7. Bounded repairs and durable recovery
A repair should use the failing test/diagnostic, the attempted change and the original acceptance criteria. Keep failures structured rather than recursively embedding the entire previous prompt. Stop on exhausted budget, deadline, repeated nonprogress or attempt limit.
This executable orchestration example uses application adapters. It returns all accumulated task edits and rolls back an interrupted attempt only inside the task's isolated workspace:
import asyncio
async def execute_with_retry(task, runtime, workspace, budget,
max_attempts=3, timeout_seconds=120):
baseline = await workspace.checkpoint()
previous_error = None
feedback = None
reason = "attempt_limit"
try:
async with asyncio.timeout(timeout_seconds):
for attempt in range(max_attempts):
if not budget.can_start_attempt():
reason = "budget_limit"
break
checkpoint = await workspace.checkpoint()
try:
checks = await runtime.edit_and_check(
task=task, feedback=feedback, budget=budget
)
except BaseException:
await runtime.stop_tools()
await workspace.restore(checkpoint)
raise
if checks.passed:
reason = "ready_for_human_review"
break
if checks.failure_signature == previous_error:
reason = "repeated_failure"
break
previous_error = checks.failure_signature
feedback = checks.failures
except TimeoutError:
reason = "time_limit"
return {
"status": reason,
"diff": await workspace.diff_since(baseline),
"checks": await workspace.latest_check_results(),
"spent_usd": budget.spent_usd,
}
The 120-second limit demonstrates an active repair-loop cap; it is not the earlier end-to-end ten-minute SLO. Admission, queue waiting, setup, checkpoints and cleanup need their own bounded deadlines. Python cancellation is cooperative: a process/sandbox supervisor must enforce hard termination and cleanup if a tool or coroutine fails to stop. Unexpected infrastructure errors propagate to the durable controller, which records the failure and recoverable artifact.
The adapters must reserve maximum allowed spend before each model/tool operation, settle actual usage, retain charges for unsuccessful work, and bind every check to its artifact. After rollback, old check results remain historical evidence and cannot certify the restored artifact. Empty/missing results do not count as a pass.
A normalized failure signature should include failing checks and meaningful error classes while excluding volatile paths/timestamps. A repeated signature suggests nonprogress; different signatures do not prove progress. Also compare artifact changes and fulfilled criteria.
Crash versus uncertain side effect
| Interruption | Recovery |
|---|---|
| Worker dies before a sandbox edit finishes | Revoke its lease/access; recover the last known artifact or discard that isolated attempt |
| Test process hangs | Supervisor terminates it; mark timeout, not failed application behavior or passed |
| Model times out after billing may have occurred | Record uncertain usage and reconcile; do not release reserved budget as if the call were free |
| PR creation times out | Look up the stable publication intent/branch/head before retrying |
| Base branch moves after approval | Detect the mismatch, merge/rebase under policy, rerun affected checks and obtain approval for the new artifact/base |
See error handling and recovery for durable state and uncertain outcomes.
8. Verification and review must test the requested behavior
Capture the baseline test state so pre-existing failures are visible. Do not simply excuse every later failure as “already broken”: compare which tests failed and whether the patch worsened them.
For the rate-limiting task, exercise:
- Requests below and above the configured limit, including the window boundary.
- Correct identity separation and resistance to spoofed forwarding headers.
- Shared enforcement across two API replicas and concurrent increments.
- The specified datastore-outage policy and permitted exemptions.
- Expected response codes/headers and unaffected authentication behavior.
- Relevant regression, type, dependency and security checks on the final patch.
A scanner such as Semgrep detects covered patterns; it does not prove business correctness or complete security. Test edits also need review: deleting an assertion can improve the pass rate while weakening the product.
For a 20-file task, record dependencies as a directed acyclic graph and process prerequisites first. Add shared middleware before wiring its consumers. A repaired prerequisite can invalidate earlier downstream results; rerun affected checks and the required final suite. Parallel edits need isolated branches/workspaces and a conflict-aware integration stage.
9. Model and cost decisions
Claude Sonnet 5 is a current candidate for repeated edit-and-test turns; compare alternatives on accepted patches in representative repositories. More capable candidates such as Claude Opus 5.5 may justify their cost for hard tasks, but escalation must fit a separate allowance or require an expanded task budget. Do not infer a model's quality as a fixed fraction of a benchmark leader. Current Claude models.
For an illustrative Sonnet 5 task at $2 input/$10 output per million tokens, uncached standard processing:
| Phase | Input / billed output tokens | Model cost or allowance |
|---|---|---|
| Planning | 8,000 / 2,000 | $0.036 |
| Retrieval embedding | Allocated example budget | $0.001 |
| Coding per attempt | 15,000 / 3,000 | $0.060 |
| Sandbox, tests, storage | Hypothetical per-task allowance | $0.030 |
| At 1.5 mean attempts | $0.036 + $0.001 + 1.5 × $0.060 + $0.030 | $0.157 |
| At three attempts | Same inputs, three coding turns | $0.247 |
The published model rates support the token calculation. Infrastructure and embedding allowances require actual workload pricing; testing is not free. Reasoning, growing tool transcripts and repeated context belong in billed tokens. The table is an average scenario, not a worst-case reservation.
If 60% of admitted tasks become accepted changes, the $0.157 mean cost yields $0.157 / 0.60 ≈ $0.262 per accepted change before human review and omitted overhead. A ten-minute review at an assumed $100/hour costs $16.67—much more than the model calls. Evaluate saved developer time and defects rather than optimizing only token price.
10. Evaluation, rollout and operational ownership
Evaluate the entire model/tool/harness configuration on repository tasks with known acceptance checks. Pin task snapshots, dependency environments and evaluation versions. Keep private holdouts and check for benchmark contamination; a public leaderboard cannot establish product readiness.
Measure accepted and retained patches, requested behavior, regressions/security findings, reviewer time, partial-result usefulness, p50/p95 completion, token/tool spend, repeated failures and cancellation cleanup. Report task mix and denominator; selected successful demos are not a success rate.
Pilot read-only explanations, then proposed edits, then reviewed publication. Operations owns queue and sandbox capacity; the platform team owns leases, budgets and credential boundaries; repository owners define meaningful checks and merge policy. Roll back model/prompt/tool configurations while retaining exact artifacts and audit records for already completed tasks.
Interview follow-ups
1. Why not start with many specialist agents? Stages give checkpoints without multi-agent coordination cost. Add specialists after measuring a real gap, and isolate parallel edits so agents cannot silently overwrite one another.
2. How do you preserve unrelated work? Capture authorized input changes in the base snapshot, work in an isolated checkout, and return a cumulative patch tied to that base. Restore only task-owned checkpoints; application to the user's workspace requires a version/conflict check.
3. What prevents an infinite repair loop? Runtime-enforced attempts, time, tokens and spend, plus repeated-failure/nonprogress detection. The model's promise to stop is not enforcement.
4. Is passing CI enough to commit? The scenario also requires exact-artifact human approval. CI provides evidence within its coverage; the acceptance criteria, changed tests and unresolved limitations still need review.
5. What if the process dies after creating a PR? The durable publication record may have an uncertain outcome. Reconcile the branch/head or provider identifier before retrying, using a stable intent to avoid duplicate PRs.
6. How would you defend the $0.50 budget? Reserve worst-case operation costs, enforce output/tool limits and stop before the next action would exceed the allowance. Report actual cost per admitted task and per accepted change; do not claim a favorable average enforces every task's cap.
60-second interview answer
I would run a bounded edit-and-test workflow in an isolated repository snapshot. The task starts with acceptance criteria, scoped tools and a reserved budget. Retrieval follows actual symbols and dependencies, while durable checkpoints preserve a cumulative patch across repairs and failures. Checks and human approval refer to the exact artifact, and a separate publisher reconciles uncertain external actions. I would measure accepted changes, defects, review time and total cost, then expand autonomy only where the evidence supports it.
Remember: Scope → Snapshot → Edit → Verify → Approve → Publish.