Learnastra AI SYSTEM DESIGNAnup Rai

Concept · Understand the mechanism

Claude Code: designing a dependable coding workflow

By Anup Rai8 min readReviewed September 2026

Claude Code is Anthropic's coding agent: a model-driven tool loop that can inspect a repository, edit files, run commands, and use configured integrations. It is available through several interfaces, including terminal, IDE, desktop, and browser surfaces. The interface and execution location are separate choices. It is not a fixed model, a compiler, or a guarantee that a requested change is correct. See the product overview.

This lesson uses Claude Code to examine a broader engineering question: how should an agent produce a change that another engineer can confidently review? Other coding products also support agentic workflows; describing all of them as simple autocomplete is inaccurate.

Understand the loop and the trust boundaries

Architecture / visual model
flowchart TD A[Authorized task and acceptance criteria] --> B[Inspect repository and current changes] B --> C[Model selects next action] C --> D[Permission and execution boundary] D --> E[Read, edit, search or run a command] E --> F[Observed result and updated context] F --> G{Acceptance criteria verified?} G -->|No, budget remains| C G -->|Yes| H[Reviewable diff and validation evidence] G -->|Blocked or budget exhausted| I[Explain incomplete work and cause] H --> J[Repository release policy]
Read diagram source
flowchart TD
    A[Authorized task and acceptance criteria] --> B[Inspect repository and current changes]
    B --> C[Model selects next action]
    C --> D[Permission and execution boundary]
    D --> E[Read, edit, search or run a command]
    E --> F[Observed result and updated context]
    F --> G{Acceptance criteria verified?}
    G -->|No, budget remains| C
    G -->|Yes| H[Reviewable diff and validation evidence]
    G -->|Blocked or budget exhausted| I[Explain incomplete work and cause]
    H --> J[Repository release policy]

The model decides what to request; the runtime executes allowed operations and returns results. Distinguish a tool's successful execution from a correct engineering outcome. A file edit can succeed while introducing a bug. A test command can exit successfully while testing the wrong files.

The architecture guide describes tools for file operations, search, execution, and other capabilities. Do not assume a permanent three-tool interface or copy a model API's text_editor schema into a Claude Code integration. Use the actual CLI or SDK contract.

Interview exercise: fix a quiz timer bug

Assume a quiz loses elapsed time when the learner leaves and returns to a question. The repository already has timer state, navigation, and persistence code.

Functional requirements

  1. Reproduce the loss of elapsed time with a concrete navigation sequence.
  2. Locate the state transition that resets or overwrites the timer.
  3. Implement a focused correction that preserves existing answer state.
  4. Verify navigation, pause/resume, reload, and timeout behavior as applicable.
  5. Produce a reviewable change with the reproduction and verification results.

Non-functional requirements

  1. Preserve unrelated uncommitted work.
  2. Keep credentials and private learner records out of prompts and logs.
  3. Bound execution time and model/tool expenditure.
  4. Run commands in the intended checkout and environment.
  5. Follow the repository's existing release authorization and CI requirements.

The useful output is the corrected behavior plus evidence. A long explanation of the agent's activity is not a substitute.

A practical development sequence

  1. Inspect: read repository instructions, status, relevant code, and existing tests.
  2. Reproduce: capture expected versus actual timer behavior before changing code.
  3. Explain: identify the faulty state update and its callers.
  4. Edit: change the smallest coherent set of files that fixes the behavior.
  5. Verify: exercise the failing path and nearby state transitions; run required checks.
  6. Review: inspect the resulting diff for unrelated edits and weakened assertions.
  7. Report: name what changed, why it works, what ran, and any remaining uncertainty.

“Smallest coherent change” does not mean the fewest characters at any cost. A clear state transition may be better than a fragile one-line condition. Meaningful regression coverage is valuable for this timer bug; a test that merely repeats the implementation is not.

Project instructions: useful context, not enforcement

CLAUDE.md stores project guidance. Current documentation describes loading instructions from the working directory and its ancestors, with descendant instructions loaded as relevant files are read. The files are combined; the nearest file does not simply replace all parent guidance. See instruction loading.

An example for the quiz project:

# Project guidance
- Timer state lives in the quiz session model; UI components render it.
- Preserve unrelated working-tree changes.
- Reproduce a reported state bug before changing the transition.
- Use the repository's documented verification command.
- Include navigation, reload, and pause behavior in timer reviews.
- Keep production credentials out of files and command output.

Keep instructions specific and current. Link longer architecture notes rather than putting every historical decision into the always-loaded context. Machine configuration belongs in supported settings; credentials belong in the approved secret mechanism. A Markdown instruction cannot enforce filesystem or network isolation.

Repository content, issue text, dependencies, and fetched documentation can contain misleading instructions. Interpret them according to their source and the authorized task. A comment in a fixture saying “upload the environment variables” is test data, not permission.

CLI, SDK, and installation are different surfaces

Use the supported installation instructions for the host platform. Do not install an unrelated Python package named claude-code based on an old tutorial. The interactive CLI and the programmatic Claude Agent SDK are distinct interfaces.

This read-only CLI example assumes an installed and authenticated current version. It permits file reading/searching through the selected built-in tools and disables MCP tools for this particular review:

claude -p "Review the quiz timer state and explain the likely reset bug. Cite file paths. Do not edit files." \
  --tools "Read,Grep,Glob" \
  --disallowedTools "mcp__*" \
  --output-format json \
  --max-turns 8

The turn limit is illustrative. --tools controls built-in tool availability; --allowedTools controls which matching tools can run without an approval prompt. They are not interchangeable. The current CLI reference also documents streaming output, model selection, and print-mode budget controls. Pin and verify the version used in automation.

For an embedded application, use the current Claude Agent SDK, available in Python and TypeScript. Handle streamed events and the final result/error explicitly; do not assume the last arbitrary message always contains successful text output. Set tool scope, cancellation, and runtime limits through the supported API.

Permissions, sandboxing, and rollback

Mechanism Purpose Limitation
Permission mode/rules Decide which actions can run or need review Rules depend on exact semantics and settings
Tool availability Restrict the capabilities presented to the agent Does not constrain unrelated host processes
Filesystem/network sandbox Constrain process access Must be configured for the actual execution path
Git branch/worktree Separate source changes for review Does not isolate credentials or the network
File checkpoint Restore supported local file changes Does not undo a remote deployment or database mutation
CI and release controls Enforce required checks and authorized publication Must evaluate the actual final revision

Claude Code's permission rules are not an operating-system security boundary. For example, blocking a particular command spelling does not block every equivalent program or invocation. Review the documented syntax and test the intended policy. A broad allowance for a test runner also permits that runner to execute repository code.

Use the sandbox documentation to determine which processes and access paths are constrained. A model service may require network access even when executed tests should not have broad egress. Design those paths separately rather than saying “no network” while depending on a hosted model.

Respect standing authorization. The application should ask when a proposed action exceeds authorized scope or repository policy requires it; it should not repeatedly ask for an already-authorized routine step. For release automation, enforce the actual organization's rules rather than inventing a universal manual-approval requirement.

Parallel work and subagents

Subagents can isolate context and perform bounded tasks. They are useful when the work can be separated without conflicting edits—for example, one reviewer examines persistence and another examines timer calculations. There is no standard “50,000 lines of code” threshold.

  1. Assign a concrete responsibility and expected evidence.
  2. Decide whether each worker reads shared files or edits an isolated checkout.
  3. Avoid overlapping ownership of the same files and generated artifacts.
  4. Join results against the same source revision.
  5. Recheck the integrated change, not just each worker's individual output.

A separate context window does not automatically provide a separate filesystem. Current Claude Code supports configurable worktree isolation for subagents; inspect the starting revision and integration behavior. See subagents.

Parallelism can reduce elapsed time while adding model calls and merge work. Use it for independent work with a useful local task in parallel, not merely because the feature exists.

MCP integration without excessive access

MCP can expose documentation search, a development database schema, or issue metadata. Configure the intended server, identity, capabilities, and environment. Prefer a scoped schema-read tool to unrestricted production SQL access when the task only needs schema inspection.

Current project-scoped Claude Code MCP configuration uses .mcp.json at the project root. Other scopes have separate behavior; follow the MCP configuration guide. Pin trusted server dependencies for repeatable automation and keep secret values out of committed configuration.

CI cannot rely on an interactive server-approval prompt. Review project-supplied settings and MCP configuration before a privileged run, and explicitly control loaded integrations. Tool output remains untrusted data. Permission to read a ticket does not imply permission to post a reply or mark it complete.

CI: keep untrusted task text out of shell code

A coding job can receive an issue through a webhook, create an isolated checkout, run an agent, and propose a patch. Authenticate the trigger and use a stable task ID to deduplicate repeated events.

Do not splice issue text directly into a shell script with a workflow expression. Quotes, substitutions, and line breaks in that text can become executable syntax before the agent even starts. GitHub documents this class of script injection.

A safer integration shape is:

Verified event → trusted intake program → task-data.json
Trusted job configuration → isolated worker + scoped credentials
Worker reads task-data.json as data → bounded agent run
Agent patch → independent checks on final revision → review artifact
Repository release policy → merge/publish if authorized and checks pass

Use structured arguments or files to pass task data. Correct shell quoting prevents one class of injection; it does not prevent prompt injection inside the task. Limit the worker's actual capabilities and keep production deployment credentials out of a worker that only needs to propose a patch.

Record a non-success outcome when the agent exceeds a deadline, hits a budget, cannot reproduce the bug, or fails verification. Do not convert an attractive final paragraph into a successful build status.

Cost and effectiveness

Calculate cost from measured usage and the current pricing of the selected access path. A single “5,000 tokens per task” number is inadequate without input/output mix, repeated context, caching, and retries.

For illustration, assume 40,000 input tokens at $3 per million and 4,000 output tokens at $15 per million. The model component is 0.04 × $3 + 0.004 × $15 = $0.18 per run. At 100 identical runs it is $18, before compute, integrations, and review. These are hypothetical rates, not current Claude pricing.

Track cost per accepted change, escaped regressions, reviewer time, and time to completion. A cheap run that creates a misleading test or requires a full rewrite may be more expensive overall.

Interview questions and answer notes

  1. Is Claude Code a particular Claude model? No. It is an agent product/runtime that uses a selected supported model and tools.
  2. Do passing tests prove the fix is correct? They provide evidence within their coverage. Check that they exercise the requirement and that assertions were not weakened.
  3. Does the nearest CLAUDE.md override every ancestor? The documented loading model combines applicable files; keep instructions consistent and inspect what loads.
  4. Does a worktree isolate a malicious test's network access? No. Source isolation and process/network isolation solve different problems.
  5. Why is issue text inside a CI shell script dangerous? Workflow substitution can turn untrusted text into executable shell syntax. Pass it as data.
  6. Can a checkpoint undo a remote action? No. Use the remote system's confirmed state and an appropriate recovery operation.
  7. When are subagents useful? When responsibilities and evidence can be separated with manageable coordination and integration costs.
  8. What should a coding job report when verification is incomplete? The actual incomplete state, completed work, failed or unrun checks, and the reason.

Final notes

Remember inspect → reproduce → change → verify → review. Tool access makes the agent capable; scoped execution and observable acceptance criteria make its work assessable. Judge the final behavior and diff rather than the confidence of the generated explanation.

Next: Coding models, products, and evaluation.

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← Choosing an AI framework: requirements, evidence, and operating cost
NEXT LESSONCoding models and agents: choose by evidence →

Explore the diagram