Learnastra AI SYSTEM DESIGNAnup Rai

Interview toolkit

Whiteboard Exercises for AI System Design

By Anup Rai34 min readReviewed September 2026

A system design diagram explains how a defined user outcome is produced under stated constraints. Use these nine authored exercises to practice the full interview: requirements → baseline → flaws → detailed design → cost/benefit → closing remarks.

All workload numbers, targets, staffing assumptions and financial worksheets below are hypothetical unless a source is explicitly cited. They are inputs to reason about, not benchmark results or reports of actual company deployments. Valued human time is an economic cost; releasing capacity does not automatically reduce payroll or produce cash savings. The proposed design is one defensible option; revise it when the interviewer changes a requirement.

How to run a mock

  1. Read the prompt and state the key assumptions before reading the worked sections.
  2. Write functional and nonfunctional requirements as separate numbered lists.
  3. Trace data preparation and one live request through a complete baseline.
  4. Find the dominant failure, then add the minimum justified repair.
  5. Work through scale, quality, recovery and the complete cost comparison.
  6. Close with the main compromise, release evidence and remaining uncertainty.

Use a 35–45 minute practice session if that matches the interview you are preparing for. The duration and this guide's self-checks are rehearsal aids, not employer scoring rules. Label arrows with the information they carry, separate synchronous/asynchronous work, and mark state ownership and trust boundaries.

Exercise Dominant interview challenge
1. Enterprise RAG Fresh, authorized evidence over a large changing corpus
2. Customer support Correct resolution and uncertain external actions
3. Code review Useful findings against an exact repository revision
4. Document processing Correct records, partial extraction and reviewer capacity
5. Moderation Tail latency, rare-event precision and appeal decisions
6. Multi-tenant platform Isolation, fair capacity and accurate metering
7. Semantic search Filtered relevance, freshness and p99 latency
8. Evaluation pipeline Reproducible release evidence within a realistic budget
9. Memory and state Correct recall, deletion and resumed execution

For concepts behind these designs, use RAG, access control, capability assessment, reliability and current cost accounting.

Exercise 1: Enterprise RAG System

Prompt: Design a knowledge assistant for 50,000 employees and 10 million continuously changing documents from SharePoint, Confluence, Google Drive, and internal wikis. Enforce document permissions at query time, support English, Spanish, and Mandarin, and target a complete response in under three seconds for 95% of queries.

Functional requirements

  1. Answer employee questions with evidence from the permitted document corpus.
  2. Ingest source creations, edits, permission changes and deletions.
  3. Provide passage-level citations and an explicit limitation when evidence is insufficient.
  4. Support English, Spanish and Mandarin, including evidence in a different language from the question.
  5. Keep the initial release read-only; do not change HR or business records.

Nonfunctional requirements

  1. Serve 50,000 employees and a corpus of 10 million documents; validate document-size and update distributions.
  2. Target p95 complete-answer latency below three seconds under the agreed peak and output-length limits.
  3. Enforce current source access before disclosure; an unavailable check cannot grant access.
  4. For this exercise, assume a 15-minute normal content-update target, with separate immediate-serving policy for observed revocations and deletions.
  5. Define supported-answer quality, unanswerable behavior and language slices before choosing models; report availability and stale-source responses separately.

Start with the baseline

Architecture / visual model
flowchart TD S[Sources and ACL changes] --> I[Parse and version chunks] I --> X[Keyword and vector indexes] U[Authenticated question] --> R[Authorized retrieval] X --> R R --> C[Evidence and token budget] C --> G[Generate and verify citations] G --> A[Answer or abstain]
Read diagram source
flowchart TD
    S[Sources and ACL changes] --> I[Parse and version chunks]
    I --> X[Keyword and vector indexes]
    U[Authenticated question] --> R[Authorized retrieval]
    X --> R
    R --> C[Evidence and token budget]
    C --> G[Generate and verify citations]
    G --> A[Answer or abstain]

Scale and initial state contracts

Assume 50,000 employees, 20% daily adoption, and five questions per active employee: 50,000 × 0.20 × 5 = 50,000 requests/day. Over eight hours, that is 50,000 / 28,800 = 1.74 requests/s; a fivefold peak is 8.68/s. At 6,000 input and 400 output tokens per request, the peak demand is about 52,083 input and 3,472 output tokens/s. This is a capacity input; benchmark the selected serving stack before converting it into GPUs.

An index record might be {doc_id: "leave", revision: 19, chunk: 4, acl_version: 8, deleted: false}. When revision 20 arrives, stage its chunks, validate completeness, and switch the active revision. At request time enforce current access, including cached answers. An old citation must not resurrect a deleted or forbidden source. Ten million documents at an assumed ten chunks each means 100 million chunks; storage and freshness are separate from request sizing.

Make the RAG latency and access checks concrete

For the three-second complete-response target, an illustrative planning allocation is 50 ms identity/permissions, 100 ms embedding, 100 ms retrieval, 150 ms reranking, 1,500 ms generation, and 100 ms network/serialization: 2,000 ms plus 1,000 ms of contingency. These are stage allocations, not measured stage p95 values whose sum proves the overall p95. Output length, queueing, and cache misses can consume the headroom. Measure the complete path under load; time to first token is a separate experience metric.

The retrieval predicate is same_tenant AND current_version AND not_deleted AND (tenant_public OR explicitly_granted_user OR permitted_group). The server obtains tenant, user, and group membership from trusted identity and policy state. Apply the predicate in both keyword and vector retrieval, then revalidate before returning evidence or a cached answer. “Public” here means public inside the authorized tenant unless a separate cross-tenant publication policy exists. A model-supplied tenant ID cannot broaden this scope. Test permission revocation as well as initial access.

Find flaws and compare repairs

Flaw Repair Benefit and cost
Latest source revision is missing Durable ingestion, revision catalog and replay Restores freshness; requires reconciliation and source ownership
Exact IDs or paraphrases retrieve poorly Compare lexical/dense/hybrid candidates on labeled queries Better evidence recall if measured; adds index and embedding work
Correct passage is poorly ordered Test reranking after confirming candidate recall Better packing; extra latency and model/service cost
An answer cache survives a revocation Track source dependencies and revalidate current scope before reuse Closes another disclosure path; reduces usable hits and adds checks

Detailed design and recovery

Architecture / visual model
flowchart TD S[Source APIs and change cursors] -->|edition and permission events| Q[Durable ingestion queue] Q -->|source revision| W[Parser and completeness checks] W -->|versioned chunks| X[Lexical and vector indexes] S -->|current edition and tombstones| C[Revision and access catalog] U[Employee] -->|authenticated question| G[Admission and deadline] G --> R[Authorized retrieval] C -->|current policy| R X -->|candidates| R R -->|permitted passages| P[Optional rerank and evidence packing] P --> M[Approved model deployment] M --> V[Evidence and citation checks] C -->|revalidate dependencies| V V --> O[Answer or explicit limitation] W --> T[Restricted traces and freshness metrics] V --> T
Read diagram source
flowchart TD
    S[Source APIs and change cursors] -->|edition and permission events| Q[Durable ingestion queue]
    Q -->|source revision| W[Parser and completeness checks]
    W -->|versioned chunks| X[Lexical and vector indexes]
    S -->|current edition and tombstones| C[Revision and access catalog]
    U[Employee] -->|authenticated question| G[Admission and deadline]
    G --> R[Authorized retrieval]
    C -->|current policy| R
    X -->|candidates| R
    R -->|permitted passages| P[Optional rerank and evidence packing]
    P --> M[Approved model deployment]
    M --> V[Evidence and citation checks]
    C -->|revalidate dependencies| V
    V --> O[Answer or explicit limitation]
    W --> T[Restricted traces and freshness metrics]
    V --> T

Use durable source cursors plus reconciliation appropriate to each source API; a notification alone need not contain the full change. A language preference is a ranking choice, not a reason to exclude the only authoritative English passage from a Spanish question. Validate multilingual retrieval and answer support.

If a source cannot provide the promised permission freshness, make that limitation explicit and change the access architecture or scope. A five-minute cached group list cannot support a zero-staleness access promise. Check access before evidence reaches the model and before a saved or cached answer is disclosed.

Full cost and benefit

At 50,000 questions per workday and 22 workdays, compare 1.1 million monthly questions. These are illustrative full incremental budgets, with a human escalation taking one minute at $60/hour ($1 each).

Monthly cost Existing search Candidate assistant
Common search/service operations $15,000 $15,000
Model calls at assumed $0.012/question $0 $13,200
Added indexing/evaluation $0 $2,000
Additional operations $0 $3,000
Implementation amortization $0 $2,000
Escalations 6% × 1.1M × $1 = $66,000 3% × 1.1M × $1 = $33,000
Total $81,000 $68,200

The $12,800/month projected saving depends on the escalation reduction and comparable accepted outcomes. The candidate's non-escalation cost is $35,200; its break-even escalation rate is about 4.16%. Validate source/index rebuild costs at this corpus size and count any different employee handling time before claiming savings.

Closing remarks

Start with one owned corpus and a read-only evidence path. The dominant risk is stale or unauthorized evidence, not choosing a fashionable database. Expand after quality, freshness, access, latency and escalation-cost measurements meet the agreed criteria.

Interviewer changes the requirement: a user's access is revoked while their previous answer remains cached. Trace how you prevent exposure from index, cache, memory, and saved artifacts. Name the owner and measurable revocation policy.

Continue the deep dive: Complete enterprise RAG interview.

Exercise 2: Customer Support Chatbot

Prompt: Design e-commerce support for 10,000 conversations/day, a one-million-product catalog, order history, and FAQs, with three languages and integration into the existing Zendesk workflow. The business proposes resolving 70% of tickets without handoff; define correct resolution and recontact before accepting that target. Support order lookup, product questions, and eligible returns/refunds, with approval for policy exceptions.

Functional requirements

  1. Answer product and policy questions across the three required languages.
  2. Look up orders through authenticated customer-scoped tools.
  3. Create eligible return/refund proposals and execute only permitted actions.
  4. Obtain approval for exact policy exceptions and preserve the decision record.
  5. Handoff to the existing Zendesk workflow with context, status and unresolved effects; support a customer's request for human help.

Nonfunctional requirements

  1. Support 10,000 conversations/day and a one-million-product catalog; size turns and tool calls separately.
  2. Treat 70% no-handoff as a proposed business target, conditional on correct resolution and measured later recontact.
  3. Prevent cross-customer access and refunds beyond the remaining refundable balance.
  4. For this exercise, target p95 read-only answers below four seconds; represent slow actions asynchronously with visible status.
  5. Preserve operation identity across retries and recover known or unknown action outcomes after worker failure.

Start with the baseline

Architecture / visual model
flowchart TD U[Identity and request] --> P[Policy and order lookup] P --> V[Validate exact proposed refund] V --> H{Approval needed} H -->|Yes| W[Durable approval wait] W --> R[Revalidate and execute with operation ID] H -->|No| R R --> C[Confirm or reconcile receipt]
Read diagram source
flowchart TD
    U[Identity and request] --> P[Policy and order lookup]
    P --> V[Validate exact proposed refund]
    V --> H{Approval needed}
    H -->|Yes| W[Durable approval wait]
    W --> R[Revalidate and execute with operation ID]
    H -->|No| R
    R --> C[Confirm or reconcile receipt]

Scale and initial state contracts

At 10,000 conversations/day, a 70% no-handoff target leaves 3,000 human-handled conversations if all remaining cases escalate. At six minutes each, this is 300 reviewer-hours/day. At six productive review-hours per person, plan 300 / 6 = 50 daily staffed equivalents before coverage, leave, and surge allowances. If a narrower 5% subset needs a specialist, size that queue separately: 500 × 6 / 60 = 50 specialist-hours/day. Deflection is not the same as saved labor if difficult cases become longer.

Persist {case_id, order_id, amount_cents, currency, policy_version, approval_digest, operation_id, status}. A timeout after submitting the refund moves SUBMITTED → UNKNOWN, not FAILED → NEW_REFUND. Query the receiver using the same operation ID; revalidate approval if amount or policy changes. Recall: propose, authorize, persist, execute, reconcile.

Validate a proposed refund before execution

A tool proposal can contain {order_id: "o17", amount_cents: 40000, currency: "USD"}. A strict schema requires an integer amount, known currency format, and no unexpected fields. The service then checks the authenticated principal's ownership, the order's currency, refundable balance, policy version, and required approval. Neither a valid schema nor a model-provided approved: true establishes authority. The server creates or retrieves the stable business operation ID and persists it before sending the payment request.

Find flaws and compare repairs

Flaw Repair Benefit and cost
Customer email is used as proof of ownership Server-derived identity and order authorization Prevents arbitrary lookup; adds identity/account integration
Model says a refund succeeded Receiver receipt determines completion Accurate state; requires durable reconciliation
Timeout triggers a new refund Stable scoped operation identity and receiver deduplication Avoids duplicate effects within the receiver contract; retention and conflict handling matter
Handoff target rewards premature ticket closure Measure correct resolution and recontact, preserve unknown state Better outcome evidence; more labeling and follow-up work

Detailed design and recovery

Architecture / visual model
flowchart TD U[Customer or ticket webhook] -->|identity and event ID| G[Authenticated intake and deduplication] G --> R[Intent and bounded workflow router] R -->|policy question| K[Authorized catalog and policy retrieval] K --> A[Evidence-grounded answer] R -->|order operation| O[Scoped order service] O --> P[Typed return or refund proposal] P --> V[Ownership, balance and policy checks] V --> H{Exact approval required?} H -->|yes| W[Durable approval wait] H -->|no| J[Persist business operation] W -->|revalidate approved digest| J J -->|same operation key| F[Refund receiver] F -->|receipt or uncertain result| C[Reconciler and status store] C -->|confirmed outcome| A R -->|request, exception or no progress| T[Human queue with full handoff] C -->|unresolved| T
Read diagram source
flowchart TD
    U[Customer or ticket webhook] -->|identity and event ID| G[Authenticated intake and deduplication]
    G --> R[Intent and bounded workflow router]
    R -->|policy question| K[Authorized catalog and policy retrieval]
    K --> A[Evidence-grounded answer]
    R -->|order operation| O[Scoped order service]
    O --> P[Typed return or refund proposal]
    P --> V[Ownership, balance and policy checks]
    V --> H{Exact approval required?}
    H -->|yes| W[Durable approval wait]
    H -->|no| J[Persist business operation]
    W -->|revalidate approved digest| J
    J -->|same operation key| F[Refund receiver]
    F -->|receipt or uncertain result| C[Reconciler and status store]
    C -->|confirmed outcome| A
    R -->|request, exception or no progress| T[Human queue with full handoff]
    C -->|unresolved| T

Bind an approval to customer, order, currency, amount, policy/version and expiry. Recheck the order and refundable balance immediately before execution. Concurrent refunds require an atomic reservation or equivalent receiver-enforced balance rule; checking a balance in two independent requests is insufficient.

A ticket webhook may be delivered more than once. Deduplicate its event and the business action independently. Do not close the case while a consequential operation remains unknown. Sentiment or model self-confidence can inform routing, but neither alone proves that an action is allowed or a case is resolved.

Full cost and benefit

The earlier capacity calculation assumes six minutes per escalated case. Stress the economics with eight minutes for the harder remaining cases, at $60/hour.

Daily cost All-human baseline Candidate
Common ticket platform $1,000 $1,000
Human work 10,000 × 6 minutes = $60,000 3,000 × 8 minutes = $24,000
AI/tool usage, assumed $0.10/conversation $0 $1,000
Added operations $0 $600
Implementation amortization $0 $400
Total $61,000 $27,000

The projected saving is $34,000/day, subject to verified resolution, recontact and staffing outcomes. The candidate requires 400 productive human-hours/day, or about 67 staffed equivalents at six productive hours each, before coverage allowances. The specialist subset is part of this total unless explicitly additional. Deflection alone cannot substantiate the saving.

Closing remarks

Separate informational answers from state-changing actions. Launch a narrow authorized workflow, reconcile payment uncertainty, and expand only when correct resolution and full human workload support the business target.

Interviewer changes the requirement: the payment API times out after it may have committed. Distinguish unknown from failed. Show reconciliation rather than a fresh refund request. Measure correct resolution and recontact, not only chatbot deflection.

Continue the deep dive: Complete support-automation interview.

Exercise 3: Code Review Assistant

Prompt: Review 50,000 pull requests/day across GitHub/GitLab, respect repository conventions, and provide specific, actionable feedback and suggested fixes without disrupting developers.

Functional requirements

  1. Review pull-request diffs and relevant repository context after GitHub/GitLab events.
  2. Produce specific findings with file/line, consequence, severity and supporting evidence.
  3. Respect repository conventions without treating repository text as new authority.
  4. Suggest patches and report which checks actually ran.
  5. Supersede stale findings when the PR head changes; keep publishing, merging and deployment as separate permissions.

Nonfunctional requirements

  1. Handle 50,000 PRs/day, with measured size and burst distributions.
  2. For this exercise, target p95 completed review below two minutes for the supported PR-size class; show status for larger work.
  3. Run untrusted code/configuration in appropriately isolated workers without production credentials.
  4. Define finding precision, severe-defect recall and developer review-time objectives before rollout.
  5. Prevent duplicate or stale comments and preserve enough versioned evidence to reproduce a finding.

Start with the baseline

Architecture / visual model
flowchart TD E[PR event and commit SHA] --> Q[Deduplicated queue] Q --> S[Isolated checkout and static checks] S --> M[Diff and scoped context to model] M --> V[Verify findings and severity] V --> H[Human review of suggestions]
Read diagram source
flowchart TD
    E[PR event and commit SHA] --> Q[Deduplicated queue]
    Q --> S[Isolated checkout and static checks]
    S --> M[Diff and scoped context to model]
    M --> V[Verify findings and severity]
    V --> H[Human review of suggestions]

Scale and initial state contracts

At 50,000 PRs/day, assume 20% require model review and two model calls per reviewed PR: 20,000 calls/day. At 10,000 input plus 1,000 output tokens per call, that is 200 million input and 20 million output tokens/day. If each selected PR occupies a sandbox for 90 seconds, the mean concurrent sandboxes over 24 hours is 10,000 × 90 / 86,400 = 10.42; a sixfold burst requires about 63 before headroom. Peak webhook bursts and repository sizes need measurement.

Key work by (repository_id, head_sha, check_version). Before posting, verify the PR still has that head SHA. A finding includes file, lines, reason, evidence, severity, and test status. An unavailable test runner is UNAVAILABLE, never PASS. Running repository scripts requires isolation and a deliberate network/secret policy.

Find flaws and compare repairs

Flaw Repair Benefit and cost
Every repository script runs with the app's credentials Constrained disposable execution and scoped network policy Limits exposure; more runtime/maintenance work
Finding refers to an old head SHA Check exact commit before posting; rerun or supersede changed work Relevant feedback; cancelled/repeated work consumes capacity
Passing tests is claimed as complete compatibility Add contract/static checks appropriate to changed interfaces More evidence; suite maintenance and execution cost
Success is counted as comments produced Measure useful findings and investigation burden Less noise; requires labels and developer feedback

Detailed design and recovery

Architecture / visual model
flowchart TD E[PR webhook] -->|repo, head SHA, event ID| Q[Deduplicated job queue] Q --> S[Isolated checkout and bounded checks] S -->|diff, contracts and test status| C[Scoped context builder] C --> M[Model proposes findings] M --> V[Evidence, severity and duplicate checks] V --> H{PR head still matches?} H -->|yes| P[Scoped comment publisher] H -->|no| N[Supersede and schedule current revision] P --> U[Developer review] U -->|accepted, rejected or unclear| F[Adjudicated evaluation data] S --> T[Job trace and resource accounting] V --> T
Read diagram source
flowchart TD
    E[PR webhook] -->|repo, head SHA, event ID| Q[Deduplicated job queue]
    Q --> S[Isolated checkout and bounded checks]
    S -->|diff, contracts and test status| C[Scoped context builder]
    C --> M[Model proposes findings]
    M --> V[Evidence, severity and duplicate checks]
    V --> H{PR head still matches?}
    H -->|yes| P[Scoped comment publisher]
    H -->|no| N[Supersede and schedule current revision]
    P --> U[Developer review]
    U -->|accepted, rejected or unclear| F[Adjudicated evaluation data]
    S --> T[Job trace and resource accounting]
    V --> T

Store the logical job identity, attempt lease, head SHA, check/model/prompt versions and finding IDs. A stale worker cannot publish after a newer attempt has superseded it. API retries need idempotent comment publication or reconciliation of an uncertain post, not simply another comment.

A suggested patch is untrusted until checked. Label tests as passed, failed, unavailable or not run. Do not infer that a renamed public field is safe from a passing unit suite that never exercises consumers. Source comments can express coding conventions, but cannot grant access to secrets or authorize a merge.

Full cost and benefit

Compare all 50,000 daily PRs, with human time valued at $60/hour. The baseline averages five minutes of review. The candidate averages 4.5 minutes of ordinary review plus separate investigation of 2,000 AI comments at two minutes each; this extra time is not included in the 4.5-minute figure.

Daily cost Baseline Candidate
Common repository/CI costs $1,000 $1,000
Ordinary human review $250,000 $225,000
Extra AI-comment investigation $0 $4,000
20,000 model calls at assumed $0.02 $0 $400
10,000 sandbox runs at assumed $0.03 $0 $300
Added operations + implementation amortization $0 $300
Total $251,000 $231,000

The projected $20,000/day reduction requires a measured time saving with no unacceptable increase in missed defects. If ordinary review stays at five minutes, the candidate instead costs $256,000/day. A noisy reviewer can be a net burden despite inexpensive inference.

Closing remarks

Start with high-confidence, evidence-backed findings on a bounded PR class. Treat repository content and code as untrusted inputs, bind results to the current commit, and measure developer time and defect outcomes before expanding.

Interviewer changes the requirement: tests pass but a public API contract changed. Explain additional contract evidence and human review. Define whether the assistant can run builds, access secrets, or publish changes.

Continue the deep dive: Code-assistant design and autonomous-coding recovery.

Exercise 4: Document Processing Pipeline

Prompt: Process 100,000 financial-services documents/day: invoices, contracts, and forms, including PDFs, scans, and handwritten notes. Leadership asks for 99% extraction accuracy; define the field/document denominator and critical-field acceptance criteria. Provide traceable evidence and human review for uncertainty. Clarify the requested SOC 2 assurance scope and whether any health-data processing actually brings HIPAA obligations into scope.

Functional requirements

  1. Accept supported invoices, contracts and forms, including scans and handwriting within a declared quality range.
  2. Extract typed fields with source page/region evidence.
  3. Validate field relationships and business identifiers without guessing missing values.
  4. Route uncertain or inconsistent results to an authorized reviewer.
  5. Track processing, review and downstream posting separately; prevent duplicate business records.

Nonfunctional requirements

  1. Process 100,000 documents/day; assume ten pages/document for initial sizing and test the actual distribution.
  2. Define the proposed 99% accuracy target by field type and whole-document correctness; critical fields need their own acceptance criteria.
  3. For this exercise, target p95 automated draft completion within five minutes for documents of at most 20 pages; review has a separate service target.
  4. Restrict source and extracted data by tenant and role; apply the actual retention and assurance requirements.
  5. Bound queues and retries, preserve progress across crashes, and expose incomplete pages instead of marking the document complete.

Start with the baseline

Architecture / visual model
flowchart TD U[Validated upload] --> Q[Job queue and content hash] Q --> O[OCR and layout] O --> E[Typed extraction with page evidence] E --> V{Arithmetic and evidence checks} V -->|Pass| A[Accepted extraction] V -->|Uncertain| H[Reviewer queue] H --> A
Read diagram source
flowchart TD
    U[Validated upload] --> Q[Job queue and content hash]
    Q --> O[OCR and layout]
    O --> E[Typed extraction with page evidence]
    E --> V{Arithmetic and evidence checks}
    V -->|Pass| A[Accepted extraction]
    V -->|Uncertain| H[Reviewer queue]
    H --> A

Scale and initial state contracts

Assume 100,000 documents/day and ten pages/document: one million pages/day. At an illustrative two worker-seconds/page, work totals two million worker-seconds/day, or 23.15 continuously busy worker equivalents. Targeting 70% utilization gives 23.15 / 0.70 = 33.1, so at least 34 comparable workers for average load, then add measured peak/deadline capacity. If 2% of documents need four minutes of review, plan 2,000 × 4 / 60 = 133.3 reviewer-hours/day separately.

Represent a field as {amount_minor: 125000, currency: "USD", scale: 2, page: 2, bbox: [40, 90, 180, 120], source_hash, status: "needs_review"}. A document can be extracted but not approved. A repeated upload can reuse parsing only when its tenant, source bytes, parser version, and retention policy allow it; a content hash alone is not permission.

Validate relationships as well as individual fields

Suppose line items total 100,000 cents, the documented discount is 5,000 cents, and tax is 7,600 cents. Under an invoice rule with no other charges, the expected total is 100000 − 5000 + 7600 = 102600 cents. An extracted total of 120,600 cents fails this check even if every field has a high extraction confidence. Use decimal arithmetic or integer minor units with the currency's correct scale. Confirm discounts, shipping, tax basis, and rounding from the document before applying the equation; a missing charge should route to review rather than be “corrected” by guessing. Show the source regions for the conflicting values.

Find flaws and compare repairs

Flaw Repair Benefit and cost
“99% accurate” hides record errors Measure field and complete-document outcomes separately A meaningful release bar; more labeling
High confidence bypasses inconsistent arithmetic Validate source-backed relationships and route conflicts Detects some errors; may add review work
Duplicate upload creates duplicate posting Separate parse reuse from business idempotency Less repeated compute and duplicate effects; stronger state contracts
Tenfold volume overwhelms reviewers Capacity planning, priority queues and bounded intake Protects critical work; delays or rejects excess intake

Detailed design and recovery

Architecture / visual model
flowchart TD U[Authenticated upload] -->|validated bytes and tenant| O[Private immutable source object] O --> J[Durable document job and revision] J --> Q[Bounded page-work queue] Q --> P[Parser, OCR and layout workers] P -->|page results and failures| A[Complete-page aggregation] A --> E[Typed extraction with source regions] E --> V[Schema, arithmetic and business checks] V -->|eligible auto-accept| D[Accepted extraction version] V -->|uncertain or conflicting| H[Authorized review queue] H -->|reviewed exact version| D D --> B[Separately authorized business posting] B --> R[Receipt or reconciliation state] J --> T[Progress, queue age and audit metrics] H --> T
Read diagram source
flowchart TD
    U[Authenticated upload] -->|validated bytes and tenant| O[Private immutable source object]
    O --> J[Durable document job and revision]
    J --> Q[Bounded page-work queue]
    Q --> P[Parser, OCR and layout workers]
    P -->|page results and failures| A[Complete-page aggregation]
    A --> E[Typed extraction with source regions]
    E --> V[Schema, arithmetic and business checks]
    V -->|eligible auto-accept| D[Accepted extraction version]
    V -->|uncertain or conflicting| H[Authorized review queue]
    H -->|reviewed exact version| D
    D --> B[Separately authorized business posting]
    B --> R[Receipt or reconciliation state]
    J --> T[Progress, queue age and audit metrics]
    H --> T

Bind review to the source digest and extraction revision. A corrected extraction invalidates an earlier approval. Represent money using integer minor units or decimal values with the currency's scale; document bounding boxes must declare the coordinate system and page dimensions.

A parser may return partial pages. Mark missing/unreadable pages explicitly and avoid claiming a complete contract extraction. Parsing reuse is scoped to authorized source bytes and parser configuration; a matching content hash alone cannot authorize access or identify a duplicate invoice business transaction.

For illustration only, if 100 required fields each independently had 99% correctness, whole-document correctness would be 0.99^100 ≈ 36.6%. Real field errors are often correlated, so measure complete-document accuracy directly. Financial-services documents are not automatically subject to HIPAA, and SOC 2 assurance is not obtained simply by selecting encryption algorithms.

Full cost and benefit

At $60/hour, a four-minute review costs $4. Compare the same 100,000 daily documents with a current parser/review baseline and an additional extraction stage.

Daily cost Baseline Candidate
Common parsing/OCR, assumed $0.04/document $4,000 $4,000
Added extraction, assumed $0.05/document $0 $5,000
Human review 10% × 100,000 × $4 = $40,000 2% × 100,000 × $4 = $8,000
Operations $1,000 $2,000
Implementation amortization $0 $1,000
Total $45,000 $20,000

The projected $25,000/day saving depends on validated auto-accept quality and review time. At tenfold volume with unchanged page work, average-load capacity at 70% utilization rises to at least 331 comparable page workers; review work at 2% becomes about 1,333 hours/day, before coverage. Better model confidence alone does not eliminate that staffing requirement.

Closing remarks

Optimize for a correct, traceable business record, not merely parsed JSON. Keep extraction acceptance separate from posting, make partial failures visible, and size the reviewer queue as deliberately as the OCR workers.

Interviewer changes the requirement: volume rises tenfold and reviewers are saturated. Prioritize queues, protect high-impact fields, and define latency/quality tradeoffs. Do not simply lower review thresholds without measuring harm.

Continue the deep dive: Document-intelligence interview.

Exercise 5: Real-Time Content Moderation

Prompt: Moderate one million posts/day across text, images, and video in ten languages. Detect policy categories such as hate, violence, adult content, and spam; support appeals. Posts should become visible within 500 ms when permitted. Define the percentile, temporary treatment of uncertain posts, and whether every media analysis must finish before publication.

Functional requirements

  1. Evaluate text, image and video posts against versioned policy categories in ten languages.
  2. Record an allow, restrict, pending-review or other explicitly defined policy action.
  3. Provide slower analysis for cases outside the fast path without hiding their temporary state.
  4. Support appeals and authorized correction of earlier decisions.
  5. Preserve the content revision and evidence associated with enforcement.

Nonfunctional requirements

  1. Handle one million posts/day; size video duration and media bytes separately from post count.
  2. For this exercise, target p99 initial decision/state below 500 ms; content needing deep review remains explicitly pending under the agreed visibility policy.
  3. Define category/language precision and recall with false-positive and false-negative consequences.
  4. Report the pending fraction and final-decision delay so moving work off the fast path cannot disguise poor completion.
  5. Bound review queues, protect reviewer access/wellbeing and preserve appeal/reversal history.

Start with the baseline

Architecture / visual model
flowchart TD U[Content and policy version] --> F[Fast deterministic and classifier checks] F --> D{Confident decision} D -->|Yes| A[Durable allow or restrict action] D -->|No| Q[Deep model or reviewer queue] Q --> A A --> P[Appeal and outcome feedback]
Read diagram source
flowchart TD
    U[Content and policy version] --> F[Fast deterministic and classifier checks]
    F --> D{Confident decision}
    D -->|Yes| A[Durable allow or restrict action]
    D -->|No| Q[Deep model or reviewer queue]
    Q --> A
    A --> P[Appeal and outcome feedback]

Scale and initial state contracts

For one million items/day at 0.1% harmful prevalence, there are 1,000 harmful and 999,000 benign items. A detector with 90% recall and 1% false-positive rate produces 900 true positives and 9,990 false positives. Precision is 900 / 10,890 = 8.26%. This explains why impressive recall can still overwhelm review. Do not confuse false-positive rate with the fraction of flagged items that are wrong.

A decision record binds content_revision, policy_version, category, evidence, reviewer/model version, and action_id. If the content changes before enforcement, reconsider the decision. For the 500 ms publication budget, measure stage tails and reserve time for the response; a slow deep review may need a provisional restricted state rather than blocking the request indefinitely.

Find flaws and compare repairs

Flaw Repair Benefit and cost
Slow video/LLM analysis is placed in every synchronous path Fast eligibility decision plus bounded deep-review work Meets initial-state deadline if measured; some posts wait
Rare harm yields mostly false flags Calibrate by prevalence and policy slice Better review precision; may trade recall or require a new model
Content changes after classification Bind action to the evaluated content revision Avoids stale enforcement; reruns consume capacity
Every successful appeal becomes a training negative Adjudicate the reason and applicable policy before dataset inclusion Better labels; expert review and delayed updates

Detailed design and recovery

Architecture / visual model
flowchart TD U[Post and media upload] -->|content revision| S[Private staging and admission] S --> F[Fast rules and supported classifiers] F --> D{Policy action determined?} D -->|yes| A[Version-bound enforcement record] D -->|no| P[Explicit pending visibility state] P --> Q[Bounded deep-analysis queue] Q --> M[Media analysis and policy evidence] M --> H[Reviewer or validated decision rule] H --> A A --> V[Visible or restricted post] V -->|appeal| R[Independent reassessment] R -->|new versioned decision| A R --> L[Adjudicated labels, not automatic ground truth] F --> T[Latency, slices, prevalence and queue metrics] H --> T
Read diagram source
flowchart TD
    U[Post and media upload] -->|content revision| S[Private staging and admission]
    S --> F[Fast rules and supported classifiers]
    F --> D{Policy action determined?}
    D -->|yes| A[Version-bound enforcement record]
    D -->|no| P[Explicit pending visibility state]
    P --> Q[Bounded deep-analysis queue]
    Q --> M[Media analysis and policy evidence]
    M --> H[Reviewer or validated decision rule]
    H --> A
    A --> V[Visible or restricted post]
    V -->|appeal| R[Independent reassessment]
    R -->|new versioned decision| A
    R --> L[Adjudicated labels, not automatic ground truth]
    F --> T[Latency, slices, prevalence and queue metrics]
    H --> T

Thresholds must be measured for the policy and error costs; a score of 0.95 is not inherently a 95% probability of correctness. Fast-path allow/restrict and pending decisions need distinct metrics. A provider outage should follow the agreed temporary visibility policy, not silently allow everything.

The no-delay requirement may be incompatible with mandatory complete analysis of long video. State that conflict, bound accepted media or change the visibility/review contract. Use an atomic content-revision check when applying enforcement; a delayed result cannot automatically apply to edited content.

Full cost and benefit

Use the same hypothetical prevalence and recall from the sizing example. If a candidate reduces false-positive rate from 1% to 0.1% while retaining 90% recall, it yields 900 true positives and 999 false positives: 1,899 flags, with precision about 47.39%. This is a target scenario, not a promised threshold adjustment.

Assume every flag needs two minutes of review at $60/hour ($2), and compare the same million daily posts.

Daily cost Baseline detector Candidate
Detector/analysis compute $1,000 $1,800
Flag review 10,890 × $2 = $21,780 1,899 × $2 = $3,798
Common operations/appeal provision $500 $500
Added implementation amortization $0 $300
Total $23,280 $6,398

The potential $16,882/day saving comes mainly from reduced false flags. Review work falls from 363 to 63.3 hours/day. Both detectors still miss 100 harmful posts under these assumptions; decide whether that is acceptable independently of savings. Measure appeals and reviewer time rather than assuming they stay constant.

Closing remarks

Define policy and temporary visibility before selecting models. Measure error costs at the actual prevalence, bind enforcement to content versions, and treat appeals as a governed correction path.

Interviewer changes the requirement: harmful content becomes rarer and false positives become more costly. Reassess threshold economics and precision, even if recall stays constant. Keep slower deep analysis off a hard synchronous path when it cannot meet the deadline.

Continue the deep dive: Content-moderation interview.

Exercise 6: Multi-Tenant AI Platform

Prompt: Serve 500+ enterprise customers, each with its own documents and models, strict tenant data isolation, usage-based billing, and pricing tiers with different capabilities. Support interactive apps and background jobs, and produce evidence for the agreed SOC 2 control scope.

Functional requirements

  1. Configure each enterprise tenant's permitted models, regions, credentials and capabilities.
  2. Serve authenticated interactive requests and durable background jobs.
  3. Apply tenant/user authorization to data, tools, caches, jobs and artifacts.
  4. Meter usage with task/attempt attribution and produce reconciled billing records.
  5. Support tier changes, revocation and administrative audit without trusting client-supplied tier labels.

Nonfunctional requirements

  1. Support at least 500 enterprise customers with declared traffic and service classes.
  2. Prevent cross-tenant and within-tenant unauthorized disclosure, including asynchronous execution and observability.
  3. For this exercise, preserve a reserved interactive share of a shared quota while bounding background queues.
  4. Keep provider/model/region hard constraints during fallback and quota exhaustion.
  5. Track spend reservations and actual charges atomically; do not treat an unknown remote charge as zero.

Start with the baseline

Architecture / visual model
flowchart TD U[Authenticated tenant] --> A[Quota and token admission] A --> I[Interactive reserved pool] A --> B[Background bounded pool] I --> P[Provider or runtime] B --> P P --> M[Per-tenant usage and outcomes]
Read diagram source
flowchart TD
    U[Authenticated tenant] --> A[Quota and token admission]
    A --> I[Interactive reserved pool]
    A --> B[Background bounded pool]
    I --> P[Provider or runtime]
    B --> P
    P --> M[Per-tenant usage and outcomes]

Scale and initial state contracts

Assume a usable limit of 600,000 tokens/minute and 6,000 total tokens/request. The theoretical ceiling is 100 requests/minute before separate input/output or request limits. Reserve 60% for interactive traffic: 60 requests/minute at this assumed size. A background tenant asking for 10,000-token jobs cannot be treated as consuming the same capacity as a 500-token chat. Use token-aware admission and measure output uncertainty.

A scheduler item holds {tenant_id, class, estimated_tokens, deadline, retry_count}. Charge actual usage afterward; reject or defer work that cannot finish before its deadline. A failed provider must not cause each tenant to retry without a shared budget. Record retries against the originating tenant while enforcing an overall limit.

Find flaws and compare repairs

Flaw Repair Benefit and cost
One global FIFO queue Class/tenant-aware fair scheduling and bounded queues Better isolation; scheduling complexity and possible idle reserves
Tenant ID is only a log field Trusted scope propagated and enforced at each data/action boundary Actual isolation; more integration tests and policy state
Counters update only after the call Atomic admission reservations with reconciliation Limits concurrent overspend; unknown outcomes retain capacity
Each failed worker retries independently Shared deadlines, retry budgets and circuit state Less amplification; some work is rejected or deferred

Detailed design and recovery

Architecture / visual model
flowchart TD C[Admin control plane] -->|versioned policy and entitlements| P[Tenant configuration store] U[User or job submitter] --> G[Authenticated gateway] G --> A[Authorization and atomic quota reservation] P --> A A -->|interactive| I[Reserved interactive scheduler] A -->|background| B[Bounded fair job scheduler] I --> W[Scoped provider/runtime workers] B --> W P -->|recheck model, region and permission| W W --> D[Permitted data and tool services] W -->|attempt usage and status| L[Append-only usage ledger] L --> R[Provider reconciliation and bill adjustments] W --> O[Scoped results and artifacts] L --> M[Tenant metrics and operational alerts]
Read diagram source
flowchart TD
    C[Admin control plane] -->|versioned policy and entitlements| P[Tenant configuration store]
    U[User or job submitter] --> G[Authenticated gateway]
    G --> A[Authorization and atomic quota reservation]
    P --> A
    A -->|interactive| I[Reserved interactive scheduler]
    A -->|background| B[Bounded fair job scheduler]
    I --> W[Scoped provider/runtime workers]
    B --> W
    P -->|recheck model, region and permission| W
    W --> D[Permitted data and tool services]
    W -->|attempt usage and status| L[Append-only usage ledger]
    L --> R[Provider reconciliation and bill adjustments]
    W --> O[Scoped results and artifacts]
    L --> M[Tenant metrics and operational alerts]

The 600,000-token/minute example above assumes one usable combined quota for teaching. Real services may separately limit input tokens, output tokens, requests, concurrency, regions or model deployments. Enforce all applicable limits; do not divide a model's advertised quota into a fictional uniform capacity pool.

Bind usage to tenant, task, attempt, model/region/tier, rate version and disjoint billed categories. Deduplicate events by their provider/event identity and append corrections rather than silently rewriting settled bills. A queue item retains the originating tenant but must revalidate current capabilities when it executes. Model memory or prompt caches require their own documented isolation guarantees.

Full cost and benefit

For the same monthly useful workload, suppose unavoidable model/serving work costs $20,000. Compare existing retry waste with a controlled scheduler. Figures are illustrative and include incremental implementation and operations.

Monthly cost Baseline Candidate
Useful serving work $20,000 $20,000
Avoidable retry work $8,000 $2,000
Added scheduling/metering operations $0 $3,000
Implementation amortization $0 $1,000
Total $28,000 $26,000

The projected $2,000/month saving requires more than $4,000 of avoided waste to cover the new fixed costs, while maintaining the required completed outcomes. Reserved capacity can also increase idle cost. A necessary isolation control may still be justified when it does not save money; state the requirement rather than inventing savings.

Closing remarks

Make tenant scope and entitlements enforceable across synchronous and asynchronous work. Separate admission, scheduling, execution and billing state, and use observed workload and failure behavior to justify shared versus dedicated capacity.

Interviewer changes the requirement: one tenant floods the queue during a provider outage. Show how retry budgets, bulkheads, queue limits, and reserved capacity preserve other tenants' objectives.

Continue the deep dive: Multi-tenant application interview and gateway design.

Exercise 7: Semantic Search at Scale

Prompt: Search 50 million e-commerce products at 100 million queries/day with p99 latency below 100 ms. Support price, category, brand, and rating filters, personalization from permitted user history, and real-time inventory updates. Include exact SKU and paraphrase queries in the relevance tests.

Functional requirements

  1. Search 50 million products using exact identifiers and natural-language queries.
  2. Apply price, category, brand and rating filters with defined semantics.
  3. Rank eligible results using permitted personalization data.
  4. Propagate inventory, price and description updates with versioned state.
  5. Migrate embedding/index versions with evaluation, cutover and rollback.

Nonfunctional requirements

  1. Support 100 million queries/day; test a provisional fivefold average peak and mixed filter selectivity.
  2. Target end-to-end p99 search latency below 100 ms for the defined region/network boundary.
  3. For this exercise, target normal event-to-search metadata freshness within five seconds and revalidate stock/price at purchase.
  4. Preserve exact-SKU and other critical relevance slices during model changes.
  5. Keep index/query encoders compatible and avoid exposing personalized results through shared caches.

Start with the baseline

Architecture / visual model
flowchart TD D[Product source of truth] -->|versioned changes| I[Lexical index and searchable metadata] Q[Query and explicit filters] --> S[Identifier or keyword search] I --> S S --> V[Current eligibility checks] V --> O[Ranked search results]
Read diagram source
flowchart TD
    D[Product source of truth] -->|versioned changes| I[Lexical index and searchable metadata]
    Q[Query and explicit filters] --> S[Identifier or keyword search]
    I --> S
    S --> V[Current eligibility checks]
    V --> O[Ranked search results]

The lexical baseline supports exact identifiers and ordinary keyword queries. Add semantic retrieval only when measured paraphrase failures justify it; the following sizing estimates that candidate.

Scale and initial state contracts

With 50 million products and an assumed two vectors per product, plan for 100 million vectors, 768 dimensions, and float32 storage. Raw vectors require 100,000,000 × 768 × 4 = 307.2 GB in decimal units. Two copies require 614.4 GB before graph edges, IDs, metadata, allocator overhead, and backups. An 8-bit representation is 76.8 GB of raw codes, but quality, index overhead, and any retained full-precision vectors still matter. Measure recall against exact search on a representative subset while tuning search effort.

Never mix vector spaces during migration. Store {embedding_model, revision, dimensions, normalization, index_generation} with the index and query encoder. Shadow queries against B, compare exact-ID and language slices, then switch an alias atomically with a rollback. RRF combines ranks, not incompatible raw BM25 and cosine scores.

Separate query throughput, freshness, and the 100 ms deadline

100,000,000 / 86,400 = 1,157.4 queries/s on average. A hypothetical fivefold peak is about 5,787 queries/s, distinct from the number of stored vectors. Load-test filtered queries at that peak, including cold queries and inventory churn. One 100 ms planning budget is 5 ms admission/cache lookup, 15 ms query encoding, 30 ms parallel keyword/vector retrieval, 15 ms fusion/filter validation, 10 ms personalization, and 25 ms response/network. Concurrent retrieval paths contribute their critical-path duration, not their sum. This allocation leaves no spare time: measured tails may require caching, a faster encoder, less reranking, or a changed requirement. Verify the end-to-end p99 instead of adding component percentiles.

Inventory and price updates can update authoritative metadata without recomputing an unchanged description embedding. Description changes need new embeddings. Stream versioned updates, measure event-to-query freshness, and check authoritative stock before checkout. Personalization may reorder eligible results; it must not override price filters, current availability, or permission checks.

Find flaws and compare repairs

Flaw Repair Benefit and cost
Dense retrieval misses exact SKUs Lexical/identifier path plus evaluated rank fusion Protects exact matches; another retrieval path
Filtering leaves too few eligible results Filter-aware retrieval and evaluated overfetch/search effort Better recall under filters; more computation
Stale cache ignores inventory changes Version/expiry-aware invalidation and current metadata checks More accurate eligibility; fewer cache hits
A new encoder queries an old space Version-bound encoder/index pair and dual-index migration Controlled rollout; temporary double storage and indexing

Detailed design and recovery

Architecture / visual model
flowchart TD D[Product source of truth] -->|versioned changes| Q[Update stream and replay cursor] Q -->|price and inventory| C[Current searchable metadata] Q -->|description changed| E[Versioned embedding workers] E --> A[Index A and encoder contract] E --> B[Staged index B and encoder contract] U[Query and permitted user context] --> G[Admission and scope-aware cache] G --> R[Release router: matching encoder and index] R --> K[Parallel lexical and vector retrieval] A --> K B --> K C -->|filter and freshness checks| K K --> F[Rank fusion and bounded personalization] F --> V[Eligible result validation] V --> O[Search results] R --> T[Isolated shadow comparison and slice metrics]
Read diagram source
flowchart TD
    D[Product source of truth] -->|versioned changes| Q[Update stream and replay cursor]
    Q -->|price and inventory| C[Current searchable metadata]
    Q -->|description changed| E[Versioned embedding workers]
    E --> A[Index A and encoder contract]
    E --> B[Staged index B and encoder contract]
    U[Query and permitted user context] --> G[Admission and scope-aware cache]
    G --> R[Release router: matching encoder and index]
    R --> K[Parallel lexical and vector retrieval]
    A --> K
    B --> K
    C -->|filter and freshness checks| K
    K --> F[Rank fusion and bounded personalization]
    F --> V[Eligible result validation]
    V --> O[Search results]
    R --> T[Isolated shadow comparison and slice metrics]

The latency allocation in the sizing section spends the full 100 ms. A revised planning allocation might be 5 ms admission, 10 ms encoding, 25 ms parallel retrieval, 10 ms fusion/filtering, 10 ms personalization and 20 ms serialization/network: 80 ms, leaving 20 ms of planning headroom. This is still not measured p99 performance. Validate cold queries, selective filters, churn and regional network delay.

Partitioning only by category can create hot shards or miss cross-category intent. Choose partition/routing behavior using the distribution and query semantics. A description edit re-embeds affected products; it does not inherently require rebuilding every product. During a whole-model migration, replay updates to both versions through cutover and preserve a rollback-ready version.

Two full index versions each with two replicas require 1.2288 TB of raw float32 vectors under the example's dimensions, before auxiliary data. The earlier two-copy estimate is not a complete replicated migration footprint.

Full cost and benefit

A better search experience may cost more. Compare the same 100 million daily queries with a lexical baseline and a candidate hybrid design.

Daily cost Baseline Candidate
Base search infrastructure $1,800 $1,800
Additional vector serving $0 $1,200
Embedding/update work $0 $300
Operations $600 $800
Implementation/migration amortization $0 $100
Total $2,400 $4,200

The candidate costs $1,800/day more. If a verified incremental purchase contributes $6 after its non-search variable costs, the candidate needs 300 additional purchases/day to cover that increase. Measure causal purchase contribution in an appropriate experiment; clicks or a benchmark gain alone do not establish it. Include temporary duplicate-index capacity in the migration budget.

Closing remarks

Preserve exact-match behavior, current eligibility and the tail-latency contract. Use compatible index releases and measured relevance/business outcomes to justify the extra serving and migration cost.

Interviewer changes the requirement: an embedding upgrade improves a benchmark but harms exact SKU lookup. Compare slices and hybrid retrieval before replacing everything. Design a dual-index migration and rollback.

Continue the deep dive: Recommendation and retrieval design and production retrieval.

Exercise 8: Evaluation Pipeline for a Production LLM Product

Prompt: Support an assistant with 50,000 daily users and weekly model/prompt releases without unacceptable quality regressions. Design offline evaluation, CI gates, judge calibration, and production monitoring within a proposed evaluation budget of 2% of inference spend. State whether that budget includes human review and engineering.

Functional requirements

  1. Evaluate an exact candidate release against a versioned baseline and datasets.
  2. Run deterministic checks and appropriate semantic/human grading.
  3. Produce per-case outcomes, slice summaries and a reproducible gate decision.
  4. Calibrate judges and retain protected holdouts separate from development iteration.
  5. Connect offline evidence to isolated shadowing, canaries, production outcomes and rollback.

Nonfunctional requirements

  1. Support an assistant with 50,000 daily users and weekly release decisions.
  2. Keep candidate execution isolated from protected labels and production write capabilities.
  3. Preserve all scheduled outcomes, including execution failures and unknown grades; a missing result cannot silently pass.
  4. Apply hard requirements separately from average quality, and report uncertainty using the actual sampling/repetition structure.
  5. Treat 2% of inference spend as a proposed budget; define whether it covers human review, operations and engineering before claiming feasibility.

Start with the baseline

Architecture / visual model
flowchart TD C[Candidate release manifest] --> T[Deterministic contracts] T --> D[Versioned labeled cases] D --> J[Calibrated grading and paired comparison] J --> G{Hard gates and quality bounds} G -->|Pass| N[Canary with rollback] G -->|Fail| F[Failure analysis and new regression cases]
Read diagram source
flowchart TD
    C[Candidate release manifest] --> T[Deterministic contracts]
    T --> D[Versioned labeled cases]
    D --> J[Calibrated grading and paired comparison]
    J --> G{Hard gates and quality bounds}
    G -->|Pass| N[Canary with rollback]
    G -->|Fail| F[Failure analysis and new regression cases]

Scale and initial state contracts

Assume 2,000 cases, two configurations, and three repetitions: 12,000 generations. At a hypothetical $0.004/generation and one $0.001 judge call per generation, the run costs $48 + $12 = $60, before labeling and infrastructure. Twenty percent expert review at 30 seconds/result adds 20 human-hours, often the dominant expense. Repetitions of one case are correlated; do not treat 12,000 generations as 12,000 independent user scenarios.

Persist {case_id, dataset_hash, candidate_digest, output, rubric_version, grade, severity, grader_version}. Block on a reproduced unauthorized action even when the mean improves. Use paired cases for quality comparisons and keep a protected holdout; inspect failures on development data without repeatedly selecting against the final test set.

Find flaws and compare repairs

Flaw Repair Benefit and cost
A mean score conceals a critical regression Named hard gates and slice criteria Protects requirements; some releases remain blocked
Repetitions are counted as new independent tasks Case-level paired analysis or an appropriate hierarchical design Honest uncertainty; may require more unique cases
Judge errors are assumed negligible Human calibration, disagreement review and adversarial checks Better measurement; review cost
A fixed score delta is called a noise band Use prespecified criteria and uncertainty appropriate to the data More defensible decisions; statistical and data work

Detailed design and recovery

Architecture / visual model
flowchart TD C[Exact candidate and baseline manifests] --> O[Evaluation orchestrator] D[Versioned cases without protected labels] --> O O --> Q[Bounded isolated execution queue] Q --> W[Candidate and baseline workers] W --> R[Immutable outputs and execution statuses] L[Restricted references and rubrics] --> G[Deterministic or calibrated semantic grading] R --> G G --> H[Human disagreement and critical-case review] G --> A[Complete aggregation and paired analysis] H --> A A --> P{Evidence meets all release criteria?} P -->|yes| V[Approved limited canary] P -->|no or inconclusive| F[Failure analysis and new candidate] V --> M[Live outcomes, stop conditions and rollback]
Read diagram source
flowchart TD
    C[Exact candidate and baseline manifests] --> O[Evaluation orchestrator]
    D[Versioned cases without protected labels] --> O
    O --> Q[Bounded isolated execution queue]
    Q --> W[Candidate and baseline workers]
    W --> R[Immutable outputs and execution statuses]
    L[Restricted references and rubrics] --> G[Deterministic or calibrated semantic grading]
    R --> G
    G --> H[Human disagreement and critical-case review]
    G --> A[Complete aggregation and paired analysis]
    H --> A
    A --> P{Evidence meets all release criteria?}
    P -->|yes| V[Approved limited canary]
    P -->|no or inconclusive| F[Failure analysis and new candidate]
    V --> M[Live outcomes, stop conditions and rollback]

Bind the decision to code, model/endpoint, prompt, tools, retrieval, policy, dataset and grader versions. A passing report for a nearby commit is not approval for a changed release. Reconcile the expected case/attempt manifest with actual results and reject duplicates, foreign run IDs or missing evidence.

Binary criteria can be easier to operationalize for specific failures, but are not inherently unbiased or perfectly reproducible. Ordinal rubrics can also be useful when defined and validated. User thumbs, repeated queries and citation clicks are behavioral signals; they are not automatic correctness labels.

Keep development failures available for diagnosis while limiting selection against the protected final holdout. A canary or shadow that can issue real external writes needs explicit isolation or authorized effect controls; duplicating live requests is not automatically harmless.

Full cost and benefit

Complete the run-cost example above: 12,000 generations and judge calls cost $60, and 20% human review at 30 seconds/result requires 20 hours. At $60/hour, plus $15 orchestration/storage, that is $1,275/run.

Assume four release runs/month, $100,000/month inference spend, $500 evaluation maintenance and $500 production monitoring. These are illustrative budgets.

Monthly cost Broad 20% result review Candidate 5% stratified result review
Four runs: model, judge and infrastructure $300 $300
Human review $4,800 $1,200
Maintenance $500 $500
Production monitoring $500 $500
Total $6,100 $2,500
Share of stated inference spend 6.1% 2.5%

Neither plan meets the proposed $2,000/month ceiling. The lower-review design saves $3,600 but is acceptable only if its calibrated sampling still supplies the required evidence; critical cases may need additional review. Negotiate budget, scope or cadence. Do not silently remove required checks or omit labor to report compliance with 2%.

Closing remarks

Design the decision and evidence contract first. Version the complete experiment, preserve missing outcomes and critical failures, and show explicitly when the requested budget cannot support the needed confidence.

Interviewer changes the requirement: average quality rises while a rare high-severity failure worsens. Explain why the average cannot override the hard gate, who decides, and what evidence is needed next.

Continue the deep dive: Complete evaluation-gated CI/CD interview.

Exercise 9: Memory and State for a Long-Running Agent

Prompt: Support one million users whose assistant remembers selected facts and preferences over months, with sessions that may last hours. Design relevant recall, correction, forgetting, and per-user isolation, alongside execution state that can wait for approvals and resume across days.

Functional requirements

  1. Retain selected user facts/preferences and relevant episode summaries for later sessions.
  2. Retrieve useful memories under current user/tenant scope.
  3. Correct, supersede and delete memories and derived artifacts.
  4. Keep durable task progress, approval state and operation receipts separate from recalled facts.
  5. Resume long-running work across days with compatible state and current authorization.

Nonfunctional requirements

  1. Support one million users with declared active-user and session distributions.
  2. For this exercise, target p95 memory retrieval below 150 ms for the defined per-user record limit; model time is separate.
  3. Prevent memory poisoning from promoting untrusted content into authority.
  4. Before confirmed deletion completes, establish how current-serving state, derived records and in-flight work are invalidated; storage/backups follow an explicit retention/deletion policy.
  5. Bound per-call context, memory writes and background jobs; evaluate false, stale and missing recall separately.

Start with the baseline

Architecture / visual model
flowchart TD R[Recent raw turns] --> S[Episode summaries with provenance] S --> F[Approved durable facts] F --> C[Relevance and current-access checks] C --> P[Bounded prompt] W[Workflow checkpoints and operation receipts] --> P D[Deletion or permission event] --> F D --> S D --> R
Read diagram source
flowchart TD
    R[Recent raw turns] --> S[Episode summaries with provenance]
    S --> F[Approved durable facts]
    F --> C[Relevance and current-access checks]
    C --> P[Bounded prompt]
    W[Workflow checkpoints and operation receipts] --> P
    D[Deletion or permission event] --> F
    D --> S
    D --> R

Scale and initial state contracts

Assume a 16,000-token context budget: 2,000 instructions/tools, 3,000 recent conversation, 5,000 retrieved evidence, 1,000 memory, 3,000 reserved output, and 2,000 safety margin. These sum to 16,000. When evidence grows, select and compress with provenance instead of silently truncating the newest user constraint. Durable storage capacity is a different budget: one million users with 20 approved 500-byte facts each is 10 GB raw, before indexes, history, and backups.

A fact has {subject, value, source_event, valid_from, expires_at, consent_scope, tombstone_version}. A checkpoint says where execution resumes; a fact says what may be recalled. On resume, recheck deletion and permissions against current stores. A summary is a lossy view of evidence and must not override a later correction or resurrect a deleted preference.

Find flaws and compare repairs

Flaw Repair Benefit and cost
Summaries become unquestioned facts Preserve provenance and distinguish candidates from confirmed user information Better correction and trust; more metadata/review
“Newer and more confident” always overwrites Apply source authority, valid time and user correction rules Fewer false supersessions; unresolved conflicts may require clarification
Checkpoint restore resurrects deleted memory Current deletion epoch/tombstones and revalidation on resume Preserves current policy; invalidates some cached/paused work
Memory and execution receipts share one vague transcript Separate fact lifecycle from operation state Reliable recovery; additional schema and migration work

Detailed design and recovery

Architecture / visual model
flowchart TD U[User events and permitted observations] --> P[Provenance and purpose checks] P --> C[Candidate extraction and episode summary] C --> V[Conflict, authority and retention validation] V --> F[Versioned approved facts and summaries] Q[Authenticated current task] --> R[Scoped relevant recall] F --> R R --> B[Context budget and evidence packing] B --> M[Model invocation] D[Correction or deletion request] --> E[Authoritative epoch and tombstones] E --> F E --> R E --> I[Invalidate derived caches and affected in-flight work] W[Durable workflow, approvals and receipts] --> A[Resume with current-policy validation] E --> A A --> B M --> O[Validate result dependencies before disclosure or action]
Read diagram source
flowchart TD
    U[User events and permitted observations] --> P[Provenance and purpose checks]
    P --> C[Candidate extraction and episode summary]
    C --> V[Conflict, authority and retention validation]
    V --> F[Versioned approved facts and summaries]
    Q[Authenticated current task] --> R[Scoped relevant recall]
    F --> R
    R --> B[Context budget and evidence packing]
    B --> M[Model invocation]
    D[Correction or deletion request] --> E[Authoritative epoch and tombstones]
    E --> F
    E --> R
    E --> I[Invalidate derived caches and affected in-flight work]
    W[Durable workflow, approvals and receipts] --> A[Resume with current-policy validation]
    E --> A
    A --> B
    M --> O[Validate result dependencies before disclosure or action]

Working, episodic, semantic and procedural memory are useful conceptual categories, not a universal L1–L4 architecture. Semantic memory may contain beliefs or claims; storage does not make them true. A model's context can contain selected working state while the complete task record remains external.

Valid time records when a fact applies in the domain. Recorded/transaction time records when the system learned or stored it. valid_from and valid_to alone describe valid-time history, not a complete bitemporal model. Preserve both histories when the application needs that distinction. See long-term memory.

A deletion must invalidate derived summaries, embeddings and caches as required, not merely remove one visible fact row. A paused task rechecks current state. An in-flight provider request cannot be made to “forget” by deleting a local row: cancel where supported, prevent stale disclosure/actions, and apply the provider's actual retention contract. An external payment receipt remains business execution state rather than a removable preference.

Full cost and benefit

Assume 10% of the million users have one session/day: 100,000 sessions/day or 3 million per 30-day month. Compare a permitted full-history baseline with selective memory, assuming equal accepted task outcomes.

Monthly cost Full-history baseline Selective memory
Model context work 3M × $0.012 = $36,000 3M × $0.004 = $12,000
Extraction/summarization, assumed $0.003/session $0 $9,000
Storage and retrieval $2,000 $3,000
Operations $2,000 $4,000
Correction/review provision $3,000 $3,000
Implementation amortization $0 $1,000
Total $43,000 $32,000

The projected saving is $11,000/month, not the $24,000 reduction in context calls alone. Measure changed false recall, correction workload and missed information. The full-history baseline is eligible only when retaining and sending that history is permitted; an ineligible privacy design cannot win on price.

Closing remarks

Retain information for an explicit purpose, with provenance and correction/deletion semantics. Keep factual recall separate from execution authority, and validate current state on every resumed action and affected result.

Interviewer changes the requirement: a user deletes a remembered preference while a workflow is paused. Trace deletion and revalidation. Explain why restoring a checkpoint must not restore revoked authority or undo an external action.

Continue the deep dive: Memory architectures, state management and durable execution.

Add the leadership round

For any exercise, assume a four-engineer team as a rehearsal constraint, then explain:

  1. Which useful scope the team delivers first.
  2. Which uncertainty the first experiment resolves.
  3. Who owns source correctness, evaluation labels, interfaces and incidents.
  4. What is purchased, built, deferred or removed.
  5. What evidence permits expansion and what result would stop the project.

The staffing number is not a claim that every system needs four people. Tie the plan to actual skills, dependencies and operating work. For individual-contributor interviews, keep your design technically concrete without inventing management responsibility.

Final summary and notes

Recall card Check on the whiteboard
Requirements Numbered behavior, measurable constraints and explicit exclusions
Baseline Complete data and live request/action paths
State Identity, version, ownership, status and operation key
Failures Trigger, detection, recovery and remaining uncertainty
Scale Units, average versus peak, queues and human capacity
Quality Defined denominator, critical slices and missing outcomes
Economics Same workload, accepted outcomes and full incremental cost
Closing Choice, compromise, evidence and next decision

Change one constraint on your second attempt: latency, data freshness, permissions, budget, reviewer availability or workload mix. Explain which decisions change and which remain valid. Use answer frameworks for the conversation structure and the question bank for further practice.

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← Common Pitfalls in AI System Design Interviews
NEXT LESSONBehavioral Interviews for AI Engineers and Engineering Leaders →

Explore the diagram