Learnastra AI SYSTEM DESIGNAnup Rai

Concept · Understand the mechanism

Reliability patterns: bound work and recover without duplicate effects

By Anup Rai11 min readReviewed September 2026

Reliability is the ability of a system to perform its required function under stated conditions over time. For an AI application, a fast HTTP response is insufficient if the answer violates the task contract or a tool action runs twice.

Availability concerns whether the service is usable when needed. Resilience concerns handling and recovering from faults. Define the user-visible function and measure the complete path rather than assuming a reachable model endpoint is a reliable product.

Interview scope: interactive feedback and batch reports

A learning service provides interactive interview feedback and produces longer progress reports in the background. Both use a model provider and a shared database. The interactive request has an illustrative three-second useful deadline; batch jobs have a separate completion target.

Functional requirements

  1. Return feedback when dependencies complete within the task budget.
  2. Queue eligible batch jobs durably and expose their state.
  3. Classify failures and recover only when repetition is safe.
  4. Offer a permitted fallback, deferred result or clear failure.
  5. Let operators contain an outage and verify recovery.

Non-functional requirements

  1. Bound queues, concurrency, attempts and total elapsed work.
  2. Isolate interactive, batch and recovery capacity.
  3. Preserve identity, data location, source and tool constraints across fallback paths.
  4. Prevent retries from duplicating consequential effects.
  5. Measure successful outcomes, tail latency, rejection and remaining error budget.

Start with one provider, request deadlines, a bounded queue and a truthful degraded response. Multi-provider failover becomes useful when its tested benefit justifies integration, capacity and operating cost.

How a slowdown becomes overload

For a stable system, Little's law relates average in-flight work to arrival rate and time: L = λW. At 20 requests/second and one second per request, about 20 requests are in flight. At ten seconds, supporting that same stable throughput requires about 200.

If capacity cannot sustain arrivals, queues grow and the stable estimate no longer predicts a bounded backlog. More application replicas cannot create provider token quota or database capacity. An unbounded queue converts visible rejection into increasingly late failures.

Choose a pattern for the failure

Pattern Meaning Benefit Cost or mistake
Timeout Maximum wait for one operation Releases local waiting resources Remote work may continue
Deadline Last useful completion time for the whole task Coordinates all stages and retries Needs propagation and response reserve
Retry Another attempt at a recoverable operation Handles some transient failures Added load; unsafe writes can duplicate effects
Backoff with jitter Increasing randomized retry delays Reduces synchronized retry bursts Waiting consumes the task deadline
Circuit breaker Temporarily rejects calls to a dependency judged unhealthy Limits repeated failures while recovery proceeds Incorrect scope/threshold can block healthy work
Bulkhead Separate bounded resource pools Contains a workload's resource exhaustion Reserved capacity may sit unused
Admission/load shedding Limit accepted work or drop lower-priority work Protects useful completion under overload Some requests are rejected/deferred
Fallback An alternate permitted result or path Preserves some product value May change quality, cost and latency

Google SRE overload guidance explains why rejected excess work can preserve overall usefulness. Rejecting work is still part of the user experience and must remain visible in the relevant service metrics.

Retry policy: classify first, budget second

Outcome Next step
Brief connection failure before an effect Retry if safe and within deadline/load budgets
Provider 429 Respect quota scope and retry guidance; wait only if the request remains useful
Invalid credentials or malformed request Fix the underlying request/configuration; do not repeat unchanged
Context overflow Change context selection within the task contract; not a blind transport retry
Policy refusal or authorization denial Follow the product's policy; do not route around it
Timeout after a possible write Reconcile the stable operation ID or use a receiver's proven deduplication contract
Partially streamed answer Handle partial delivery explicitly; do not silently append a different model's answer

One application request with three SDK attempts inside three application attempts can make nine downstream attempts. Define one coordinated attempt policy, an end-to-end deadline and a service-wide retry budget. A breaker opening should stop attempts through that dependency, not trigger immediate retries against the same open circuit.

For full jitter, one common policy after attempt index k is a delay sampled uniformly from 0 to min(cap, base × 2^k). With base 100 ms and cap 1,000 ms, the ranges begin 0–100, 0–200, 0–400, 0–800 and 0–1,000 ms. Clip waits and calls to the remaining deadline; provider guidance may require a different minimum wait. Random delay does not make an unsafe operation safe to repeat.

Circuit-breaker recovery is a state machine

Architecture / visual model
stateDiagram-v2 [*] --> Closed Closed --> Open: failure policy triggers Open --> HalfOpen: cooldown expires HalfOpen --> Closed: bounded probes establish recovery HalfOpen --> Open: probe failure
Read diagram source
stateDiagram-v2
    [*] --> Closed
    Closed --> Open: failure policy triggers
    Open --> HalfOpen: cooldown expires
    HalfOpen --> Closed: bounded probes establish recovery
    HalfOpen --> Open: probe failure

In closed, eligible traffic flows. In open, calls fail quickly or use an allowed alternate path. In half-open, only a bounded set of probes tests recovery. Atomically reserve probe capacity: checking a counter without incrementing it can admit every concurrent request.

Scope health to the actual failure domain, such as endpoint/model/region. One tenant's bad key is not evidence that all tenants should lose the provider. Define which failures count, the observation window, minimum sample volume and whether you count individual attempts or final request failures. Circuit-breaker pattern.

Bulkheads and cancellation must release the right resources

Separate worker, connection or concurrency pools for workloads that need isolation. Bound the waiting queue and its maximum age as well as active work. Distributed instances may each enforce local limits; verify their total remains below a shared provider or database constraint. Bulkhead pattern.

Acquire and release every permit exactly once, including on cancellation, timeout and an exception inside the protected operation. A cancellation while waiting must not leak a queue slot; a timeout raised by work inside the context must not cause a queue permit to be released twice.

In Python, asyncio.wait_for() cancels its awaited task on timeout. Using it to wait for a hedge delay can accidentally cancel the primary attempt you intended to keep racing. asyncio.wait() returns completed/pending tasks without cancelling pending tasks on timeout. Whichever design you use, handle cancellation, collect task results and clean up explicitly. Python task semantics.

Hedging, failover and honest degradation

A hedge starts an additional suitable attempt while the original is still pending to reduce tail latency. A failover moves work to an alternate path after failure or unavailability. Return the first acceptable successful result, not merely the first task to finish; that task could have failed.

Hedging increases traffic and may worsen an overloaded dependency. Use a budget and spare capacity. Cancel losers where possible, but track work/billing that may continue. Do not hedge consequential writes without an explicit receiver-side deduplication contract and a reason to incur the complexity.

Failed capability Permitted degraded result Broken promise to avoid
Feedback generation Save the attempt and offer delayed feedback Present a canned score as evaluated feedback
Policy retrieval Show an applicable verified template or say evidence is unavailable Remove retrieval and invent current policy
Report provider Durable queued job with expected status updates Pretend an in-memory task is durably accepted
Payment/action result Pending/reconciliation state Say “failed” and issue a new independent action
Primary model Validated eligible alternate with available quota Ignore residency, tool or output-contract restrictions

Two providers can share cloud, identity, network or data dependencies. Failover needs capacity and periodic testing; adding a dormant secondary endpoint does not establish availability. See gateway routing.

Set objectives and close the design

A service-level indicator (SLI) measures a specified aspect of service. An SLO sets its target over a defined window. An error budget is the permitted amount of failure under that target. Track quality, availability and latency separately when one blended score hides a serious defect.

For 100,000 eligible requests and a 99.9% success target, the budget is 100 unsuccessful requests. Define eligibility and treatment of rejected, abandoned and partial work in advance. That budget is not permission for a severe data leak or unauthorized action; critical invariants have separate containment rules.

p99 is a high percentile of the measured distribution, not a per-request maximum. Adding component p99 values does not generally produce end-to-end p99. Measure the full path under realistic load, including retries and queueing.

Test quota exhaustion, slow/stalled streams, worker crashes, duplicate deliveries, partial writes and fallback incompatibility. Assign authority to shed load or disable a capability, preserve incident evidence, and verify recovery before reopening traffic. Prefer the smallest tested architecture that meets the stated outcome and recovery requirements.

Spend a deadline once, and release the capacity

For an illustrative three-second end-to-end deadline, reserve 200 ms for final response handling. At 0–100 ms, authenticate and admit; at 100–300 ms, retrieve; at 300–1,300 ms, attempt the model; at 1,300–1,400 ms, cancel/close that failed read attempt; at 1,400–2,400 ms, use an evaluated fallback; by 2,600 ms, validate and finish. There are 400 ms left. Giving the fallback a fresh three-second timeout would violate the original budget.

On timeout, release local semaphore permits in a finally block and cancel/close the transport where supported. If upstream work cannot be confirmed stopped, track it as outstanding; freeing a local permit does not prove the provider released its GPU. Do not automatically retry an ambiguous write as if it were a failed read.

A simple adaptive read timeout policy is min(remaining_deadline − response_reserve, clamp(rolling_p99 + network_margin, lower, upper)). For a measured 800 ms p99, 200 ms margin, and bounds 500–1,500 ms, the desired attempt timeout is 1,000 ms. If only 700 ms remains and response reserve is 200 ms, the maximum is 500 ms; skip the attempt if it cannot plausibly complete. Use minimum sample counts and a sane default after startup. Timeout-censored observations and outages can bias the estimator; cap it rather than letting degraded latency grow every deadline.

For bulkheads, suppose 40 local concurrency slots are allocated as 24 interactive, eight batch, and eight recovery/administrative. A batch flood cannot borrow interactive capacity without an explicit lending policy that preserves its guarantee. Separate queues with maximum age make overload visible. These numbers are assumptions to load-test, not universal allocations. Measure rejection, queue delay, and completion rate by class; an empty reserved pool may be an intentional price of isolation.

Interview questions with developed answers

Q1: How do you design for high availability in an LLM system?

Sample answer: I begin with the user outcome and permitted degraded behavior. I bound queue depth, concurrency, and request deadlines so a provider slowdown cannot consume all resources. I use classified, budgeted retries for safe transient failures, a circuit breaker for an unhealthy dependency, and separate capacity for critical and background work. A validated alternate provider may help, but I also plan for shared dependencies and unavailable fallback capacity. For a support system, a verified status response or human handoff can be the appropriate fallback. I measure task completion and quality alongside availability, because returning wrong answers quickly is not success.

Follow-up: Which dependencies could take both providers down together?

Q2: What is the difference between a circuit breaker and a retry?

Sample answer: A retry gives one operation another chance after a recoverable failure. A circuit breaker prevents many operations from repeatedly hitting a dependency that appears unhealthy. They can work together: a safe call gets a bounded retry policy while the breaker monitors service health. When open, the breaker fails calls quickly; later it admits only a small number of recovery probes. Neither pattern makes an unsafe payment retry safe. For that I still need a stable operation ID and a receiver contract or reconciliation path.

Follow-up: Where do you count failures? Define whether the breaker sees individual attempts or final request failures, then test that policy under load.

Q3: Why can autoscaling make an outage worse?

Sample answer: More application replicas increase the ability to send work, but they do not necessarily increase the capacity of the bottleneck. If every replica shares the same provider token quota, adding replicas can increase rate-limit errors and retry traffic. If the database is saturated, more workers can deepen its queue. I would identify the constrained resource, control admitted traffic, and protect essential operations. Scaling is useful when it adds capacity at that bottleneck; otherwise load shedding or queueing within deadlines may protect the system better.

Follow-up: Which metric distinguishes insufficient workers from exhausted external quota?

Q4: How would you choose timeouts and test them?

Sample answer: I start with the user's useful end-to-end deadline and allocate time across queueing, connections, retrieval, generation, and tools. Each stage receives the remaining budget, so retries cannot reset the clock. I measure real latency distributions and test slow responses, stalled streams, and cancellation. A timeout is evidence that we stopped waiting, not proof that remote execution stopped. For side effects, I preserve the operation identity and resolve the result separately. I also monitor abandoned work that continues consuming resources after the user is gone.

Follow-up: Why not simply add component p99s? Their tail events do not necessarily occur on the same requests; measure the complete path.

Q5: What should a manager require before automatic failover?

Sample answer: The alternate path must meet quality, permissions, residency, schema, and tool requirements and have enough tested capacity. I would ask for a realistic failure exercise, clear activation and recovery criteria, and an owner who can disable the switch. We should know the effect on cost and user experience. If those requirements cannot be met, a smaller honest fallback is preferable to an untested model taking consequential actions. I would periodically exercise failover because an unused secondary path can drift out of compatibility.

Follow-up: Does returning to the primary require another canary? It may, especially after a provider or configuration change.

60-second interview answer

I define reliability from the user's outcome: a fast HTTP response is not enough if the answer is wrong or an action is duplicated. I use deadlines, bounded retries, admission control, circuit breakers, and isolation to keep a failing dependency from exhausting the service. A retry is allowed only when its semantics are safe. Fallback models must pass the relevant quality and privacy gates. When safe completion is impossible, I return a clear partial result or handoff. I measure latency, task success, duplicate effects, and error-budget consumption together.

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← Ensembles: use multiple outputs only when they improve the decision
NEXT LESSONAI governance and compliance: turn obligations into operated controls →

Explore the diagram