Concept lesson · Foundations
Idempotency, retries, and timeouts
Start here
Definition
An operation is idempotent when repeating the same logical request has the same intended effect as performing it once. A retry is another attempt at that request; a timeout only says the caller stopped waiting and does not establish whether the effect happened.
Why it matters: Networks can lose the response after a server commits. A client needs a way to recover the original result without accidentally creating another order or charge.
The server binds the caller and idempotency key to the request fingerprint and committed result. A retry must not create a second business operation.
Read the diagram step by step
- An authenticated client U9 sends POST /orders with idempotency key buy-204. The server commits order O17 and the result for (U9,buy-204) together.
- The reply is lost. Retrying the same request and identity returns the stored O17 result.
- Reusing buy-204 with a changed payload is rejected. A distinct intended purchase uses a new key.
- Bound retries with backoff, jitter and an end-to-end deadline; retry safety is not overload control.
Worked example
U9 submits buy-204 and the server commits order O17, but the reply is lost. A retry with the same caller, key, and payload returns O17 instead of creating O18.
Key takeaways
- A timeout means unknown outcome, not proven failure.
- Commit the request identity, the local business change, and its saved result together so a crash cannot separate them.
- Bound retry attempts and elapsed time; external effects need their own recovery contract.
You will learn to
- Distinguish a failed operation from an unknown outcome.
- Design an atomic idempotency record with a request fingerprint.
- Budget retries and isolate overloaded dependencies.
Practice in this chapter
8 interview questions with model answers and follow-ups.
Go to interview practiceUseful foundations: HTTP APIs and request lifecycle · Transaction isolation
Workload and timing examples are interview assumptions.
Dotted concept links open the relevant explanation in a new tab.
01Idempotency, retry, timeout, and deadline: definitions
Idempotency means that repeating the same logical operation has the same intended business effect as performing it once. A retry is another attempt at that operation. A timeout is a maximum waiting duration: if no response arrives, the caller does not know whether the server received the request, committed it, or lost its reply. A deadline is an absolute point in time by which a call should finish. Propagating one deadline bounds the total waiting budget across a call chain; it does not prove that remote work stopped or undo committed effects.
Idempotency does not require every low-level network message to occur once. Without a stable operation identity, another attempt can accidentally create a second business intent. Our target is one order O17 for checkout attempt buy-204, even if three HTTP attempts arrive.
This distinction appears in payments, file uploads, job queues, webhooks, and agent tool calls. First identify the business operation and which transaction or external service durably records its result. Then decide how a retry finds that result.
02Lost-response retry: one order from two attempts
Bind an idempotency key, a stable identifier for one logical operation, to the authenticated caller and a normalized representation of the request. For example, POST /orders uses Idempotency-Key: buy-204, user U9, item B2, quantity 1, and quote Q8. The following trace isolates the failure between committing O17 and returning its response.
| Step | Server state | Caller-visible state |
|---|---|---|
| 1 | No record for (U9, buy-204) |
Request sent |
| 2 | Transaction creates O17 and saves the request result | Still waiting |
| 3 | Transaction commits | Still waiting |
| 4 | Response is lost | Timeout; outcome unknown |
| 5 | Same key arrives again | Same logical operation retried |
| 6 | Server returns saved O17 result | Purchase confirmed once |
The request identity must come from a stable retryable intent, not a fresh random key on every network attempt. A distinct second purchase should use a new key. Reusing a key with a different item should be rejected rather than silently returning a result for the wrong request.
- 1 → 2first attempt or retryU9: buy-204 → Order API
- 2 → 3claim caller/key atomicallyOrder API → Unique request-result record
- 3 → 4commit order and result togetherUnique request-result record → Order O17
- 4 → 5durable external attempt identityOrder O17 → Payment attempt pay-204
- 3 → 6retrieve committed outcomeUnique request-result record → Retry: return O17
03Idempotency key, payload fingerprint, and atomic result storage
Store RequestResult(callerId, key, payloadHash, state, resourceId, response) with a unique (callerId, key) constraint. A payload hash is a fingerprint computed from the fields that define the operation, such as item, quantity, and quote. Canonical means these fields are normalized consistently before hashing, so equivalent inputs produce the same representation. It detects reuse of the same key for a different intent; it is not authorization.
When work cannot finish in one short database transaction, persist its progress and give a worker temporary ownership, often through a lease. Expiry lets a replacement take over, but recovery must still account for requests the previous worker may already have sent. This is why a long operation needs more states than simply “key absent” or “completed.”
Specify the key namespace: the group within which an idempotency key must be unique, such as all requests by one caller. The example uses caller-wide keys, so the fingerprint includes the operation and target as well as item fields. A tenant or service that uses separate namespaces must include that scope in the unique identity. Replaying a saved response still requires current permission; an old idempotency key must not expose a resource after access is revoked.
| Stored state | Same identity and payload | Unsafe reaction |
|---|---|---|
| No record | Atomically create the effect and outcome, or durably claim a long operation | Check absence and create outside one protected boundary |
| In progress | Return status, wait within budget, or recover ownership | Launch another uncoordinated worker |
| Completed | Return the recorded effect identity and an authorized result | Repeat the business mutation |
| External outcome unknown | Reconcile the original external operation | Treat timeout as rejection and choose a fresh key |
| Same key, different fingerprint | Reject the conflict | Return an unrelated old result |
A lease lets a replacement worker take over after a deadline. The old worker may resume later, so the store must atomically check the current ownership version (epoch) and expected state before saving a result. That check cannot undo an external request already sent. The receiving service still needs duplicate protection, or a way to check and resolve the uncertain result.
04External effects and uncertain payment outcomes
Suppose checkout calls a payment provider after creating an order. The provider charges successfully, but its reply is lost before local state records success. Repeating a new provider request can double-charge even if the local order insert was idempotent.
Use one stable provider attempt key for the payment, record it durably before or as part of scheduling the attempt, and reconcile the provider's status after uncertainty. A webhook may report the result, but duplicate and reordered webhooks need their own identity/state checks. Only finalize the local purchase once the confirmed result satisfies the state machine.
If the provider has no safe retry or status lookup, an uncertain payment may need manual investigation. Explain that limit. To claim duplicate protection, identify the exact action protected, how long its request ID is remembered and which failures are covered. “Exactly once” alone explains none of those.
Read the provider's actual contract rather than copying a generic retry recipe. For example, Stripe documents replaying the first saved status and body for an idempotency key, including a saved 500. Reusing that key can therefore replay an error without proving that no effect occurred; using a fresh key simply to escape the saved error can duplicate work. Keep the operation pending and use the supported recovery path.
05End-to-end deadlines and retry amplification
A deadline is an absolute point in time by which a call should finish; a timeout is a maximum waiting duration, often for one step. If the user allows two seconds for checkout, giving three nested services independent two-second timeouts can exceed that budget. Propagate the deadline or its remaining time budget through the call chain and stop work that is no longer useful when safe to do so.
Every parent branches into three total attempts, including the original. Read from top to bottom.
Remember: Retry budgets multiply across layers.
Read the diagram
- Count the branching levels: 1, 3, 9, 27.
- Three caller attempts each permit three middle-layer attempts.
- Each of those nine can permit three storage attempts, producing 27 in the worst case.
Try from memoryIf only the outer layer permits three attempts, how many storage attempts can one request cause?
At most three in this simplified chain, assuming each inner layer makes one attempt per call.
Retry transient transport failures or documented retryable responses when the operation is safe and time remains. Do not repeatedly retry invalid input, denied permission, or a business condition that is no longer satisfied, such as an expired reservation. Respect server retry guidance.
Budget connection setup, queueing, processing, backoff and response transfer within the same end-to-end limit. If Retry-After asks for a wait beyond the remaining interactive budget, return a retryable/pending result instead of sleeping and then starting an already-expired attempt. HTTP and RPC clients may have their own automatic retries, so inventory them before multiplying attempts. gRPC clients also need an explicit realistic deadline; deadline propagation and cancellation handling vary by language and application code.
06Exponential backoff, jitter, circuit breakers, and bulkheads
Bounding the number of retries still leaves two problems: many clients may retry together, and slow calls may occupy every available resource. The controls below address different parts of that load: when to retry, whether to call a failing dependency, and which workloads share a resource pool.
Exponential backoff increases the waiting interval between retries. For a base interval of 100 ms, caps might be 100, 200, 400, and 800 ms. Jitter randomizes each wait, for example choosing a value between zero and the current cap. Ten thousand clients then avoid retrying at exactly the same instant.
A circuit breaker stops calls temporarily after sufficient failure evidence and later allows limited probes. It reduces repeated futile work; it does not repair the dependency or authorize dropping important writes. A bulkhead gives workloads separate concurrency/resource pools so a slow image export cannot consume every checkout connection.
| Control | What it bounds | Example |
|---|---|---|
| Deadline | Total useful elapsed time | Stop interactive checkout attempts after its budget |
| Retry budget | Additional attempts | At most one extra request at the API layer |
| Backoff + jitter | Timing of repeated work | Spread recovery attempts |
| Concurrency limit | Work in flight | Only 50 simultaneous provider calls |
| Circuit breaker | Calls into known failure | Limited recovery probes |
Queues also need limits. If work arrives faster than it can complete indefinitely, an ever-growing queue delays the failure while consuming memory or storage; it does not add processing capacity.
07Interview walkthrough: safe checkout retries
Interviewer: “The customer presses Buy twice because the first request timed out. How do you prevent two orders?”
Candidate: “Both attempts use buy-204 for the same user and request. I save the unique request-result record and order in one transaction. If the reply is lost after commit, the retry returns O17. For an external payment, I reuse the provider’s request key and check uncertain results. I stop retries at the user’s deadline and spread them with backoff and jitter so an outage does not trigger a flood.”
The answer is grounded because it names the durable state before and after the lost response, rather than assuming the network delivers exactly once.
Practise the interview questions
Say your answer aloud before opening the model answer. Then answer the follow-up and compare the reasoning.
What does idempotency mean, and why does a timeout make it useful?
Reveal a model answer
Idempotency means repeating the same logical operation has the same intended effect as doing it once. A timeout does not prove failure: O17 may have committed before its response was lost. U9 retries buy-204 with the same caller and payload so the service returns O17 instead of making another purchase.
Interviewer follow-up
Does idempotency require the response bytes and every network message to be identical?
Reveal the follow-up answer
No. It concerns the intended effect. Retries can produce additional network messages or different status details while still referring to the same one business operation.
What the answer must demonstrate: Distinguish one business effect from one transport attempt.
What exactly does the idempotency key identify?
Reveal a model answer
“One caller’s logical intent, such as U9’s purchase attempt buy-204. I scope it to the authenticated caller and compare a canonical payload fingerprint. A fresh network retry reuses it; a new intended purchase gets a different key.”
Interviewer follow-up
What if the payload changes under the same key?
Reveal the follow-up answer
Reject a changed operation, target or canonical payload under the same scoped key. The stored result still needs current authorization; knowledge of the key is not permission to inspect another resource.
What the answer must demonstrate: Separate caller, intent, and payload.
Two requests both see no saved result. How is one order guaranteed?
Reveal a model answer
“The claim and business effect must share an atomic boundary, such as a unique request record and order insert in one transaction. A separate check-then-insert can let both proceed. The losing concurrent attempt waits for or retrieves the winner’s outcome.”
Interviewer follow-up
What if a long-running task is in progress?
Reveal the follow-up answer
Return its status or wait up to a limit. If a new worker takes over, give it a new ownership version and atomically reject old versions when saving local results. For external calls already sent, use the receiving service’s duplicate protection or check their outcomes.
What the answer must demonstrate: Show the atomic boundary.
Can idempotency records expire after a minute?
Reveal a model answer
“Only if the contract prevents valid retries after that minute or another durable identity prevents repetition. Deleting the record can make a delayed duplicate look like a new operation. I align retention with the retry horizon, business identifiers, and downstream retention.”
Interviewer follow-up
What if the provider retains keys for less time than we do?
Reveal the follow-up answer
Our workflow must stop blind retries outside the provider guarantee and reconcile through a durable provider resource ID or another supported status path.
What the answer must demonstrate: Treat deduplication retention as part of correctness.
Why doesn’t a local transaction make the external charge exactly once?
Reveal a model answer
“The provider is outside that transaction. It can charge and lose its reply before we save the result. I use one durable provider attempt key, query or reconcile its outcome, and process duplicate notifications safely. The local order and remote charge have separate commit boundaries.”
Interviewer follow-up
What if the provider lacks those capabilities?
Reveal the follow-up answer
I cannot invent the guarantee. I would state the residual uncertainty and design reconciliation, compensation, or an operational resolution path.
What the answer must demonstrate: Avoid blanket exactly-once claims.
Three layers each make three attempts. What reaches the bottom?
Reveal a model answer
“In the worst simple nesting, up to 27 calls for one user operation. That extra work can keep a struggling dependency down. I choose one retry layer or a shared budget, cap attempts and total time, and stop retrying permanent failures.”
Interviewer follow-up
Why add jitter?
Reveal the follow-up answer
Without it, synchronized clients can retry at the same intervals. Jitter distributes those attempts in time, reducing repeated spikes.
What the answer must demonstrate: Show the multiplication and the bound.
How do timeouts relate to a two-second user budget?
Reveal a model answer
“The two-second budget becomes an absolute deadline two seconds after the request starts. Each downstream call gets at most the remaining time as its timeout, including planned retries. Independent two-second waits at every layer can greatly exceed it and keep doing work after the user has left.”
Interviewer follow-up
Does cancellation undo a completed action?
Reveal the follow-up answer
No. Cancellation can stop unnecessary pending work, but committed effects still need normal reconciliation or compensation.
What the answer must demonstrate: Distinguish stopping work from reversing it.
How do you stop a slow export dependency from taking down checkout?
Reveal a model answer
“I separate concurrency pools so exports cannot consume all checkout workers or connections. I bound queues and use deadlines, then reject or defer lower-priority work when capacity is exhausted. A breaker can limit calls to the unhealthy dependency while probing recovery.”
Interviewer follow-up
Would an infinite queue preserve availability?
Reveal the follow-up answer
It accepts work without a credible completion time and can exhaust resources. Availability must include a meaningful service contract, not merely enqueueing forever.
What the answer must demonstrate: Protect a finite resource and explain overload behavior.
Blank-page exercise · 20 minutes
Build the answer yourself
Draw a purchase that commits before its response is lost. Add a concurrent retry and an uncertain payment outcome.
- Name the request key, caller, and payload fingerprint.
- Show the transaction and external-effect boundaries.
- Specify duplicate handling and retention.
- Calculate retry amplification and set a bounded budget.
Check that each component and design decision follows from your requirements and workload.
Recall the key ideas
Answer from memory before opening each card. Explain why the choice works and what it costs. Revisit missed cards tomorrow.
Idempotency, retries, and timeoutsTimeoutRecall first, then reveal
The caller stopped waiting; the durable outcome may already exist.
Unknown is not failed.
Return to lessonIdempotency, retries, and timeoutsThe order was saved, but the reply was lost. What makes the retry safe?Recall first, then reveal
The same authenticated caller, request key and matching request recover the saved order and result. A changed request using that key must be rejected.
Same request → same saved result.
Return to lessonIdempotency, retries, and timeoutsSafe retriesRecall first, then reveal
Safe operation + bounded budget + backoff/jitter + reconciliation.
Retry with a reason and a limit.
Return to lessonFinal revision
Summary and interview notes
Retry the same logical operation only when its effect can be recovered safely and the remaining budget justifies another attempt. A durable operation identity prevents duplicate local mutations; also define how external results are checked, how obsolete workers are prevented from publishing, and how long saved results remain available for retries.
Remember these points
- Bind the key to the caller, operation, target and normalized request fields. Reuse it when retrying the same request.
- Save the request claim, business change and result in one atomic transaction.
- A timeout leaves the result unknown. After ownership changes, the store must reject results from the former worker.
- Deduplication retention and provider key lifetime bound safe retry; an expired record can make an old request look new.
- Three retrying layers with three total attempts each can create 27 downstream calls.
Interview tips
- Draw the crash after commit but before reply, then add two concurrent retries.
- Show how a duplicate returns an already committed result before re-running create-time validation.
- Count automatic SDK/proxy retries and include connection, queue and backoff time in the deadline.
Important qualifications
- Saved outcomes still require current resource authorization.
- Cancellation and circuit breakers reduce future work; they do not reverse an external effect already committed.
Technical references
- AWS: making retries safe with idempotent APIsOperation identity and safe retry design.
- Stripe idempotent requestsA concrete provider-specific idempotency contract; retention and error behavior are not universal.
- Google SRE cascading failuresOverload feedback, resource exhaustion, and bounded recovery behavior.
- gRPC: DeadlinesExplicit deadlines, propagation, clock-skew handling and application responsibility for stopped work.
Practice marks stay in this browser.