System designby Learnastra

System-design interview · Extended interviews

Design a webhook delivery platform

By Anup Rai

Deliver signed events through bounded retries, explain a lost receiver acknowledgment, and isolate slow endpoints without claiming transport can guarantee arbitrary business effects.

You will learn to

  • Distinguish business event, logical delivery, transport attempt and receiver processing.
  • Build durable sender and receiver flows with explicit retry and transaction boundaries.
  • Defend destination security, fair scheduling, ordering and retention tradeoffs.

Practice in this chapter

8 interview questions with model answers and follow-ups.

Go to interview practice

Useful foundations: Message queues, event logs, delivery guarantees, and backpressure · Design a distributed message log · Databases, data models, and ACID transactions · Production readiness: SLI, SLO, observability, and recovery

Workload and timing examples are interview assumptions.

01Define acceptance before promising delivery

A webhook is an HTTP event sent to a customer-controlled endpoint. The sender records a business fact; the receiver decides what to do with it. In this design, a receiver’s successful response means it durably accepted the event, not that every downstream business action completed. An unreachable endpoint can exhaust a bounded retry contract.

Ask what the receiver’s success response must mean: durable acceptance or completed business work, and whether events need ordering. This design chooses durable acceptance and best-effort order, with bounded retries and inspectable failures.

Use shipment event E402: order O901 changed to shipped at object version 7. It matches endpoint EP 9 and creates logical delivery D22. The first HTTP attempt is A1. If the receiver commits E402 to storage but its 202 reply disappears, the sender does not know whether acceptance happened. Attempt A2 must preserve E402 and D22 while carrying a new attempt identity.

The smallest complete flow is shipment transaction plus pending event → delivery worker → signed HTTP request → receiver durable receipt → acknowledgment → later business processing. A pending sender record is an outbox. A receiver’s record of accepted events is an inbox. These solve opposite sides of the handoff, and no local transaction spans both companies’ databases.

02Functional requirements

Agree on what the service must do before choosing its components.

  1. Manage destinations. Support endpoint registration, subscriptions, configuration changes and signing-key rotation. Pin queued work to an endpoint/configuration version so a URL edit does not silently rewrite past attempts.
  2. Deliver business events. Send the recorded event to each matching customer endpoint and record durable receiver acceptance separately from completion of its downstream work.
  3. Inspect and control delivery. Expose status, bounded attempt history, next retry and terminal reason; support authorized pause and cancellation.
  4. Redrive retained events. Allow an authorized request to send a retained historical event again. Preserve its event identity and record the reason; today’s business state cannot replace expired historical bytes.

03Non-functional requirements

Use these hypothetical requirements for the worked interview. Confirm the assumptions with the interviewer; the numerical targets require testing and are not measured results.

  1. Workload and initial delay. Assume ten million events/day with three matching endpoints/event and a tenfold peak. Start healthy first attempts within five seconds for at least 95% of deliveries.
  2. Bounded retry and retention. Retain shared event payloads for seven days. For this exercise, allow at most ten automatic attempts in the initial run. An authorized redrive starts a separately audited run with the same ten-attempt ceiling; it preserves event/delivery identity and cannot extend the original seven-day payload lifetime. Cap total HTTP attempt duration at five seconds and capture at most 16 KiB of response data. Retain sanitized attempt/audit metadata for 30 days. A permanently offline receiver ends in an inspectable exhausted state.
  3. Durability and effects. A committed business change must not lose its pending event. Retries preserve event/delivery identity; the receiver must durably accept before acknowledging. One local business effect depends on the receiver’s own transaction and retained identity, not a sender-only exactly-once promise.
  4. Ordering. Use best-effort order. Consumers rely on object versions or reconciliation. Optional per-object serial dispatch can block later events behind a failure and does not establish global order across customers.
  5. Isolation and fair capacity. A slow endpoint must not occupy every socket or worker. Bound endpoint concurrency, tenant shares and global transport work; receiver replay protection must cover the supported redrive horizon.
  6. Authenticity and safe delivery. Authenticate tenant operations, sign documented bytes, enforce timestamp replay policy and validate the actual outbound destination on every attempt. Signature validity is distinct from schema validation, authorization and business idempotency.

04Build one durable sender and one receiver inbox

One periodic worker scans due deliveries. In a short database transaction it claims D22, records attempt A1 and assigns a lease token. A lease gives temporary ownership of local processing; its token lets the database recognize which attempt may update current state. Commit and release locks before making the remote HTTP call.

The worker loads the pinned event/configuration, validates the destination and signs a fresh request. On 202 it records accepted only if its token and expected state are still current. On a timeout it records an uncertain outcome and schedules a bounded retry. It cannot roll back the shipment or infer that the event never arrived.

The receiver verifies signature and schema, then atomically inserts a uniquely scoped inbox E402 and durable processing work. It returns 202 only after this commits. Its worker later applies the local business mutation and marks the inbox done in one local transaction. A process crash after acceptance is then recoverable internally.

This baseline supports a small service. Its one blocking worker at 0.5 seconds/request handles only two attempts/s, so throughput is the first scaling limit. Its lost-response behavior is already the core correctness problem; more workers will not remove that uncertainty.

Design diagramDurable event before outbound delivery

The shipment transaction ends before the customer HTTP call begins.

Durable event before outbound deliveryThe shipment transaction ends before the customer HTTP call begins. business to db: Commit O901 and E402; db to worker: Claim D22 with lease; worker to receiver: Signed E402 attempt; receiver to inbox: Durable unique receipt; receiver to worker: 202 after receipt commitCommit O901 and E402Claim D22 with leaseSigned E402 attemptDurable uniquereceipt202 after receipt commitSERVICEShipment applicationSTOREBusiness / outbox /delivery DBWORKERDue delivery workerEXTERNALCustomer endpointSTOREReceiver inbox andjobsyncasyncreturn
Read each connection in order
  1. syncCommit O901 and E402Shipment application → Business / outbox / delivery DB
  2. asyncClaim D22 with leaseBusiness / outbox / delivery DB → Due delivery worker
  3. syncSigned E402 attemptDue delivery worker → Customer endpoint
  4. syncDurable unique receiptCustomer endpoint → Receiver inbox and job
  5. return202 after receipt commitCustomer endpoint → Due delivery worker

05Size attempts and sockets rather than events alone

Ten million events with three destinations create 30 million logical deliveries/day, about 347/s average. At a tenfold peak that is roughly 3,470 first attempts/s. If each logical delivery averages 1.2 attempts, peak transport work becomes approximately 4,167/s. Retry rate is a capacity assumption to measure, not an unlimited allowance.

Quantity Calculation Implication
Normal in-flight first attempts 3,470/s × 0.5 s ≈ 1,735 Request duration sets socket demand
All-slow retry-adjusted peak 4,167/s × 5 s ≈ 20,835 Timeouts can dominate memory/connections
Shared payload retention 10M × 1 KB × 7 = 70 GB Store one body and reference it
Base delivery metadata 30M × 200 B × 7 = 42 GB Add attempts, indexes and copies
One-hour first-attempt backlog 3,470/s × 3,600 ≈ 12.5M Recovery needs staggered spare capacity

At an extra 1,000 attempts/s, that backlog takes about 3.5 hours to drain, if receivers permit the rate. More workers cannot responsibly exceed a customer’s accepted concurrency. Separate endpoint limits, tenant fairness and global socket budgets so a backlog cannot turn into a recovery stampede.

06Give every level a stable identity

The business transaction updates O901 and inserts immutable outbox event E402 together. The planner creates D22 with this unique delivery key:

(tenant, eventId, endpointId, configVersion)

Including endpoint ID matters: two different endpoints can both have configuration version 3 and both need the event.

Sender records

Record Fields and observations Purpose
Event Body, schema version and checksum Store immutable event bytes once.
Delivery Endpoint/configuration version, status, next-attempt time, retry count, current lease token and terminal reason Track one destination’s obligation and current scheduling state.
Attempt Start time, bounded response class and timeout observations Explain one transport attempt without replacing delivery truth.

A due-time index finds pending work without scanning every historical attempt.

Event envelope fields

Field Example or meaning
Event ID E402
Type Shipment event
Object ID/version Order O901, version 7
Payload The recorded shipment business fact.

Per-attempt headers

Header information Purpose
Delivery ID D22 remains stable across transport retries.
Attempt ID A1 and A2 identify different HTTP attempts.
Signing timestamp and key ID Identify the attempt’s signing context and replay-age boundary.
Signature Authenticate the documented raw bytes and fields.

Define exactly which raw bytes and fields are signed. Re-parsing and reserializing JSON can change bytes, so verification operates on the received representation before trusting its content.

Inspection returns tenant-scoped state and a bounded history cursor. Redrive requires authorization and retained original bytes. An expired payload cannot be reconstructed from today’s order state and called the same historical event. Store secret references rather than signing keys in ordinary delivery records or logs.

Record a scoped redrive request ID and its run ID atomically before making exhausted work eligible again. Repeating that request returns the same run rather than resetting its budget repeatedly. Track automatic-attempt count per run; do not reset the event identity or original payload-expiry time.

07Classify failures and schedule them fairly

Retry eligibility asks whether another attempt could help; timing asks when. Network errors, timeouts and many 5xx responses receive bounded retries. A 429 indicates receiver load control and may supply retry guidance that must be validated and capped. Document permanent endpoint/configuration errors rather than assuming every 4xx has identical meaning.

Exponential backoff increases delay after repeated failures. Jitter varies that delay across deliveries to avoid synchronized retries. Stop when the current run reaches ten automatic attempts or the original retained event reaches its seven-day expiry, whichever comes first. An authorized redrive is a new audited run, not an automatic counter reset. Expose next retry and terminal reason so “we retry” is an observable contract.

Add bounded parallel workers, with per-endpoint concurrency, per-tenant scheduling shares and a global socket ceiling. Limits cannot simply be independent per-worker counters, or adding workers multiplies the advertised endpoint cap. Expensive retries must not starve first attempts or other tenants.

When database polling becomes costly, use a ready queue containing delivery IDs and a durable due-time schedule for delayed work. The database remains authoritative. Workers recheck state after dequeue, and periodic due scans repair a lost enqueue. Queue duplication is expected and must not create a second logical delivery. This adds scheduling machinery while preserving the simple baseline’s stored obligations.

Design diagramFair dispatch separates durable obligations from HTTP attempts

The shipment transaction records the event. Due scans schedule delivery IDs; bounded workers claim current database state before sending. The receiver acknowledges its durable inbox, then applies its own local effect. Lost acknowledgments can repeat the same event.

Fair dispatch separates durable obligations from HTTP attemptsThe shipment transaction records the event. Due scans schedule delivery IDs; bounded workers claim current database state before sending. The receiver acknowledges its durable inbox, then applies its own local effect. Lost acknowledgments can repeat the same event. business to db: Commit shipment and event; db to due: Find due deliveries; due to queue: Enqueue delivery IDs; queue to workers: Dispatch fairly; workers to db: Claim and save outcome; workers to receiver: Signed HTTP attempt; receiver to inbox: Commit unique receipt; inbox to process: Pending local work; process to inbox: Commit effect and doneCommit shipment and eventFind due deliveriesEnqueue delivery IDsDispatch fairlyClaim and save outcomeSigned HTTP attemptCommit unique receiptPending local workCommit effect and doneSERVICEShipment applicationSTOREOutbox and deliveryDBWORKERDue-time schedulerQUEUEReady delivery IDsWORKERBounded senderworkersEXTERNALCustomer endpointSTOREReceiver inbox andresultsWORKERReceiver workersyncasync
Read each connection in order
  1. syncCommit shipment and eventShipment application → Outbox and delivery DB
  2. asyncFind due deliveriesOutbox and delivery DB → Due-time scheduler
  3. asyncEnqueue delivery IDsDue-time scheduler → Ready delivery IDs
  4. asyncDispatch fairlyReady delivery IDs → Bounded sender workers
  5. syncClaim and save outcomeBounded sender workers → Outbox and delivery DB
  6. syncSigned HTTP attemptBounded sender workers → Customer endpoint
  7. syncCommit unique receiptCustomer endpoint → Receiver inbox and results
  8. asyncPending local workReceiver inbox and results → Receiver worker
  9. syncCommit effect and doneReceiver worker → Receiver inbox and results

08Follow the lost acknowledgment through both databases

A1 sends E402/D22. NorthHarbor’s endpoint verifies the raw signed bytes and commits inbox E402 plus its processing job. Its 202 response is lost. The sender eventually records timeout and keeps the logical delivery eligible for retry; it has no observation distinguishing this history from a request that never arrived.

A2 receives a new attempt ID, timestamp, signature and lease token, while event/body identity remains stable. The receiver’s unique inbox lookup finds the already accepted fingerprint and returns 202 again. It does not create another processing job. D22 becomes accepted under A2’s valid token. The receiver’s separate processor may already have completed or may still be queued.

Now let A1’s old worker resume and report retrying after A2 succeeded. A guarded sender update requires its token to match the current delivery token and its expected state. The stale update fails instead of undoing accepted. Attempt history can still record a correctly associated late observation without replacing delivery truth.

The lease-token check stops A1 from overwriting newer sender state. The receiver’s saved event identity stops repeated arrivals from creating another local effect. Both are needed: rejecting A1’s late database write cannot undo an HTTP request already sent.

Request traceSame event, second transport attempt

Receiver acceptance and sender knowledge are independent facts.

Same event, second transport attemptReceiver acceptance and sender knowledge are independent facts. sender to receiver: A1: E402 / D22; receiver to db: Commit inbox + processing job; receiver to sender: 202 lost; sender to receiver: A2: same E402 / D22; receiver to db: Find matching saved receipt; receiver to sender: 202 acceptedPARTICIPANTSenderPARTICIPANTReceiverPARTICIPANTInbox database1. A1: E402 / D222. Commit inbox + processing job3. 202 lost4. A2: same E402 / D225. Find matching saved receipt6. 202 acceptedsyncblockedreturn
Read each connection in order
  1. syncA1: E402 / D22Sender → Receiver
  2. syncCommit inbox + processing jobReceiver → Inbox database
  3. blocked202 lostReceiver → Sender
  4. syncA2: same E402 / D22Sender → Receiver
  5. syncFind matching saved receiptReceiver → Inbox database
  6. return202 acceptedReceiver → Sender

09Prove the receiver’s effect and explain ordering

Use a trusted (sender, tenant, eventId) inbox key and immutable payload fingerprint. A receipt transaction inserts the inbox and job under that unique identity; a duplicate verifies the same fingerprint. Different bytes under the same identity conflict rather than silently replacing the event. Identity comes from the verified credential mapping, not an unsigned tenant header.

For a local database effect, use one processing transaction:

  1. Lock the inbox row. Concurrent processors serialize through that row.
  2. Check completion. Return without repeating the effect if the inbox is already done.
  3. Apply the effect. Make the local business change and mark the inbox done.
  4. Commit both changes. A crash before commit preserves neither; a retry after commit sees done.

This proves one local effect under the stated database boundary despite multiple HTTP attempts.

An external charge or email is outside that transaction. Persist a stable outgoing action identity and use the destination’s idempotency/status contract. Represent uncertainty and reconcile before inventing another action. An inbox alone cannot make an arbitrary remote API atomic.

For full object snapshots, atomically install only a newer object version. For changes such as “subtract one unit,” an older unapplied event still matters. Detect missing events and replay them or fetch complete state. Serial sending through each 202 orders receipt, while the receiver must separately preserve processing order if business actions require it.

10Treat outbound URLs as a trust boundary

A customer-controlled URL can trick the service into requesting internal or metadata endpoints: this is server-side request forgery. Restrict schemes, validate resolved IPv4/IPv6 destinations at connection time, and either disable redirects or validate every hop. Registration-time DNS checks cannot establish where a later connection will go.

The HTTP client must connect to the address that passed validation. Resolve, choose an allowed address and connect to that address while retaining the original hostname for TLS certificate verification. Do not validate one resolution and let the HTTP client perform a separate unchecked lookup. Add network egress restrictions as another boundary and repeat checks on every attempt.

Sign exact body bytes under a documented timestamp/key-ID scheme. The receiver verifies authenticity and rejects excessive timestamp age under a replay policy. Fresh retry signatures remain compatible with stable event identity. A valid signature does not remove schema checks or business authorization.

Rotate keys with a bounded overlap and explicit retirement rules. Limit body size, decompressed size, response bytes, connection time and total attempt duration. Sanitize inspection and logs because a receiver can return secrets in an error body. Pause/delete policies must be rechecked before sending even when a stale ready-queue entry exists.

11Operate a bounded promise through outages

Monitor first-attempt delay, eventual acceptance, oldest pending age, attempts per delivery, endpoint sockets, lease expiry, backlog age and payload-retention margin. Successful eventual delivery can hide hours of delay; separate it from the five-second healthy-first-attempt objective.

During an endpoint outage, back off that endpoint while other tenants continue. Reconnection does not authorize an unbounded drain. Keep failed deliveries inspectable and offer redrive only within the payload and destination-ownership contract. Receiver inbox retention must cover the allowed replay/redrive horizon, or the service must disclose a weaker duplicate-effect guarantee for older work.

Test sender death after POST, receiver death after inbox commit, local processor death around its transaction, stale lease updates, duplicate redrive, key rotation and DNS changes between attempts. Test one noisy tenant alongside healthy ones, not merely a global throughput benchmark.

Shared payloads reduce repeated storage, while attempt history can become a major cost under persistent failure. Bound retention and response capture. If the product requires proof of eventual business completion, add a separate authenticated status protocol; changing the meaning of HTTP accepted in marketing does not implement that requirement.

12Check the design against its requirements

Before closing, check the final design against the agreed requirements. FR means functional requirement and NFR means non-functional requirement; the numbers refer to the lists above. These are proposed validation checks, not test results.

Requirement Mechanism in the final design Validation and remaining limit
FR 1, 2; NFR 3 The sender outbox preserves the business event; pinned configuration and the receiver inbox preserve handoff identity. Crash after the shipment commit and after receiver acceptance. Require recoverable D22 and one receiver job after repeated E402, not proof of downstream completion.
FR 3, 4; NFR 2, 4 Due-time state, backoff, attempt limits and retained history make retry/redrive observable. Keep an endpoint offline through exhaustion, change its URL and request redrive after payload expiry. Do not promise eventual acceptance or global order.
NFR 1, 5 Bounded parallel workers and shared endpoint/tenant/global limits protect healthy traffic. Load-test first-attempt delay during one tenant’s timeouts and backlog drain; measure the five-second/95% goal without exceeding destination limits.
NFR 3, 6 Current lease-token updates protect sender state; signatures and scoped inbox transactions protect receiver processing. Resume A1 after A2 succeeds, replay a signed duplicate and conflict its fingerprint; reject stale updates and duplicate local effects. External effects need their own protocol.
NFR 6 Connection-bound address validation and network egress controls restrict customer URLs. Change DNS between retries, test redirects and rotate keys; verify that inspection never exposes credentials or unbounded response bodies.

13Rapid revision

Rehearse the numbered functional requirements and non-functional targets first. Use this table to recall the mechanisms, then close with the requirements check above.

Remember: Lost reply → same event, new attempt.

Prompt Recall
What survives a business crash? The business change and its outbox record commit together
What identifies a retry? Keep the event and delivery IDs; create a new transport attempt
What does 202 prove here? The receiver saved the event durably; its business work may still be pending
Why retry an uncertain timeout? The sender cannot tell whether the request or only the reply was lost
What stops an old sender overwriting newer state? Accept its update only if its lease token and expected state still match
What prevents a repeated local effect? Deduplicate the verified event in the inbox; commit the effect and DONE status together
What about external effects? Keep one action identity and recover its outcome under the external provider’s contract
What protects other customers? Limit work per endpoint, per tenant and across the service
What protects outbound access? On every attempt, validate the address the HTTP client actually connects to

Close with: “I preserve the shipment event before sending, schedule bounded fair attempts, and keep event identity stable through uncertain transport. The receiver commits an inbox before acknowledgment and protects its own business boundary. I can guarantee durable attempts and explain outcomes, but cannot force an unreachable customer to accept or make an arbitrary external action exactly once.”

Practise the interview questions

Say your answer aloud before opening the model answer. Then answer the follow-up and compare the reasoning.

Foundation · Question 1

What does the receiver’s 202 mean?

Reveal a model answer

Under this contract it means the event and processing work were durably accepted. The receiver can respond before completing its business workflow, which preserves low endpoint latency without losing work on process restart. A sender needing proof of completed processing requires a separate status or callback protocol.

What the answer must demonstrate: Under this contract it means the event and processing work were durably accepted.

Foundation · Question 2

Which fields change between A1 and A2?

Reveal a model answer

Event E402, delivery D22 and immutable body remain stable. Attempt identity, timestamp, signature and sender lease token change. This lets diagnostics distinguish network attempts while the receiver recognizes the same business event. A new event ID would undermine deduplication.

What the answer must demonstrate: Event E402, delivery D22 and immutable body remain stable.

Applied · Question 3

Can the sender avoid retries after a lost reply?

Reveal a model answer

It cannot infer acceptance from the timeout. Refusing every retry loses events in the history where the request never arrived; retrying can repeat arrival when receipt committed. Receiver inbox identity makes that repetition safe under the agreed contract, or a supported receipt lookup can resolve it.

What the answer must demonstrate: It cannot infer acceptance from the timeout.

Applied · Question 4

A1 reports timeout after A2 succeeds. What changes?

Reveal a model answer

The delivery row changes only if the result update carries the current lease token and expected in-flight state. A stale A1 cannot overwrite A2’s accepted state. Its observation can be retained separately in attempt history. This local fence does not prevent remote duplicate POSTs.

What the answer must demonstrate: The delivery row changes only if the result update carries the current lease token and expected in-flight state.

Applied · Question 5

Version 8 arrives before version 7. May version 7 be ignored?

Reveal a model answer

If events carry complete snapshots, a version guard can safely retain the newer state. If events are dependent deltas, dropping version 7 may lose a necessary effect; detect gaps and replay or fetch authoritative complete state. Waiting for each 202 orders acceptance, not necessarily receiver processing.

What the answer must demonstrate: If events carry complete snapshots, a version guard can safely retain the newer state.

Follow-up · Question 6

How does a five-second failing endpoint affect capacity?

Reveal a model answer

At the retry-adjusted peak, thousands of attempts per second can occupy more than 20,000 sockets. Per-endpoint concurrency and tenant scheduling shares prevent one receiver using the whole pool; global limits protect the fleet. Backoff retains obligations without repeatedly hammering the destination.

What the answer must demonstrate: At the retry-adjusted peak, thousands of attempts per second can occupy more than 20,000 sockets.

Applied · Question 7

Why validate DNS again on each attempt?

Reveal a model answer

A URL can resolve to a different destination after registration. Validate the actual chosen address and connect to it while checking TLS for the original hostname. Otherwise DNS rebinding or an unchecked second resolution can reach internal resources. Redirects require the same validation or must be disabled.

What the answer must demonstrate: A URL can resolve to a different destination after registration.

Follow-up · Question 8

Can an operator safely replay a three-month-old event?

Reveal a model answer

This design retains payloads for seven days, so it rejects unavailable history. A longer archive contract must preserve original bytes and destination ownership, and receiver deduplication or business reconciliation must cover that horizon. Reconstructing the current object is a new event, not faithful replay.

What the answer must demonstrate: This design retains payloads for seven days, so it rejects unavailable history.

Blank-page exercise · 45 minutes

Build the answer yourself

Deliver E402 to EP 9, lose the first 202, stall the endpoint, and recover without repeating the receiver’s local effect.

  • Agree the numbered functional requirements and non-functional targets, including receiver acceptance, ordering, bounded retry, fairness and security.
  • Estimate attempts, sockets and retention.
  • Trace sender outbox and receiver inbox.
  • Prove stale-worker and duplicate-effect handling.
  • Explain ordering, redrive and key retention.
  • Check delivery controls, replay, delay, retention, ordering and destination security against the numbered requirements; name external-effect and offline-receiver limits.

Check that each component and design decision follows from your requirements and workload.

Recall the key ideas

Answer from memory before opening each card. Explain why the choice works and what it costs. Revisit missed cards tomorrow.

Design a webhook delivery platformNorthHarbor saves E402, but its 202 reply disappears. What changes on retry?Recall first, then reveal

Keep E402 and delivery D22; create a new transport attempt. NorthHarbor finds its saved inbox record, so repeated arrival need not repeat the local business effect.

Lost reply → same event, new attempt.

Return to lesson
Design a webhook delivery platformHow does the receiver stop a repeated event from repeating a local update?Recall first, then reveal

In one transaction, change business state and mark the verified inbox event DONE. A retry then sees that completion.

Effect and DONE commit together.

Return to lesson
Design a webhook delivery platformWhy can the sender not promise that every event finishes its business work?Recall first, then reveal

The receiver may stay unreachable, retention may expire, and completion depends on the receiver and any external provider’s guarantees.

Bound attempts; disclose outcomes.

Return to lesson

Final revision

Summary and interview notes

The sender saves events and retries uncertain deliveries within limits. The receiver saves accepted events and prevents duplicate local business updates. A sender timeout alone cannot reveal whether the receiver accepted the event.

Remember these points

  • Agree the numbered functional requirements and non-functional targets before designing components; validate the final design against them.
  • Commit business state and outbox together.
  • Separate event, delivery and attempt identity.
  • Acknowledge only durable inbox acceptance.
  • Fence sender results and transact local receiver effects.
  • Bound fairness, retry lifetime and destination access.

Interview tips

  • Draw both timeout histories.
  • State receiver processing order separately from receipt order.

Important qualifications

  • External effects need their own recovery protocol.
  • A forever-offline receiver cannot be promised successful acceptance.

Continue after the core interview

Explore the advanced version

The advanced lesson keeps the full detailed design. Use these sections when you want to examine the stronger requirements and failure cases.

Technical references

  • Stripe webhook documentationProvider-specific examples of signature verification, duplicate deliveries, event ordering, and retry behavior.
  • Stripe idempotent requestsDistinguishes outbound API retry keys from an application’s webhook inbox and business deduplication policy.
  • OWASP SSRF prevention guidancePrimary security guidance on destination validation, unsafe address ranges and redirect handling for outbound requests to user-controlled URLs.

Practice marks stay in this browser.