System designby Learnastra

System-design interview · Extended interviews

Design a notification service

By Anup Rai

Separate notification intents, channel deliveries and provider attempts; design preferences, provider quotas, callbacks and uncertain-outcome recovery.

You will learn to

  • Separate a business notification from channel-specific delivery attempts.
  • Use durable scheduling without claiming exactly-once external delivery.
  • Explain priority, opt-out, quiet hours, and ambiguous provider outcomes.

Practice in this chapter

8 interview questions with model answers and follow-ups.

Go to interview practice

Useful foundations: Message queues, event logs, delivery guarantees, and backpressure · Databases, data models, and ACID transactions · Production readiness: SLI, SLO, observability, and recovery

Workload and timing examples are interview assumptions.

01Problem and scope

A notification service accepts business events and delivers approved messages through email, push, SMS or an in-app inbox. Distinguish the notification intent, one recipient/channel delivery, and each transport attempt. Give the intent, each recipient/channel delivery and each provider attempt separate IDs so retries can reuse the right operation and status can report which channels succeeded. Event ship-o81 is an example: one intent creates email and in-app deliveries while SMS is suppressed by preference. Provider acceptance, device delivery and user reading are different outcomes.

Interviewer: “Make notifications reliable.” Candidate: “Do we mean durably accepted by our service, accepted by a provider, delivered to a device, or read by a person?” Interviewer: “Track those separately. Transactional order updates are urgent; marketing may wait and must respect opt-outs.” This prevents us from reporting that a person received a message merely because a provider accepted it.

We build a shared notification service used by authenticated product services. They submit approved templates and business event IDs rather than arbitrary destinations and free-form scripts. We support one recipient per intent; a campaign service expands an audience into individual recipient intents while obeying the notification API's admission limits. We do not design an email server, mobile push network or marketing audience builder from scratch.

The delivery guarantee depends on which system performs the final action. Email and in-app differ here: an in-app item can be atomically inserted into our database, whereas an email provider performs a separate effect. A useful design promises one logical intent and recoverable delivery state, then states exactly where duplicate external delivery can remain possible.

02Functional requirements

  1. Accept an intent. Accept a notification intent for a tenant, event, recipient, category and versioned template. A repeated business event returns the original n44; changed parameters under that identity are a conflict or an explicitly new revision.
  2. Select eligible channels. Determine eligible channels and destinations from trusted user profiles and preferences. For ship-o81, create email d-email-44 and in-app d-app-44, and record why SMS is suppressed.
  3. Schedule delivery. Schedule immediately or at a specified instant/time-zone policy. Quiet hours may defer eligible work; urgent exceptions must be part of the recipient/category contract.
  4. Execute each channel. Deliver in-app items durably and invoke channel providers. Expose queued, dispatching, provider-accepted, known-delivered, failed, suppressed and unknown meanings per channel.
  5. Apply provider receipts. Accept authenticated provider receipts, update known outcomes, and let callers inspect partial success. Email success does not imply push success.
  6. Manage preferences and history. Change preferences, cancel unstarted work and inspect user-visible notification history. Cancellation after an external effect is not guaranteed recall.

Acceptance boundaries

Destinations are references to verified profile records, not untrusted phone numbers supplied by every caller. Templates have typed parameters and immutable versions so retries can reproduce the intended message. A rendering error is a terminal configuration problem to investigate, not a reason to retry a provider request indefinitely. The service should explain why a message did not send, not only whether a worker ran.

03Non-functional requirements

  1. Admission latency and availability. Assume p95 intent admission below 150 ms and 99.95% eligible intent-admission availability. Both exclude downstream delivery.
  2. Transactional delivery deadline. For this exercise, 99% of eligible, immediately due transactional channel deliveries receive confirmed provider acceptance within 30 seconds of durable intent acceptance. Scheduled marketing can wait.
  3. Durability. Acknowledged intents survive one database-zone failure under the configured replication protocol.
  4. Retention. Assume thirty days of detailed attempt metadata and ninety days of user-visible in-app history. Sensitive rendered bodies may have shorter retention or no persistence.
  5. Authorization and content safety. Enforce permitted channels/categories, versioned templates and typed parameters. Callers cannot inject arbitrary markup or change destinations without authorization.
  6. Consent ordering. An opt-out committed before the send-authorization transaction prevents that authorization. Recheck near dispatch; already authorized or externally accepted work may remain in flight and cannot always be recalled.

Measure the right delivery outcome

The 30-second denominator excludes policy-suppressed channels and separately scheduled work; provider errors and unknown outcomes count as misses at the deadline. In-app posting has a separate completion metric because it invokes no provider. Provider acceptance, device delivery and human reading are distinct facts. A bounced email and an SMS delivery receipt have different channel semantics; store them separately.

Category and time-zone policy

Support transactional updates, scheduled reminders and marketing over email, push, SMS and in-app channels. Bulk audience selection can remain a separate campaign system emitting bounded intents. Quiet hours use the recipient's time zone, not the worker's clock zone; define daylight-saving transitions and user time-zone changes. Only a specific user-approved urgent security exception can bypass the applicable preference. Putting marketing in a high-priority queue does not grant that exception.

Unknown external outcomes

An external timeout leaves two possibilities: the provider did nothing, or it acted but the reply was lost. Recovery depends on whether the provider can recognize a repeated operation or look up the original result.

Provider capability Required category policy
Idempotency or reliable lookup exists Recover under that documented contract
Neither capability exists Explicitly balance duplicates against missed delivery: a transactional update may tolerate a reviewed duplicate, while marketing may prefer no retry

This external-effect contract is separate from the uptime percentage.

04Capacity estimates

Treat attempts, channel deliveries and business intents as different units. The ratios below are workload assumptions.

Estimate Arithmetic Boundary
Provider attempts 100 million/day ÷ 86,400 ≈ 1,157/s Includes normal retries
Attempt burst Tenfold average ≈ 11,600/s Still bounded by each provider's contracted quota
Deliveries and intents At 1.1 attempts/external delivery and two eligible external-channel deliveries/intent: 90.9 million external deliveries and 45.5 million intents/day About 526 average intents/s and 5,260/s tenfold peak
Attempt metadata 100 million × 1 KB = 100 GB/day; 3 TB for 30 days Before indexes and replicas
Provider payload 11,600/s × 2 KB = 23.2 MB/s Before protocol overhead
In-app reads 20 million/day ≈ 231/s; tenfold peak ≈ 2,315/s Provision independently from sending
Failed-provider backlog 500 attempts/s × ten minutes = 300,000 attempts Retain the 300,000 pending attempts and pace provider calls within quota
Backlog drain 1,000/s capacity − 500/s new arrivals = 500/s net Another ten minutes to drain 300,000

Quotas and traffic mix

These averages count provider-backed channels only; in-app deliveries add local database work without a provider call. The ship-o81 example has one email and one in-app delivery, so its mix differs from this fleet-wide assumption. If each intent produces two external deliveries, its rate can be half the external-delivery rate; retry storms can reverse the relationship between intent and provider-call load. Measure failed addresses, suppressed channels and retries separately. Provider quota, reputation and cost usually constrain sending before raw bandwidth. More workers cannot exceed a downstream quota.

Retention and recovery budgets

Store sensitive rendered bodies only when required under appropriate retention. Queued work can reference immutable templates and safe parameter records instead of duplicating a full rendered body. Spread recovery retries instead of launching the whole backlog at once.

Separate transactional and marketing capacity, bound per-tenant queued volume, and reject impossible deadlines before acceptance. A campaign must not consume every send slot.

05APIs and contracts

The application submits a business intent; channel adapters later obtain provider results. These endpoints expose those different stages, so callers can distinguish acceptance of the notification request from progress toward delivery.

API Example Response
Create intent POST /notifications with {"eventId":"ship-o81","userId":"u7","template":"shipped-v3","params":{"orderId":"o81"}} n44, accepted
Inspect GET /notifications/n44 Per-channel queued, accepted, delivered, failed, or unknown
Preference PUT /users/u7/preferences Versioned channel/category policy
Provider callback Signed provider event with message ID Durable acknowledgment of callback ingestion

Create requests include a category, template version, optional schedule, and a stable event identity. The response is 202 with n44 and per-channel planning status; it is not a delivery receipt. Return 409 when the same identity contains different canonical parameters, 403 for an unauthorized sender/category and 429 when the caller exceeds its accepted-work budget. A permanently invalid template is rejected before committing work when detectable.

Preference updates carry an expected version. A successful response returns policy version 19, and a later dispatch authorization must use current state rather than trusting the planner’s old version 18. Cancellation addresses a logical intent or delivery and reports whether it prevented authorization, requested best-effort cancellation of in-flight work, or arrived after a known external effect.

History uses GET /users/u7/inbox?after=(createdAt,itemId)&limit=50, with verified user/tenant scope. Provider callbacks go to dedicated endpoints that verify the provider-specific signature and bind message references to the correct channel/account. They return success after durable inbox acceptance; asynchronous parsing/application can then retry without relying on the provider keeping the HTTP connection open.

Before the first send authorization, revalidate the destination and freeze the exact recipient, template version and rendered parameters for that delivery. Once any provider attempt may have happened, never change those bytes under the same provider key. An address change cannot turn recovery of the old delivery into a send to a new address. A deliberate replacement is a new logical delivery only after applying the category’s policy to the old delivery’s unresolved outcome.

Push registration and device fanout

A push provider routes messages using a registration token for an app installation. A user can have several installations, and their tokens can change. Registration operations maintain the account-to-installation binding before the planner chooses push destinations.

Device operation Contract
Register or refresh Authenticated PUT /users/me/devices/{installationId} submits the provider registration/token and platform; the server binds it to the current account and returns a destination version.
Sign out or unregister Disable that account/device binding. A device identifier or push token alone is not proof of account ownership.
Plan push Expand one recipient's push intent into deliveries for its eligible registered devices; each delivery names one immutable destination ID and version.
Resolve provider rejection A definitive invalid-registration response disables only the matching registration version; an old response cannot invalidate a newer refreshed token.

A user may have several devices, and reinstall or provider rotation may change a device's registration. Keep last-refreshed time and prune stale registrations under the chosen provider/product policy. Revalidate registration eligibility during send authorization, then freeze the exact registration for that delivery just like an email address. Never replace the token on an uncertain attempt under the old provider key. Expiration/TTL and collapse keys are channel policies: use them for replaceable updates when appropriate, not as proof that independent business alerts were delivered. Provider acceptance still does not mean the device displayed or the user read a push. FCM registration management.

06Data model and access patterns

These records preserve three kinds of information: what we planned to send, whether it was authorized, and what actually happened. An outbox stores pending dispatch events in the delivery transaction; an inbox durably accepts incoming provider events before they are processed. Both let a handoff resume after a process crash.

Record Key and important fields Query or invariant
Intent unique (tenant,eventId,userId,category); payloadHash, templateVersion Deduplicate ship-o81 and reproduce its meaning
Delivery (intentId,channel,destinationId,destinationVersion); state, dueAt, generation One logical destination across attempts
Attempt (deliveryId,attemptNumber); providerKey, authorizationVersion, outcome Explain calls and unknown responses
Push registration (tenant,userId,installationId); destinationId, version, provider registration reference, lastRefreshedAt, enabled Multiple devices per user; version-guarded refresh and invalidation
Preference (tenant,userId); categories/channels, version, timeZone Current permission at dispatch authorization
Template (tenant,name,version); typed schema, approved body Immutable rendering and safe reuse
Inbox/outbox unique event/dispatch identities and processing state Recoverable internal handoffs
In-app item unique deliveryId; recipient, contentRef, createdAt Retry-safe insertion and ordered user reads

Partition user-facing intent, preference, delivery and in-app state by tenant plus recipient bucket. That lets send authorization serialize against the relevant preference update within one owner. Provider quotas group work differently from recipient storage: queues can be grouped by provider/channel while each message carries the key needed to locate its delivery record.

Index pending deliveries by (state,nextAttemptAt,deliveryId), recipient history by (userId,createdAt,itemId), and provider references for callbacks. Queues contain IDs and minimal routing metadata, not the only copy of content or authorization. Sensitive destination and parameter records have restricted access and retention; metrics never use arbitrary email addresses as labels.

07Basic working design

Start with one authenticated API, one database, and a worker that polls an indexed pending-delivery table. The order service submits ship-o81 after its own order transaction through its outbox or equivalent reliable publication. Our API cannot atomically commit with the caller’s separate order database, so the caller must retry the same event identity until it knows acceptance.

The API transaction creates n44 and a planning task. The planner reads recipient U7’s preferences and template shipped-v3, creates two delivery rows and records SMS suppression. For in-app delivery, the worker locks the current preference and delivery rows, checks eligibility, and commits the unique d-app-44 item and completed delivery state in that same transaction. The email provider call occurs outside database locks after its separate send authorization. It saves the provider’s returned reference and reported acceptance. Recipient U7’s order remains accessible throughout; checkout does not wait for email.

This baseline is useful before a broker or many workers exist. Its pending rows survive process restart, duplicate intent submission returns n44, and the in-app channel has a clear commit boundary. It still needs a stated policy for uncertain email outcomes. A single worker is not a guarantee against duplicates: a crash after an external effect but before saving its reference already creates uncertainty.

Measure due-work scan cost, queue age and provider latency. The first benchmark should include slow responses and invalid destinations, because they consume worker slots differently from a stream of immediate successful calls.

architecture · baselineBaseline: durable intent then provider work

Recipient U7’s order does not wait for external email; the notification intent is durable before acceptance.

Baseline: durable intent then provider workRecipient U7’s order does not wait for external email; the notification intent is durable before acceptance. caller to api: Submit ship-o81; api to db: Commit n44 and planning task; db to worker: Find pending channel work; worker to db: Authorize send / commit in-app item; worker to provider: Send d-email-44 outside transaction; worker to db: Record provider fact or unknownSubmit ship-o81Commit n44 and planning taskFind pending channel workAuthorize send / commit in-appitemSend d-email-44 outsidetransactionRecord provider fact orunknownACTOROrder serviceSERVICENotification APISTOREIntent / delivery /preference DBWORKERPlanning and channelworkerEXTERNALEmail providersyncasync
Read each connection in order
  1. syncSubmit ship-o81Order service → Notification API
  2. syncCommit n44 and planning taskNotification API → Intent / delivery / preference DB
  3. asyncFind pending channel workIntent / delivery / preference DB → Planning and channel worker
  4. syncAuthorize send / commit in-app itemPlanning and channel worker → Intent / delivery / preference DB
  5. syncSend d-email-44 outside transactionPlanning and channel worker → Email provider
  6. syncRecord provider fact or unknownPlanning and channel worker → Intent / delivery / preference DB

08Find the baseline flaws

Assume a worker performs provider calls sequentially and each takes 200 ms. Its throughput is about five attempts/s. The assumed 11,600/s peak requires roughly 2,320 concurrent calls at that average latency if downstream quotas permit it. Merely adding a second worker cannot satisfy the target, and unrestricted concurrency may violate provider limits.

Polling all pending rows without a due-time index becomes expensive as the ten-minute outage adds 300,000 records. Repeatedly scanning that set can overwhelm the database even when no call is currently allowed. We need scheduled eligible work and provider-specific admission, not increasingly aggressive polling.

Now recipient U7 opts out of marketing at policy version 19 after the planner used version 18. If the worker sends using the old planned snapshot, the queued job has effectively granted permanent consent. Dispatch authorization must re-evaluate the current policy at the chosen owner and record the version used. The boundary must also admit that opt-out can race with an already authorized network call.

Finally, the provider accepts d-email-44 but the response disappears. A retry with a new attempt-based idempotency key can send twice. The correct external identity remains d-email-44; attempts are diagnostics, not new user intent. Database uniqueness solves the in-app case, but cannot force an arbitrary provider to remember that identity.

09Improve the design, step by step

First, separate durable planning from provider execution. Rising pending-table scans and several thousand concurrent calls trigger an outbox-to-queue relay and bounded provider workers. The planner commits each delivery and dispatch event together; the relay publishes its stable ID. More workers can now execute deliveries, while slow provider calls do not hold up new intent requests. The cost is broker storage, duplicate deliveries and another service to monitor and recover. A lost publish acknowledgment can cause repetition, so consumers recheck durable delivery state. Indexed database polling remains simpler when it meets throughput and scheduling needs; we add a broker when measurements show a need to scale workers separately.

Second, reserve urgency and tenant fairness. A large campaign can consume all provider tokens while ship-o81 misses its thirty-second objective. Divide work into transactional, scheduled and retry lanes, with weighted scheduling and reserved transactional capacity within each provider quota. At 1,000 allowed attempts/s, reserving an illustrative 600 for transactional work prevents a marketing burst from taking those slots; unused capacity can be borrowed under a bounded rule. Urgent messages get protected capacity; the provider’s total quota stays the same. Costs are scheduler complexity and potentially delayed marketing. Strict priority risks starving low-priority work; FIFO is simpler but fails the urgent target under campaigns. Select the policy from agreed deadlines and monitor age by lane.

Third, partition state by recipient ownership. Several thousand peak intents/s and ninety-day inbox history motivate tenant/recipient buckets. A directory routes recipient U7’s intent and preference operations to the same owner, keeping the permission check and dispatch claim in a local transaction. Queue assignment still follows provider rate domains. Each database holds a smaller active dataset and handles a different group of recipients; routing and migration become more complex. The new risk is authorizing against stale preference state during a move. Move preferences, deliveries and pending authorizations together, then transfer ownership using a new owner version. A larger single database is preferable when it still meets targets; splitting preference and delivery authorities prematurely would complicate the consent boundary.

Fourth, add durable receipts and recovery. Lost responses and reordered callbacks motivate a verified callback inbox plus an outcome reconciler. They save provider facts and follow each channel’s status rules; an old sent receipt cannot overwrite a delivered status. This improves visibility and recovers missed acknowledgments; it costs provider queries and additional retained evidence. A malformed or incorrectly matched receipt is the new risk, so match provider account, message reference and delivery identity. Polling every message forever is a rejected alternative: use provider capabilities, age-based recovery and terminal-state rules. Where neither query nor idempotency exists, preserve unknown status and the explicit duplicate-versus-miss policy.

Template caches improve rendering only after these boundaries are correct. An immutable version is safe to reuse; a stale preference cache is a different kind of data and cannot inherit that lifetime.

10Detailed architecture

The scaling changes produce two different groupings: recipient-owned records keep consent checks atomic, while provider-oriented queues enforce sending quotas. A queued delivery carries enough routing information to return to its recipient owner before it is authorized.

Recipient-owned state

The ingress API authenticates product services and authorizes template/category use. It routes tenant/recipient state through the owner directory. The owner database stores the intent, policy, deliveries, in-app items and outbox. Its synchronous replicas protect accepted state under the stated zone-failure model.

Planning and quota-aware dispatch

A planner expands approved intents into channel work. It uses immutable templates and a planning policy snapshot, but the dispatch claim rechecks current permission. The outbox relay publishes delivery IDs to a durable queue. A quota-aware dispatcher groups those IDs by provider/channel and urgency, then releases work only within per-provider and per-tenant budgets.

Channel execution and recovery

Channel adapters render the approved version, claim a send authorization at the recipient owner and invoke the external provider. The in-app adapter instead commits a unique inbox item locally. A callback receiver verifies and persists provider evidence before acknowledgment; a reconciler queries recoverable unknown outcomes. Both use the same delivery state transition rules.

Timing boundaries and implementation

The diagram separates the database that owns a recipient's preferences from the queue that schedules provider calls. Grouping work by provider helps enforce provider quotas, but send permission still comes from the recipient's current preference record. The API replies once the intent is durably saved. Planning, delivery, callback processing and recovery are asynchronous. Status reads expose these stages individually so “queued,” “accepted” and “delivered” never collapse into one convenient but misleading boolean.

One implementation starts with PostgreSQL transactions for recipient policy, deliveries and outbox, plus indexed due-work polling. Add SQS standard queues when independent worker scaling and buffering justify them, retaining the database as the authority. Standard queues can repeat deliveries; their visibility timeout is a worker-coordination aid, not an external-send guarantee. Provider adapters isolate channel-specific status, authentication and quota rules instead of forcing all providers into one success flag.

architecture · finalFinal: recipient ownership and provider quotas

The recipient database decides whether a send is permitted. Provider/channel queues control when permitted calls run within quota; queue priority cannot override an opt-out.

Final: recipient ownership and provider quotasThe recipient database decides whether a send is permitted. Provider/channel queues control when permitted calls run within quota; queue priority cannot override an opt-out. caller to api: 1. Intent, preference, status or inbox; api to directory: 2. Resolve tenant/recipient owner; api to db: 3. Commit intent / policy; read status; db to planner: 4. Recover pending planning; planner to templates: Load approved immutable template; planner to db: Commit deliveries + dispatch outbox; db to relay: 5. Read committed dispatch IDs; relay to queue: Publish stable delivery ID; queue to dispatch: 6. Select due category/provider work; dispatch to adapters: 7. Release within quota budget; adapters to db: 8. Atomic policy/send authorization; adapters to provider: 9. Send stable delivery key; adapters to db: Commit in-app item / outcome; provider to receipts: Signed channel outcome; receipts to db: 10. Persist receipt; apply state; reconcile to provider: Query original external reference; reconcile to db: Resolve unknown under state rules1. Intent, preference, status orinbox2. Resolve tenant/recipientowner3. Commit intent / policy; readstatus4. Recover pending planningLoad approved immutabletemplateCommit deliveries + dispatchoutbox5. Read committed dispatchIDsPublish stable delivery ID6. Select due category/providerwork7. Release within quota budget8. Atomic policy/sendauthorization9. Send stable delivery keyCommit in-app item / outcomeSigned channel outcome10. Persist receipt; apply stateQuery original externalreferenceResolve unknown under staterulesACTORProduct services /user clientsG1SERVICEAuthenticatednotification APIG1STORERecipient-ownerdirectoryG1STORERecipient DB +synchronous replicasG2WORKERPreference/templateplannerG2CACHEVersioned templatestore/cacheG2WORKEROutbox relayG2QUEUEDurable deliveryqueuesG3SERVICEQuota and fairnessdispatcherG3WORKERChannel adapters /in-app workerG3EXTERNALExternal channelprovidersG4SERVICEVerified receiptinbox/applierG3WORKERUnknown-outcomereconcilerG3syncasyncG1 Authenticated callers and routingG2 Recipient authority and planningG3 Quota-controlled recoverable executionG4 Independent provider effects
Read each connection in order
  1. sync1. Intent, preference, status or inboxProduct services / user clients → Authenticated notification API
  2. sync2. Resolve tenant/recipient ownerAuthenticated notification API → Recipient-owner directory
  3. sync3. Commit intent / policy; read statusAuthenticated notification API → Recipient DB + synchronous replicas
  4. async4. Recover pending planningRecipient DB + synchronous replicas → Preference/template planner
  5. syncLoad approved immutable templatePreference/template planner → Versioned template store/cache
  6. syncCommit deliveries + dispatch outboxPreference/template planner → Recipient DB + synchronous replicas
  7. async5. Read committed dispatch IDsRecipient DB + synchronous replicas → Outbox relay
  8. asyncPublish stable delivery IDOutbox relay → Durable delivery queues
  9. async6. Select due category/provider workDurable delivery queues → Quota and fairness dispatcher
  10. async7. Release within quota budgetQuota and fairness dispatcher → Channel adapters / in-app worker
  11. sync8. Atomic policy/send authorizationChannel adapters / in-app worker → Recipient DB + synchronous replicas
  12. sync9. Send stable delivery keyChannel adapters / in-app worker → External channel providers
  13. syncCommit in-app item / outcomeChannel adapters / in-app worker → Recipient DB + synchronous replicas
  14. asyncSigned channel outcomeExternal channel providers → Verified receipt inbox/applier
  15. sync10. Persist receipt; apply stateVerified receipt inbox/applier → Recipient DB + synchronous replicas
  16. syncQuery original external referenceUnknown-outcome reconciler → External channel providers
  17. syncResolve unknown under state rulesUnknown-outcome reconciler → Recipient DB + synchronous replicas

11Write path and acknowledgement

Durable intent acceptance precedes asynchronous planning and channel execution. Each external attempt retains its logical delivery identity across retries.

  1. The API authenticates the order service and inserts intent n44 for ship-o81. A retry with the same event identity returns n44; conflicting content is rejected or versioned explicitly.
  2. A planner evaluates recipient U7's preferences and template shipped-v3, producing email delivery d-email-44 and in-app delivery d-app-44. Their rows and dispatch events commit together. SMS is suppressed with an auditable reason.
  3. The in-app worker locks and checks current preferences, suppressing a disallowed delivery or atomically inserting the unique d-app-44 inbox item with its completed outcome. Retries return the existing outcome. The email worker claims due d-email-44, rechecks current consent/address validity, records attempt 1, and calls the provider using a stable idempotency key if that provider supports one.
  4. The provider returns message ID provider-902. The worker records accepted; a later authenticated callback changes the known delivery result. A queued callback arriving after a delivered callback must not blindly move the delivery backward.
  5. Recipient U7 can inspect the order immediately while channel work proceeds. The notification API reports each channel's actual state rather than turning partial success into one misleading boolean.

The planner’s delivery insertion and dispatch outbox share a transaction, so a crash before publication cannot strand an accepted plan. A repeated planner task observes the existing stable deliveries. For delayed work, a scheduler releases it when due rather than repeatedly calling a provider before quiet hours end.

The dispatch transaction records whether the message is authorized and the policy/destination versions used. It then releases database locks before the external call. Owner versions can reject stale workers inside our service. They cannot stop an email provider that does not check those versions. Therefore the stable delivery key and provider-supported status recovery remain necessary for the uncertain interval.

If the provider call succeeds but the database update fails, the delivery remains recoverable as dispatching/unknown. The worker does not mark it definitively failed merely to make a retry easier. The worker acknowledges the queue only after saving its progress. If that acknowledgment is lost, the next worker checks the saved delivery before deciding whether it can call the provider.

12Read and delivery path

Status reports known facts per channel, while inbox reads enforce recipient ownership and cursor ordering.

  1. The order service queries n44 using its tenant-scoped credentials, or recipient U7 queries the recipient-owned inbox with a user credential. Each path checks resource ownership rather than relying on an opaque ID.
  2. The router resolves the recipient bucket. An immediate status request can use the owner to avoid returning not-found for an intent already acknowledged; history may use a replica only under an explicit freshness policy.
  3. The API combines intent status with channel facts: in-app committed, email provider-accepted, SMS suppressed by policy. An unknown attempt is exposed separately from a definitive invalid-address failure.
  4. For the inbox, the store seeks (u7,createdAt,itemId) and returns the next fifty entries. The cursor carries the last tuple. New arrivals may appear on a refresh; a cursor alone is not a frozen snapshot.
  5. Opening an item retrieves permitted template/content references and applies current visibility rules. A deleted or tenant-restricted order should not leak through an old notification preview; retain only the necessary message content and links.

Read/open tracking, when supported and permitted, is another event with its own semantics. A mail tracking pixel is not a universal proof a human understood a message. The service returns only what the channel can establish. Provider receipts can be late, so status responses include update time and distinguish absence of evidence from evidence of failure.

13Correctness deep dive

The recipient database orders preference changes and send-authorization transactions together. This determines whether an opt-out committed before permission to send; a provider call already authorized may still complete afterward.

authorizeSend(deliveryId, worker):
  begin transaction at recipient owner
  lock Preference(user), Delivery(deliveryId)
  if delivery already terminal: return stored result
  if delivery already dispatching/unknown: return recovery-required
  if current policy forbids category/channel:
      mark suppressed with policy version; commit; return no-send
  if not due or destination invalid: defer/reject; commit; return no-send
  record dispatching, policyVersion, destinationVersion,
         stable providerKey=deliveryId, attempt identity
  commit; return authorized immutable send parameters
Race or crash Durable fact Safe next action
W1 and W2 both claim queued d-email-44 One transaction changes it to dispatching Other worker gets recovery-required, not a new send
W1 pauses after authorization Dispatching intent persists Investigate/reconcile; do not infer no call occurred
Provider accepts, reply lost Remote effect may exist; local outcome unknown Same delivery key or stable-reference lookup where supported
Provider has neither mechanism No protocol proves whether it sent Apply declared duplicate-versus-miss policy, record decision
Duplicate in-app task Unique inbox delivery ID already exists Return that item; no duplicate notification

For in-app delivery, there is no external-send gap: lock the same preference and delivery rows, check current policy, then insert the unique inbox item and mark the delivery complete in one transaction. An opt-out that commits first suppresses this insertion too. A repeated task returns the already-committed outcome rather than inserting a second item.

sequence · send-raceOpt-out and dispatch meet at the owner

Marketing is suppressed after opt-out version 19. The separately permitted shipping update proceeds; a lost provider reply leaves its outcome unknown. Previously authorized external work retains the stated in-flight limitation.

Opt-out and dispatch meet at the ownerMarketing is suppressed after opt-out version 19. The separately permitted shipping update proceeds; a lost provider reply leaves its outcome unknown. Previously authorized external work retains the stated in-flight limitation. user to owner: Opt-out via authenticated API; owner to owner: Commit preference version 19; worker to owner: Claim queued marketing delivery; owner to owner: Lock policy + delivery; policy forbids; owner to worker: Suppressed; no send authorization; worker to owner: Claim separate permitted ship-o81; owner to worker: Commit authorization; key d-email-44; worker to provider: Send with stable d-email-44; provider to provider: Accept message; provider to worker: Reply lost; worker to owner: Record unknown; preserve delivery ID; worker to provider: Query/retry same key if supportedPARTICIPANTRecipient U7PARTICIPANTChannel workerPARTICIPANTRecipient authorityAPI + DBPARTICIPANTProvider1. Opt-out via authenticated API2. Commit preferenceversion 193. Claim queued marketingdelivery4. Lock policy + delivery;policy forbids5. Suppressed; no sendauthorization6. Claim separate permittedship-o817. Commit authorization; keyd-email-448. Send with stable d-email-449. Accept message10. Reply lost11. Record unknown;preserve delivery ID12. Query/retry same key if supportedsyncreturnblocked
Read each connection in order
  1. syncOpt-out via authenticated APIRecipient U7 → Recipient authority API + DB
  2. syncCommit preference version 19Recipient authority API + DB → Recipient authority API + DB
  3. syncClaim queued marketing deliveryChannel worker → Recipient authority API + DB
  4. syncLock policy + delivery; policy forbidsRecipient authority API + DB → Recipient authority API + DB
  5. returnSuppressed; no send authorizationRecipient authority API + DB → Channel worker
  6. syncClaim separate permitted ship-o81Channel worker → Recipient authority API + DB
  7. returnCommit authorization; key d-email-44Recipient authority API + DB → Channel worker
  8. syncSend with stable d-email-44Channel worker → Provider
  9. syncAccept messageProvider → Provider
  10. blockedReply lostProvider → Channel worker
  11. syncRecord unknown; preserve delivery IDChannel worker → Recipient authority API + DB
  12. syncQuery/retry same key if supportedChannel worker → Provider

14Failure and recovery

Before choosing a recovery action, distinguish a confirmed provider rejection from an unknown outcome. A timeout belongs to the second category until evidence resolves it; retrying or switching providers must respect that uncertainty.

Failure or race Required response and boundary
Provider accepted; response lost The failure sequence is: provider sends or accepts the email, network fails, and our worker sees a timeout before saving provider-902. Repeating the call may send another email. If the provider offers idempotency, retry the same delivery key within its documented retention window. If it offers lookup by a stable client reference, reconcile first. If neither exists, mark the result unknown and choose a product policy balancing duplicate risk against missed delivery. No local queue setting removes this uncertainty.
Unknown-provider failover and retry limits Do not immediately fail over an unknown send to another provider: the second provider cannot deduplicate the first provider's side effect. Confirm failure where possible or accept and document the duplicate risk for that category. A terminal invalid address should not retry forever; temporary rate limits use exponential backoff, random jitter, and provider retry hints. Exhausted work goes to a reviewable dead-letter queue with reason and safe replay controls.
Duplicate or reordered callbacks Callbacks may repeat or arrive out of order. When a provider supplies a stable event ID, deduplicate it in the callback inbox. Otherwise make repeated facts idempotent using the documented account/message/status fields; do not assume every callback API has a unique event ID. Verify signatures, persist raw event facts under limited retention, and apply a channel-specific state model. A callback timeout should cause safe redelivery, not loss.
Consent authority unavailable If a network partition prevents the recipient database from safely committing writes, new send authorizations stop. Workers must not use an old policy cache to continue marketing merely because the provider is reachable. Already authorized in-flight effects can still complete; their callbacks are durably accepted when the service can do so and otherwise retried under the provider contract. In-app posting waits for its database authority rather than duplicating state elsewhere.
Provider backlog and overload During overload, separate retry and new-work budgets. At 500/s continuing arrivals and 1,000/s provider capacity, the 300,000 backlog needs ten minutes to drain. If the provider recovers with only 500/s capacity, the backlog never decreases. Defer marketing, preserve transactional reserve and reject impossible new deadlines before making an acceptance promise. Oldest eligible age is the useful recovery signal, not simply worker count.
Rollback or callback-processing outage A configuration rollback must not replay already accepted external messages with new IDs. Preserve intent/delivery identities through redeployment and dead-letter replay. A callback processor outage may delay status without delaying actual sends; distinguish those incidents so operators do not trigger a duplicate-delivery campaign while trying to repair missing status.

15Operations, security, and cost

Encrypt destination records and restrict template editing separately from sending. Validate parameter types and escape content for the output channel; an order service cannot turn a template field into arbitrary executable markup. Keep secrets and personal message bodies out of queue names, logs and high-cardinality metric labels. Verify provider signatures and preserve enough evidence to diagnose a disputed transition without retaining every sensitive body indefinitely.

Quiet-hour scheduling uses a named time zone, such as America/New_York, and an explicit daylight-saving policy so the scheduled date determines the applicable UTC offset. A local reminder time that does not exist can move to the next valid time or be skipped under the product contract; an ambiguous repeated time must identify one or both occurrences deliberately. Store the resolved scheduled instant and rule version so a worker restart does not reinterpret the same reminder differently.

Estimate costs from each channel’s attempt count and contracted per-attempt price, then add metadata storage, queue/worker work and recovery calls. At 100 million attempts/day, reducing unnecessary retries by 5% removes five million attempts/day, but only if the retry reduction does not worsen the promised outcome. Thirty-day metadata already occupies roughly 3 TB logical; three copies imply about 9 TB before indexes. Immutable template reuse avoids storing many identical large bodies.

Monitor oldest eligible age and the 30-second transactional-provider objective by tenant/channel, unknown outcomes, bounce rates, suppression decisions and retry amplification. Alert through an independent channel so this system’s outage cannot suppress its own incident notification. Canary template and adapter changes on controlled recipients; test opt-out races, duplicated callbacks, DST scheduling and provider recovery before expanding.

16Decision ledger and limitations

The design keeps acceptance and consent decisions in our database while delegating email, SMS and push transport to providers. These tradeoffs follow from that split in control and from the need to protect urgent messages during bursts.

Choice Benefit Cost
Durable intent plus delivery rows Traceable retries and partial results More state than one fire-and-forget call
Separate priority queues Transactional latency protection Fairness and reservation policy
Recheck preferences before send Honors recent opt-outs Additional policy lookup
Provider idempotency/reconciliation Reduces duplicate sends Depends on documented external behavior
Additional decision Consequence Reconsider when
Recipient-owned policy and delivery state Local authorization transaction; provider queues need routing metadata A user’s state becomes too hot or cross-region requirements change
Stable external delivery identity Retries can use a provider’s dedupe contract Provider retention/capability cannot meet the retry horizon
Unknown outcome as first-class state Honest visibility and safe recovery Product explicitly chooses duplicate risk over a missed message
Authorize before network call Opt-out order is testable Harder recall requirements need a provider-supported cancellation protocol
Distinct acceptance/delivery/read facts Avoids false success claims A channel adds a verified new receipt type

The next bottleneck is likely a provider quota or one merchant’s campaign fairness, not generic API CPU. Adding workers should follow that measurement. A delayed marketing message may be the correct outcome when the alternative is missing transactional deadlines or bypassing current preferences.

17Interview closing

“I distinguish the business intent, each channel delivery and each provider attempt. The API durably accepts the notification intent, and a planner records stable delivery IDs with an outbox. A quota-aware dispatcher reserves transactional capacity and schedules work by provider and tenant.

“Before sending, the recipient owner serializes current policy with a dispatch authorization. In-app insertion is unique by delivery ID. Email is an external effect: a lost response becomes unknown, and I reuse a supported provider idempotency key or reconcile the original reference. I never fail over an unknown attempt blindly and call that exactly once. Verified callbacks update channel-specific facts without regressing known delivery.

“The assumed hundred million daily attempts drive quota, retention and retry costs. During a provider outage, recovery capacity must exceed new arrivals. I will measure eligible queue age, unknown-outcome age and per-channel expense. Some deliveries remain pending or unknown until evidence arrives; reporting success earlier would mislead the caller.”

Interviewer: “Marketing must stop the instant a user opts out.” Candidate: “I can stop authorizations ordered after the opt-out. For already authorized or accepted external sends, I need provider-supported cancellation and a defined acknowledgment boundary. Without that capability, I cannot promise recall merely by deleting a queued row.”

Practise the interview questions

Say your answer aloud before opening the model answer. Then answer the follow-up and compare the reasoning.

Foundation · Question 1

Why do you need three IDs for one notification?

Reveal a model answer

The order event identifies the logical intent, the channel delivery identifies where it goes, and the attempt identifies a provider call. Retries should add attempts without inventing new intents or channels.

What the answer must demonstrate: Retry identity must survive retries.

Applied · Question 2

The email provider times out after accepting your request. What do you do?

Reveal a model answer

I treat the result as unknown. I retry the same idempotency key or query a stable provider reference if supported. Without either capability, I apply the category’s explicit duplicate-versus-miss policy.

What the answer must demonstrate: Do not infer rejection from lack of response.

Applied · Question 3

A process dies after inserting a delivery but before queueing it. How is the message sent?

Reveal a model answer

The delivery and its outbox record were committed together. A dispatcher resumes reading unsent outbox records and publishes the stable delivery ID. Repeated queue messages are safe because workers use that identity.

What the answer must demonstrate: Explain both sides of the handoff gap.

Foundation · Question 4

A recipient opts out after a campaign is queued. Do you still send?

Reveal a model answer

I recheck the current applicable preference near sending and suppress work that is no longer permitted. The queued snapshot helps explain planning, but it does not grant permanent consent.

What the answer must demonstrate: Explain whether the opt-out committed before or after send authorization.

Follow-up · Question 5

A provider outage lasts ten minutes. Why might recovery take much longer?

Reveal a model answer

New arrivals continue while the backlog drains. Drain rate is provider capacity minus new traffic, not total capacity. If there is no spare quota, the queue never catches up.

What the answer must demonstrate: Calculate net drain rate.

Follow-up · Question 6

A delivered callback arrives before a sent callback. What happens?

Reveal a model answer

I persist both facts and apply a channel-specific state rule that does not regress delivered to sent. I verify the signature and bind the message to its provider account and delivery. If there is no provider event ID, repeated facts still apply idempotently using the documented message/status identity.

What the answer must demonstrate: Avoid one generic success boolean.

Applied · Question 7

What exact guarantee can you make when a recipient opts out during dispatch?

Reveal a model answer

“The preference update and send authorization serialize at the recipient owner. If opt-out commits first, authorization is suppressed. If authorization already committed, work may be in flight; cancellation is best effort unless the provider offers a stronger protocol.”

What the answer must demonstrate: Identify the transaction that grants permission to send, then explain why opt-out cannot always recall the provider call that follows.

Follow-up · Question 8

A campaign fills the queue while order updates miss their deadline. How do you change scheduling?

Reveal a model answer

“I reserve provider capacity for transactional work and use bounded weighted scheduling across tenants and categories. At a 1,000/s provider quota, an assumed 600/s reserve protects urgent work; unused capacity can be borrowed without erasing the reserve when urgent traffic returns.”

What the answer must demonstrate: Protect urgency without pretending to increase provider quota.

Blank-page exercise · 45 minutes

Build the answer yourself

Design order and marketing notifications, then handle an email-provider timeout after the provider may have accepted recipient U7’s message.

  • Clarify intent, acceptance, delivery and read semantics.
  • Calculate attempts, per-provider quotas and net backlog drain.
  • Draw recipient ownership separately from provider scheduling.
  • Trace opt-out versus send authorization, show why an in-app retry cannot insert a duplicate, and recover a provider call whose outcome is unknown.
  • Compare fairness, provider failover and per-channel cost; give a natural closing answer.

Check that each component and design decision follows from your requirements and workload.

Recall the key ideas

Answer from memory before opening each card. Explain why the choice works and what it costs. Revisit missed cards tomorrow.

Design a notification serviceWhat is the difference between intent, delivery, and attempt?Recall first, then reveal

An intent is the business message, a delivery targets one channel/destination, and an attempt is one provider call.

Business message → recipient/channel delivery → provider call.

Return to lesson
Design a notification serviceA provider call times out. Is it safe to change providers?Recall first, then reveal

Not automatically. The first provider may already have accepted the message, so cross-provider retry can duplicate it.

Unknown is not failed.

Return to lesson
Design a notification serviceWhen must opt-out be checked?Recall first, then reveal

At planning and again in the send-authorization transaction. An opt-out committed before authorization suppresses the send; cancellation after authorization is best effort.

Queueing does not freeze consent.

Return to lesson

Final revision

Summary and interview notes

A notification service durably records intent before independently executing each channel delivery. The recipient database orders opt-outs and send authorizations. Stable delivery IDs let workers recover provider results; an unresolved timeout stays unknown.

Remember these points

  • Keep business intent, channel delivery and transport attempt identities distinct.
  • Check preferences and authorize sending in one ordered transaction; a provider call authorized earlier may be impossible to recall.
  • Freeze recipient and payload before the first external attempt, and retain them across retries of that delivery.
  • Provider acceptance, channel delivery and user reading are different facts.
  • The backlog shrinks only when workers can complete more deliveries than continue to arrive.

Interview tips

  • Start by asking which observable outcome the reliability target measures.
  • Trace an opt-out race and a provider timeout after acceptance; a queue does not solve either by itself.
  • Estimate traffic and expense per channel attempt, including retry amplification.

Important qualifications

  • A provider without idempotency or reliable status lookup requires an explicit duplicate-versus-miss policy.
  • Providers identify callbacks and report delivery differently; some do not supply unique event IDs.

Technical references

Practice marks stay in this browser.