System-design interview · Extended interviews
Design a notification service
Separate notification intents, channel deliveries and provider attempts; design preferences, provider quotas, callbacks and uncertain-outcome recovery.
You will learn to
- Separate a business notification from channel-specific delivery attempts.
- Use durable scheduling without claiming exactly-once external delivery.
- Explain priority, opt-out, quiet hours, and ambiguous provider outcomes.
Practice in this chapter
8 interview questions with model answers and follow-ups.
Go to interview practiceUseful foundations: Message queues, event logs, delivery guarantees, and backpressure · Databases, data models, and ACID transactions · Production readiness: SLI, SLO, observability, and recovery
Workload and timing examples are interview assumptions.
Dotted concept links open the relevant explanation in a new tab.
01Problem and scope
A notification service accepts business events and delivers approved messages through email, push, SMS or an in-app inbox. Distinguish the notification intent, one recipient/channel delivery, and each transport attempt. Give the intent, each recipient/channel delivery and each provider attempt separate IDs so retries can reuse the right operation and status can report which channels succeeded. Event ship-o81 is an example: one intent creates email and in-app deliveries while SMS is suppressed by preference. Provider acceptance, device delivery and user reading are different outcomes.
Interviewer: “Make notifications reliable.” Candidate: “Do we mean durably accepted by our service, accepted by a provider, delivered to a device, or read by a person?” Interviewer: “Track those separately. Transactional order updates are urgent; marketing may wait and must respect opt-outs.” This prevents us from reporting that a person received a message merely because a provider accepted it.
We build a shared notification service used by authenticated product services. They submit approved templates and business event IDs rather than arbitrary destinations and free-form scripts. We support one recipient per intent; a campaign service expands an audience into individual recipient intents while obeying the notification API's admission limits. We do not design an email server, mobile push network or marketing audience builder from scratch.
The delivery guarantee depends on which system performs the final action. Email and in-app differ here: an in-app item can be atomically inserted into our database, whereas an email provider performs a separate effect. A useful design promises one logical intent and recoverable delivery state, then states exactly where duplicate external delivery can remain possible.
02Functional requirements
- Accept an intent. Accept a notification intent for a tenant, event, recipient, category and versioned template. A repeated business event returns the original n44; changed parameters under that identity are a conflict or an explicitly new revision.
- Select eligible channels. Determine eligible channels and destinations from trusted user profiles and preferences. For ship-o81, create email d-email-44 and in-app d-app-44, and record why SMS is suppressed.
- Schedule delivery. Schedule immediately or at a specified instant/time-zone policy. Quiet hours may defer eligible work; urgent exceptions must be part of the recipient/category contract.
- Execute each channel. Deliver in-app items durably and invoke channel providers. Expose queued, dispatching, provider-accepted, known-delivered, failed, suppressed and unknown meanings per channel.
- Apply provider receipts. Accept authenticated provider receipts, update known outcomes, and let callers inspect partial success. Email success does not imply push success.
- Manage preferences and history. Change preferences, cancel unstarted work and inspect user-visible notification history. Cancellation after an external effect is not guaranteed recall.
Acceptance boundaries
Destinations are references to verified profile records, not untrusted phone numbers supplied by every caller. Templates have typed parameters and immutable versions so retries can reproduce the intended message. A rendering error is a terminal configuration problem to investigate, not a reason to retry a provider request indefinitely. The service should explain why a message did not send, not only whether a worker ran.
03Non-functional requirements
- Admission latency and availability. Assume p95 intent admission below 150 ms and 99.95% eligible intent-admission availability. Both exclude downstream delivery.
- Transactional delivery deadline. For this exercise, 99% of eligible, immediately due transactional channel deliveries receive confirmed provider acceptance within 30 seconds of durable intent acceptance. Scheduled marketing can wait.
- Durability. Acknowledged intents survive one database-zone failure under the configured replication protocol.
- Retention. Assume thirty days of detailed attempt metadata and ninety days of user-visible in-app history. Sensitive rendered bodies may have shorter retention or no persistence.
- Authorization and content safety. Enforce permitted channels/categories, versioned templates and typed parameters. Callers cannot inject arbitrary markup or change destinations without authorization.
- Consent ordering. An opt-out committed before the send-authorization transaction prevents that authorization. Recheck near dispatch; already authorized or externally accepted work may remain in flight and cannot always be recalled.
Measure the right delivery outcome
The 30-second denominator excludes policy-suppressed channels and separately scheduled work; provider errors and unknown outcomes count as misses at the deadline. In-app posting has a separate completion metric because it invokes no provider. Provider acceptance, device delivery and human reading are distinct facts. A bounced email and an SMS delivery receipt have different channel semantics; store them separately.
Category and time-zone policy
Support transactional updates, scheduled reminders and marketing over email, push, SMS and in-app channels. Bulk audience selection can remain a separate campaign system emitting bounded intents. Quiet hours use the recipient's time zone, not the worker's clock zone; define daylight-saving transitions and user time-zone changes. Only a specific user-approved urgent security exception can bypass the applicable preference. Putting marketing in a high-priority queue does not grant that exception.
Unknown external outcomes
An external timeout leaves two possibilities: the provider did nothing, or it acted but the reply was lost. Recovery depends on whether the provider can recognize a repeated operation or look up the original result.
| Provider capability | Required category policy |
|---|---|
| Idempotency or reliable lookup exists | Recover under that documented contract |
| Neither capability exists | Explicitly balance duplicates against missed delivery: a transactional update may tolerate a reviewed duplicate, while marketing may prefer no retry |
This external-effect contract is separate from the uptime percentage.
04Capacity estimates
Treat attempts, channel deliveries and business intents as different units. The ratios below are workload assumptions.
| Estimate | Arithmetic | Boundary |
|---|---|---|
| Provider attempts | 100 million/day ÷ 86,400 ≈ 1,157/s | Includes normal retries |
| Attempt burst | Tenfold average ≈ 11,600/s | Still bounded by each provider's contracted quota |
| Deliveries and intents | At 1.1 attempts/external delivery and two eligible external-channel deliveries/intent: 90.9 million external deliveries and 45.5 million intents/day | About 526 average intents/s and 5,260/s tenfold peak |
| Attempt metadata | 100 million × 1 KB = 100 GB/day; 3 TB for 30 days | Before indexes and replicas |
| Provider payload | 11,600/s × 2 KB = 23.2 MB/s | Before protocol overhead |
| In-app reads | 20 million/day ≈ 231/s; tenfold peak ≈ 2,315/s | Provision independently from sending |
| Failed-provider backlog | 500 attempts/s × ten minutes = 300,000 attempts | Retain the 300,000 pending attempts and pace provider calls within quota |
| Backlog drain | 1,000/s capacity − 500/s new arrivals = 500/s net | Another ten minutes to drain 300,000 |
Quotas and traffic mix
These averages count provider-backed channels only; in-app deliveries add local database work without a provider call. The ship-o81 example has one email and one in-app delivery, so its mix differs from this fleet-wide assumption. If each intent produces two external deliveries, its rate can be half the external-delivery rate; retry storms can reverse the relationship between intent and provider-call load. Measure failed addresses, suppressed channels and retries separately. Provider quota, reputation and cost usually constrain sending before raw bandwidth. More workers cannot exceed a downstream quota.
Retention and recovery budgets
Store sensitive rendered bodies only when required under appropriate retention. Queued work can reference immutable templates and safe parameter records instead of duplicating a full rendered body. Spread recovery retries instead of launching the whole backlog at once.
Separate transactional and marketing capacity, bound per-tenant queued volume, and reject impossible deadlines before acceptance. A campaign must not consume every send slot.
05APIs and contracts
The application submits a business intent; channel adapters later obtain provider results. These endpoints expose those different stages, so callers can distinguish acceptance of the notification request from progress toward delivery.
| API | Example | Response |
|---|---|---|
| Create intent | POST /notifications with {"eventId":"ship-o81","userId":"u7","template":"shipped-v3","params":{"orderId":"o81"}} |
n44, accepted |
| Inspect | GET /notifications/n44 |
Per-channel queued, accepted, delivered, failed, or unknown |
| Preference | PUT /users/u7/preferences |
Versioned channel/category policy |
| Provider callback | Signed provider event with message ID | Durable acknowledgment of callback ingestion |
Create requests include a category, template version, optional schedule, and a stable event identity. The response is 202 with n44 and per-channel planning status; it is not a delivery receipt. Return 409 when the same identity contains different canonical parameters, 403 for an unauthorized sender/category and 429 when the caller exceeds its accepted-work budget. A permanently invalid template is rejected before committing work when detectable.
Preference updates carry an expected version. A successful response returns policy version 19, and a later dispatch authorization must use current state rather than trusting the planner’s old version 18. Cancellation addresses a logical intent or delivery and reports whether it prevented authorization, requested best-effort cancellation of in-flight work, or arrived after a known external effect.
History uses GET /users/u7/inbox?after=(createdAt,itemId)&limit=50, with verified user/tenant scope. Provider callbacks go to dedicated endpoints that verify the provider-specific signature and bind message references to the correct channel/account. They return success after durable inbox acceptance; asynchronous parsing/application can then retry without relying on the provider keeping the HTTP connection open.
Before the first send authorization, revalidate the destination and freeze the exact recipient, template version and rendered parameters for that delivery. Once any provider attempt may have happened, never change those bytes under the same provider key. An address change cannot turn recovery of the old delivery into a send to a new address. A deliberate replacement is a new logical delivery only after applying the category’s policy to the old delivery’s unresolved outcome.
Push registration and device fanout
A push provider routes messages using a registration token for an app installation. A user can have several installations, and their tokens can change. Registration operations maintain the account-to-installation binding before the planner chooses push destinations.
| Device operation | Contract |
|---|---|
| Register or refresh | Authenticated PUT /users/me/devices/{installationId} submits the provider registration/token and platform; the server binds it to the current account and returns a destination version. |
| Sign out or unregister | Disable that account/device binding. A device identifier or push token alone is not proof of account ownership. |
| Plan push | Expand one recipient's push intent into deliveries for its eligible registered devices; each delivery names one immutable destination ID and version. |
| Resolve provider rejection | A definitive invalid-registration response disables only the matching registration version; an old response cannot invalidate a newer refreshed token. |
A user may have several devices, and reinstall or provider rotation may change a device's registration. Keep last-refreshed time and prune stale registrations under the chosen provider/product policy. Revalidate registration eligibility during send authorization, then freeze the exact registration for that delivery just like an email address. Never replace the token on an uncertain attempt under the old provider key. Expiration/TTL and collapse keys are channel policies: use them for replaceable updates when appropriate, not as proof that independent business alerts were delivered. Provider acceptance still does not mean the device displayed or the user read a push. FCM registration management.
06Data model and access patterns
These records preserve three kinds of information: what we planned to send, whether it was authorized, and what actually happened. An outbox stores pending dispatch events in the delivery transaction; an inbox durably accepts incoming provider events before they are processed. Both let a handoff resume after a process crash.
| Record | Key and important fields | Query or invariant |
|---|---|---|
| Intent | unique (tenant,eventId,userId,category); payloadHash, templateVersion |
Deduplicate ship-o81 and reproduce its meaning |
| Delivery | (intentId,channel,destinationId,destinationVersion); state, dueAt, generation |
One logical destination across attempts |
| Attempt | (deliveryId,attemptNumber); providerKey, authorizationVersion, outcome |
Explain calls and unknown responses |
| Push registration | (tenant,userId,installationId); destinationId, version, provider registration reference, lastRefreshedAt, enabled |
Multiple devices per user; version-guarded refresh and invalidation |
| Preference | (tenant,userId); categories/channels, version, timeZone |
Current permission at dispatch authorization |
| Template | (tenant,name,version); typed schema, approved body |
Immutable rendering and safe reuse |
| Inbox/outbox | unique event/dispatch identities and processing state | Recoverable internal handoffs |
| In-app item | unique deliveryId; recipient, contentRef, createdAt | Retry-safe insertion and ordered user reads |
Partition user-facing intent, preference, delivery and in-app state by tenant plus recipient bucket. That lets send authorization serialize against the relevant preference update within one owner. Provider quotas group work differently from recipient storage: queues can be grouped by provider/channel while each message carries the key needed to locate its delivery record.
Index pending deliveries by (state,nextAttemptAt,deliveryId), recipient history by (userId,createdAt,itemId), and provider references for callbacks. Queues contain IDs and minimal routing metadata, not the only copy of content or authorization. Sensitive destination and parameter records have restricted access and retention; metrics never use arbitrary email addresses as labels.
07Basic working design
Start with one authenticated API, one database, and a worker that polls an indexed pending-delivery table. The order service submits ship-o81 after its own order transaction through its outbox or equivalent reliable publication. Our API cannot atomically commit with the caller’s separate order database, so the caller must retry the same event identity until it knows acceptance.
The API transaction creates n44 and a planning task. The planner reads recipient U7’s preferences and template shipped-v3, creates two delivery rows and records SMS suppression. For in-app delivery, the worker locks the current preference and delivery rows, checks eligibility, and commits the unique d-app-44 item and completed delivery state in that same transaction. The email provider call occurs outside database locks after its separate send authorization. It saves the provider’s returned reference and reported acceptance. Recipient U7’s order remains accessible throughout; checkout does not wait for email.
This baseline is useful before a broker or many workers exist. Its pending rows survive process restart, duplicate intent submission returns n44, and the in-app channel has a clear commit boundary. It still needs a stated policy for uncertain email outcomes. A single worker is not a guarantee against duplicates: a crash after an external effect but before saving its reference already creates uncertainty.
Measure due-work scan cost, queue age and provider latency. The first benchmark should include slow responses and invalid destinations, because they consume worker slots differently from a stream of immediate successful calls.
Recipient U7’s order does not wait for external email; the notification intent is durable before acceptance.
Read each connection in order
- syncSubmit ship-o81Order service → Notification API
- syncCommit n44 and planning taskNotification API → Intent / delivery / preference DB
- asyncFind pending channel workIntent / delivery / preference DB → Planning and channel worker
- syncAuthorize send / commit in-app itemPlanning and channel worker → Intent / delivery / preference DB
- syncSend d-email-44 outside transactionPlanning and channel worker → Email provider
- syncRecord provider fact or unknownPlanning and channel worker → Intent / delivery / preference DB
08Find the baseline flaws
Assume a worker performs provider calls sequentially and each takes 200 ms. Its throughput is about five attempts/s. The assumed 11,600/s peak requires roughly 2,320 concurrent calls at that average latency if downstream quotas permit it. Merely adding a second worker cannot satisfy the target, and unrestricted concurrency may violate provider limits.
Polling all pending rows without a due-time index becomes expensive as the ten-minute outage adds 300,000 records. Repeatedly scanning that set can overwhelm the database even when no call is currently allowed. We need scheduled eligible work and provider-specific admission, not increasingly aggressive polling.
Now recipient U7 opts out of marketing at policy version 19 after the planner used version 18. If the worker sends using the old planned snapshot, the queued job has effectively granted permanent consent. Dispatch authorization must re-evaluate the current policy at the chosen owner and record the version used. The boundary must also admit that opt-out can race with an already authorized network call.
Finally, the provider accepts d-email-44 but the response disappears. A retry with a new attempt-based idempotency key can send twice. The correct external identity remains d-email-44; attempts are diagnostics, not new user intent. Database uniqueness solves the in-app case, but cannot force an arbitrary provider to remember that identity.
09Improve the design, step by step
First, separate durable planning from provider execution. Rising pending-table scans and several thousand concurrent calls trigger an outbox-to-queue relay and bounded provider workers. The planner commits each delivery and dispatch event together; the relay publishes its stable ID. More workers can now execute deliveries, while slow provider calls do not hold up new intent requests. The cost is broker storage, duplicate deliveries and another service to monitor and recover. A lost publish acknowledgment can cause repetition, so consumers recheck durable delivery state. Indexed database polling remains simpler when it meets throughput and scheduling needs; we add a broker when measurements show a need to scale workers separately.
Second, reserve urgency and tenant fairness. A large campaign can consume all provider tokens while ship-o81 misses its thirty-second objective. Divide work into transactional, scheduled and retry lanes, with weighted scheduling and reserved transactional capacity within each provider quota. At 1,000 allowed attempts/s, reserving an illustrative 600 for transactional work prevents a marketing burst from taking those slots; unused capacity can be borrowed under a bounded rule. Urgent messages get protected capacity; the provider’s total quota stays the same. Costs are scheduler complexity and potentially delayed marketing. Strict priority risks starving low-priority work; FIFO is simpler but fails the urgent target under campaigns. Select the policy from agreed deadlines and monitor age by lane.
Third, partition state by recipient ownership. Several thousand peak intents/s and ninety-day inbox history motivate tenant/recipient buckets. A directory routes recipient U7’s intent and preference operations to the same owner, keeping the permission check and dispatch claim in a local transaction. Queue assignment still follows provider rate domains. Each database holds a smaller active dataset and handles a different group of recipients; routing and migration become more complex. The new risk is authorizing against stale preference state during a move. Move preferences, deliveries and pending authorizations together, then transfer ownership using a new owner version. A larger single database is preferable when it still meets targets; splitting preference and delivery authorities prematurely would complicate the consent boundary.
Fourth, add durable receipts and recovery. Lost responses and reordered callbacks motivate a verified callback inbox plus an outcome reconciler. They save provider facts and follow each channel’s status rules; an old sent receipt cannot overwrite a delivered status. This improves visibility and recovers missed acknowledgments; it costs provider queries and additional retained evidence. A malformed or incorrectly matched receipt is the new risk, so match provider account, message reference and delivery identity. Polling every message forever is a rejected alternative: use provider capabilities, age-based recovery and terminal-state rules. Where neither query nor idempotency exists, preserve unknown status and the explicit duplicate-versus-miss policy.
Template caches improve rendering only after these boundaries are correct. An immutable version is safe to reuse; a stale preference cache is a different kind of data and cannot inherit that lifetime.
10Detailed architecture
The scaling changes produce two different groupings: recipient-owned records keep consent checks atomic, while provider-oriented queues enforce sending quotas. A queued delivery carries enough routing information to return to its recipient owner before it is authorized.
Recipient-owned state
The ingress API authenticates product services and authorizes template/category use. It routes tenant/recipient state through the owner directory. The owner database stores the intent, policy, deliveries, in-app items and outbox. Its synchronous replicas protect accepted state under the stated zone-failure model.
Planning and quota-aware dispatch
A planner expands approved intents into channel work. It uses immutable templates and a planning policy snapshot, but the dispatch claim rechecks current permission. The outbox relay publishes delivery IDs to a durable queue. A quota-aware dispatcher groups those IDs by provider/channel and urgency, then releases work only within per-provider and per-tenant budgets.
Channel execution and recovery
Channel adapters render the approved version, claim a send authorization at the recipient owner and invoke the external provider. The in-app adapter instead commits a unique inbox item locally. A callback receiver verifies and persists provider evidence before acknowledgment; a reconciler queries recoverable unknown outcomes. Both use the same delivery state transition rules.
Timing boundaries and implementation
The diagram separates the database that owns a recipient's preferences from the queue that schedules provider calls. Grouping work by provider helps enforce provider quotas, but send permission still comes from the recipient's current preference record. The API replies once the intent is durably saved. Planning, delivery, callback processing and recovery are asynchronous. Status reads expose these stages individually so “queued,” “accepted” and “delivered” never collapse into one convenient but misleading boolean.
One implementation starts with PostgreSQL transactions for recipient policy, deliveries and outbox, plus indexed due-work polling. Add SQS standard queues when independent worker scaling and buffering justify them, retaining the database as the authority. Standard queues can repeat deliveries; their visibility timeout is a worker-coordination aid, not an external-send guarantee. Provider adapters isolate channel-specific status, authentication and quota rules instead of forcing all providers into one success flag.
The recipient database decides whether a send is permitted. Provider/channel queues control when permitted calls run within quota; queue priority cannot override an opt-out.
Read each connection in order
- sync1. Intent, preference, status or inboxProduct services / user clients → Authenticated notification API
- sync2. Resolve tenant/recipient ownerAuthenticated notification API → Recipient-owner directory
- sync3. Commit intent / policy; read statusAuthenticated notification API → Recipient DB + synchronous replicas
- async4. Recover pending planningRecipient DB + synchronous replicas → Preference/template planner
- syncLoad approved immutable templatePreference/template planner → Versioned template store/cache
- syncCommit deliveries + dispatch outboxPreference/template planner → Recipient DB + synchronous replicas
- async5. Read committed dispatch IDsRecipient DB + synchronous replicas → Outbox relay
- asyncPublish stable delivery IDOutbox relay → Durable delivery queues
- async6. Select due category/provider workDurable delivery queues → Quota and fairness dispatcher
- async7. Release within quota budgetQuota and fairness dispatcher → Channel adapters / in-app worker
- sync8. Atomic policy/send authorizationChannel adapters / in-app worker → Recipient DB + synchronous replicas
- sync9. Send stable delivery keyChannel adapters / in-app worker → External channel providers
- syncCommit in-app item / outcomeChannel adapters / in-app worker → Recipient DB + synchronous replicas
- asyncSigned channel outcomeExternal channel providers → Verified receipt inbox/applier
- sync10. Persist receipt; apply stateVerified receipt inbox/applier → Recipient DB + synchronous replicas
- syncQuery original external referenceUnknown-outcome reconciler → External channel providers
- syncResolve unknown under state rulesUnknown-outcome reconciler → Recipient DB + synchronous replicas
11Write path and acknowledgement
Durable intent acceptance precedes asynchronous planning and channel execution. Each external attempt retains its logical delivery identity across retries.
- The API authenticates the order service and inserts intent
n44forship-o81. A retry with the same event identity returnsn44; conflicting content is rejected or versioned explicitly. - A planner evaluates recipient U7's preferences and template
shipped-v3, producing email deliveryd-email-44and in-app deliveryd-app-44. Their rows and dispatch events commit together. SMS is suppressed with an auditable reason. - The in-app worker locks and checks current preferences, suppressing a disallowed delivery or atomically inserting the unique
d-app-44inbox item with its completed outcome. Retries return the existing outcome. The email worker claims dued-email-44, rechecks current consent/address validity, records attempt 1, and calls the provider using a stable idempotency key if that provider supports one. - The provider returns message ID
provider-902. The worker recordsaccepted; a later authenticated callback changes the known delivery result. A queued callback arriving after a delivered callback must not blindly move the delivery backward. - Recipient U7 can inspect the order immediately while channel work proceeds. The notification API reports each channel's actual state rather than turning partial success into one misleading boolean.
The planner’s delivery insertion and dispatch outbox share a transaction, so a crash before publication cannot strand an accepted plan. A repeated planner task observes the existing stable deliveries. For delayed work, a scheduler releases it when due rather than repeatedly calling a provider before quiet hours end.
The dispatch transaction records whether the message is authorized and the policy/destination versions used. It then releases database locks before the external call. Owner versions can reject stale workers inside our service. They cannot stop an email provider that does not check those versions. Therefore the stable delivery key and provider-supported status recovery remain necessary for the uncertain interval.
If the provider call succeeds but the database update fails, the delivery remains recoverable as dispatching/unknown. The worker does not mark it definitively failed merely to make a retry easier. The worker acknowledges the queue only after saving its progress. If that acknowledgment is lost, the next worker checks the saved delivery before deciding whether it can call the provider.
12Read and delivery path
Status reports known facts per channel, while inbox reads enforce recipient ownership and cursor ordering.
- The order service queries n44 using its tenant-scoped credentials, or recipient U7 queries the recipient-owned inbox with a user credential. Each path checks resource ownership rather than relying on an opaque ID.
- The router resolves the recipient bucket. An immediate status request can use the owner to avoid returning not-found for an intent already acknowledged; history may use a replica only under an explicit freshness policy.
- The API combines intent status with channel facts: in-app committed, email provider-accepted, SMS suppressed by policy. An unknown attempt is exposed separately from a definitive invalid-address failure.
- For the inbox, the store seeks
(u7,createdAt,itemId)and returns the next fifty entries. The cursor carries the last tuple. New arrivals may appear on a refresh; a cursor alone is not a frozen snapshot. - Opening an item retrieves permitted template/content references and applies current visibility rules. A deleted or tenant-restricted order should not leak through an old notification preview; retain only the necessary message content and links.
Read/open tracking, when supported and permitted, is another event with its own semantics. A mail tracking pixel is not a universal proof a human understood a message. The service returns only what the channel can establish. Provider receipts can be late, so status responses include update time and distinguish absence of evidence from evidence of failure.
13Correctness deep dive
The recipient database orders preference changes and send-authorization transactions together. This determines whether an opt-out committed before permission to send; a provider call already authorized may still complete afterward.
authorizeSend(deliveryId, worker):
begin transaction at recipient owner
lock Preference(user), Delivery(deliveryId)
if delivery already terminal: return stored result
if delivery already dispatching/unknown: return recovery-required
if current policy forbids category/channel:
mark suppressed with policy version; commit; return no-send
if not due or destination invalid: defer/reject; commit; return no-send
record dispatching, policyVersion, destinationVersion,
stable providerKey=deliveryId, attempt identity
commit; return authorized immutable send parameters
| Race or crash | Durable fact | Safe next action |
|---|---|---|
| W1 and W2 both claim queued d-email-44 | One transaction changes it to dispatching | Other worker gets recovery-required, not a new send |
| W1 pauses after authorization | Dispatching intent persists | Investigate/reconcile; do not infer no call occurred |
| Provider accepts, reply lost | Remote effect may exist; local outcome unknown | Same delivery key or stable-reference lookup where supported |
| Provider has neither mechanism | No protocol proves whether it sent | Apply declared duplicate-versus-miss policy, record decision |
| Duplicate in-app task | Unique inbox delivery ID already exists | Return that item; no duplicate notification |
For in-app delivery, there is no external-send gap: lock the same preference and delivery rows, check current policy, then insert the unique inbox item and mark the delivery complete in one transaction. An opt-out that commits first suppresses this insertion too. A repeated task returns the already-committed outcome rather than inserting a second item.
Marketing is suppressed after opt-out version 19. The separately permitted shipping update proceeds; a lost provider reply leaves its outcome unknown. Previously authorized external work retains the stated in-flight limitation.
Read each connection in order
- syncOpt-out via authenticated APIRecipient U7 → Recipient authority API + DB
- syncCommit preference version 19Recipient authority API + DB → Recipient authority API + DB
- syncClaim queued marketing deliveryChannel worker → Recipient authority API + DB
- syncLock policy + delivery; policy forbidsRecipient authority API + DB → Recipient authority API + DB
- returnSuppressed; no send authorizationRecipient authority API + DB → Channel worker
- syncClaim separate permitted ship-o81Channel worker → Recipient authority API + DB
- returnCommit authorization; key d-email-44Recipient authority API + DB → Channel worker
- syncSend with stable d-email-44Channel worker → Provider
- syncAccept messageProvider → Provider
- blockedReply lostProvider → Channel worker
- syncRecord unknown; preserve delivery IDChannel worker → Recipient authority API + DB
- syncQuery/retry same key if supportedChannel worker → Provider
14Failure and recovery
Before choosing a recovery action, distinguish a confirmed provider rejection from an unknown outcome. A timeout belongs to the second category until evidence resolves it; retrying or switching providers must respect that uncertainty.
| Failure or race | Required response and boundary |
|---|---|
| Provider accepted; response lost | The failure sequence is: provider sends or accepts the email, network fails, and our worker sees a timeout before saving provider-902. Repeating the call may send another email. If the provider offers idempotency, retry the same delivery key within its documented retention window. If it offers lookup by a stable client reference, reconcile first. If neither exists, mark the result unknown and choose a product policy balancing duplicate risk against missed delivery. No local queue setting removes this uncertainty. |
| Unknown-provider failover and retry limits | Do not immediately fail over an unknown send to another provider: the second provider cannot deduplicate the first provider's side effect. Confirm failure where possible or accept and document the duplicate risk for that category. A terminal invalid address should not retry forever; temporary rate limits use exponential backoff, random jitter, and provider retry hints. Exhausted work goes to a reviewable dead-letter queue with reason and safe replay controls. |
| Duplicate or reordered callbacks | Callbacks may repeat or arrive out of order. When a provider supplies a stable event ID, deduplicate it in the callback inbox. Otherwise make repeated facts idempotent using the documented account/message/status fields; do not assume every callback API has a unique event ID. Verify signatures, persist raw event facts under limited retention, and apply a channel-specific state model. A callback timeout should cause safe redelivery, not loss. |
| Consent authority unavailable | If a network partition prevents the recipient database from safely committing writes, new send authorizations stop. Workers must not use an old policy cache to continue marketing merely because the provider is reachable. Already authorized in-flight effects can still complete; their callbacks are durably accepted when the service can do so and otherwise retried under the provider contract. In-app posting waits for its database authority rather than duplicating state elsewhere. |
| Provider backlog and overload | During overload, separate retry and new-work budgets. At 500/s continuing arrivals and 1,000/s provider capacity, the 300,000 backlog needs ten minutes to drain. If the provider recovers with only 500/s capacity, the backlog never decreases. Defer marketing, preserve transactional reserve and reject impossible new deadlines before making an acceptance promise. Oldest eligible age is the useful recovery signal, not simply worker count. |
| Rollback or callback-processing outage | A configuration rollback must not replay already accepted external messages with new IDs. Preserve intent/delivery identities through redeployment and dead-letter replay. A callback processor outage may delay status without delaying actual sends; distinguish those incidents so operators do not trigger a duplicate-delivery campaign while trying to repair missing status. |
15Operations, security, and cost
Encrypt destination records and restrict template editing separately from sending. Validate parameter types and escape content for the output channel; an order service cannot turn a template field into arbitrary executable markup. Keep secrets and personal message bodies out of queue names, logs and high-cardinality metric labels. Verify provider signatures and preserve enough evidence to diagnose a disputed transition without retaining every sensitive body indefinitely.
Quiet-hour scheduling uses a named time zone, such as America/New_York, and an explicit daylight-saving policy so the scheduled date determines the applicable UTC offset. A local reminder time that does not exist can move to the next valid time or be skipped under the product contract; an ambiguous repeated time must identify one or both occurrences deliberately. Store the resolved scheduled instant and rule version so a worker restart does not reinterpret the same reminder differently.
Estimate costs from each channel’s attempt count and contracted per-attempt price, then add metadata storage, queue/worker work and recovery calls. At 100 million attempts/day, reducing unnecessary retries by 5% removes five million attempts/day, but only if the retry reduction does not worsen the promised outcome. Thirty-day metadata already occupies roughly 3 TB logical; three copies imply about 9 TB before indexes. Immutable template reuse avoids storing many identical large bodies.
Monitor oldest eligible age and the 30-second transactional-provider objective by tenant/channel, unknown outcomes, bounce rates, suppression decisions and retry amplification. Alert through an independent channel so this system’s outage cannot suppress its own incident notification. Canary template and adapter changes on controlled recipients; test opt-out races, duplicated callbacks, DST scheduling and provider recovery before expanding.
16Decision ledger and limitations
The design keeps acceptance and consent decisions in our database while delegating email, SMS and push transport to providers. These tradeoffs follow from that split in control and from the need to protect urgent messages during bursts.
| Choice | Benefit | Cost |
|---|---|---|
| Durable intent plus delivery rows | Traceable retries and partial results | More state than one fire-and-forget call |
| Separate priority queues | Transactional latency protection | Fairness and reservation policy |
| Recheck preferences before send | Honors recent opt-outs | Additional policy lookup |
| Provider idempotency/reconciliation | Reduces duplicate sends | Depends on documented external behavior |
| Additional decision | Consequence | Reconsider when |
|---|---|---|
| Recipient-owned policy and delivery state | Local authorization transaction; provider queues need routing metadata | A user’s state becomes too hot or cross-region requirements change |
| Stable external delivery identity | Retries can use a provider’s dedupe contract | Provider retention/capability cannot meet the retry horizon |
| Unknown outcome as first-class state | Honest visibility and safe recovery | Product explicitly chooses duplicate risk over a missed message |
| Authorize before network call | Opt-out order is testable | Harder recall requirements need a provider-supported cancellation protocol |
| Distinct acceptance/delivery/read facts | Avoids false success claims | A channel adds a verified new receipt type |
The next bottleneck is likely a provider quota or one merchant’s campaign fairness, not generic API CPU. Adding workers should follow that measurement. A delayed marketing message may be the correct outcome when the alternative is missing transactional deadlines or bypassing current preferences.
17Interview closing
“I distinguish the business intent, each channel delivery and each provider attempt. The API durably accepts the notification intent, and a planner records stable delivery IDs with an outbox. A quota-aware dispatcher reserves transactional capacity and schedules work by provider and tenant.
“Before sending, the recipient owner serializes current policy with a dispatch authorization. In-app insertion is unique by delivery ID. Email is an external effect: a lost response becomes unknown, and I reuse a supported provider idempotency key or reconcile the original reference. I never fail over an unknown attempt blindly and call that exactly once. Verified callbacks update channel-specific facts without regressing known delivery.
“The assumed hundred million daily attempts drive quota, retention and retry costs. During a provider outage, recovery capacity must exceed new arrivals. I will measure eligible queue age, unknown-outcome age and per-channel expense. Some deliveries remain pending or unknown until evidence arrives; reporting success earlier would mislead the caller.”
Interviewer: “Marketing must stop the instant a user opts out.” Candidate: “I can stop authorizations ordered after the opt-out. For already authorized or accepted external sends, I need provider-supported cancellation and a defined acknowledgment boundary. Without that capability, I cannot promise recall merely by deleting a queued row.”
Practise the interview questions
Say your answer aloud before opening the model answer. Then answer the follow-up and compare the reasoning.
Why do you need three IDs for one notification?
Reveal a model answer
The order event identifies the logical intent, the channel delivery identifies where it goes, and the attempt identifies a provider call. Retries should add attempts without inventing new intents or channels.
Interviewer follow-up
Which ID belongs in a provider idempotency key?
Reveal the follow-up answer
The stable delivery ID, not a fresh attempt ID, assuming the provider’s documented contract supports that usage.
What the answer must demonstrate: Retry identity must survive retries.
The email provider times out after accepting your request. What do you do?
Reveal a model answer
I treat the result as unknown. I retry the same idempotency key or query a stable provider reference if supported. Without either capability, I apply the category’s explicit duplicate-versus-miss policy.
Interviewer follow-up
Would switching providers solve it?
Reveal the follow-up answer
No. Another provider cannot know the first one sent the email, so failover may create a duplicate.
What the answer must demonstrate: Do not infer rejection from lack of response.
A process dies after inserting a delivery but before queueing it. How is the message sent?
Reveal a model answer
The delivery and its outbox record were committed together. A dispatcher resumes reading unsent outbox records and publishes the stable delivery ID. Repeated queue messages are safe because workers use that identity.
Interviewer follow-up
Could the dispatcher publish twice?
Reveal the follow-up answer
Yes. Its publish acknowledgment can also be lost. The design tolerates repeated work instead of assuming a perfect handoff.
What the answer must demonstrate: Explain both sides of the handoff gap.
A recipient opts out after a campaign is queued. Do you still send?
Reveal a model answer
I recheck the current applicable preference near sending and suppress work that is no longer permitted. The queued snapshot helps explain planning, but it does not grant permanent consent.
Interviewer follow-up
What if the provider already accepted it?
Reveal the follow-up answer
Then cancellation may be impossible. The UI and audit trail must represent that boundary honestly.
What the answer must demonstrate: Explain whether the opt-out committed before or after send authorization.
A provider outage lasts ten minutes. Why might recovery take much longer?
Reveal a model answer
New arrivals continue while the backlog drains. Drain rate is provider capacity minus new traffic, not total capacity. If there is no spare quota, the queue never catches up.
Interviewer follow-up
What would you do before adding workers?
Reveal the follow-up answer
Check rate limits, prioritize urgent work, reduce or defer marketing, and bound retry traffic. More workers can just increase throttling.
What the answer must demonstrate: Calculate net drain rate.
A delivered callback arrives before a sent callback. What happens?
Reveal a model answer
I persist both facts and apply a channel-specific state rule that does not regress delivered to sent. I verify the signature and bind the message to its provider account and delivery. If there is no provider event ID, repeated facts still apply idempotently using the documented message/status identity.
Interviewer follow-up
Does every channel have a read receipt?
Reveal the follow-up answer
No. I expose only the states that channel and provider can actually establish. Accepted, delivered, and read are not interchangeable.
What the answer must demonstrate: Avoid one generic success boolean.
What exact guarantee can you make when a recipient opts out during dispatch?
Reveal a model answer
“The preference update and send authorization serialize at the recipient owner. If opt-out commits first, authorization is suppressed. If authorization already committed, work may be in flight; cancellation is best effort unless the provider offers a stronger protocol.”
Interviewer follow-up
Why not just check preferences again?
Reveal the follow-up answer
“Another check narrows the window but cannot atomically include an independent provider effect. I must define the real boundary rather than promise an impossible instant recall.”
What the answer must demonstrate: Identify the transaction that grants permission to send, then explain why opt-out cannot always recall the provider call that follows.
A campaign fills the queue while order updates miss their deadline. How do you change scheduling?
Reveal a model answer
“I reserve provider capacity for transactional work and use bounded weighted scheduling across tenants and categories. At a 1,000/s provider quota, an assumed 600/s reserve protects urgent work; unused capacity can be borrowed without erasing the reserve when urgent traffic returns.”
Interviewer follow-up
Would strict priority be simpler?
Reveal the follow-up answer
“Yes, but sustained urgent work can starve marketing. I monitor age and define a lower-priority service policy rather than assuming one FIFO queue provides every deadline.”
What the answer must demonstrate: Protect urgency without pretending to increase provider quota.
Blank-page exercise · 45 minutes
Build the answer yourself
Design order and marketing notifications, then handle an email-provider timeout after the provider may have accepted recipient U7’s message.
- Clarify intent, acceptance, delivery and read semantics.
- Calculate attempts, per-provider quotas and net backlog drain.
- Draw recipient ownership separately from provider scheduling.
- Trace opt-out versus send authorization, show why an in-app retry cannot insert a duplicate, and recover a provider call whose outcome is unknown.
- Compare fairness, provider failover and per-channel cost; give a natural closing answer.
Check that each component and design decision follows from your requirements and workload.
Recall the key ideas
Answer from memory before opening each card. Explain why the choice works and what it costs. Revisit missed cards tomorrow.
Design a notification serviceWhat is the difference between intent, delivery, and attempt?Recall first, then reveal
An intent is the business message, a delivery targets one channel/destination, and an attempt is one provider call.
Business message → recipient/channel delivery → provider call.
Return to lessonDesign a notification serviceA provider call times out. Is it safe to change providers?Recall first, then reveal
Not automatically. The first provider may already have accepted the message, so cross-provider retry can duplicate it.
Unknown is not failed.
Return to lessonDesign a notification serviceWhen must opt-out be checked?Recall first, then reveal
At planning and again in the send-authorization transaction. An opt-out committed before authorization suppresses the send; cancellation after authorization is best effort.
Queueing does not freeze consent.
Return to lessonFinal revision
Summary and interview notes
A notification service durably records intent before independently executing each channel delivery. The recipient database orders opt-outs and send authorizations. Stable delivery IDs let workers recover provider results; an unresolved timeout stays unknown.
Remember these points
- Keep business intent, channel delivery and transport attempt identities distinct.
- Check preferences and authorize sending in one ordered transaction; a provider call authorized earlier may be impossible to recall.
- Freeze recipient and payload before the first external attempt, and retain them across retries of that delivery.
- Provider acceptance, channel delivery and user reading are different facts.
- The backlog shrinks only when workers can complete more deliveries than continue to arrive.
Interview tips
- Start by asking which observable outcome the reliability target measures.
- Trace an opt-out race and a provider timeout after acceptance; a queue does not solve either by itself.
- Estimate traffic and expense per channel attempt, including retry amplification.
Important qualifications
- A provider without idempotency or reliable status lookup requires an explicit duplicate-versus-miss policy.
- Providers identify callbacks and report delivery differently; some do not supply unique event IDs.
Technical references
- Twilio outbound status trackingOfficial example of channel-specific delivery callbacks and signature validation.
- Transactional outbox patternExplains atomic durable state and later event dispatch.
- Amazon SQS at-least-once deliveryDocuments duplicate delivery as a consumer design concern.
- FCM registration managementOfficial guidance on registration freshness, refreshing legacy tokens where used, and removing invalid or stale destinations. Provider-specific registration lifecycles belong in the push adapter.
Practice marks stay in this browser.