System-design interview · Extended interviews
Design a notification service
Accept one notification intent, plan eligible channel deliveries and recover provider outcomes while respecting current preferences and protecting urgent traffic.
You will learn to
- Distinguish a business intent, a channel delivery and a provider attempt.
- Explain opt-out ordering and unknown external outcomes without claiming universal exactly-once delivery.
- Scale due work, provider quotas and backlog recovery with visible per-channel status.
Practice in this chapter
8 interview questions with model answers and follow-ups.
Go to interview practiceUseful foundations: Message queues, event logs, delivery guarantees, and backpressure · Databases, data models, and ACID transactions · Production readiness: SLI, SLO, observability, and recovery
Workload and timing examples are interview assumptions.
Dotted concept links open the relevant explanation in a new tab.
01Define what reliable notification means
Build a shared service for email, push, SMS and an in-app inbox. Trusted product services submit approved templates and business event identities. Audience selection for bulk campaigns, operating an email network and mobile push transport itself are outside scope. A campaign service can expand an audience into individual recipient intents subject to admission limits.
An order service submits event ship-o81 for user U7. The resulting intent N44 plans email delivery D-email and in-app delivery D-app; SMS may be suppressed by preference. The intent is the business request, a delivery is one channel/destination outcome, and an attempt is one provider call. Keeping those identities distinct makes partial success and retries understandable.
A successful API response means the intent is durable. Provider acceptance means a downstream service accepted responsibility; it does not prove device delivery or human reading. Expose those facts separately rather than returning one misleading delivered flag.
Target provider acceptance within thirty seconds for 99% of eligible, immediately due transactional deliveries, measured from durable intent acceptance; marketing can wait. State the denominator: eligible, immediately due messages. Suppressed or scheduled work has a different outcome, and unknown or failed provider calls count as missed delivery targets rather than disappearing from the metric.
Ask which channels and urgency classes matter, what “delivered” should mean, and how each category should handle an uncertain provider result. This design accepts durable intents, exposes channel-specific outcomes and prioritizes eligible transactional messages; it does not promise that a human reads them.
02Functional requirements
Submit notification intents. Accept an authenticated business event, recipient and approved template, returning one durable intent for matching retries.
Plan and schedule deliveries. Choose email, SMS, push or in-app delivery using recipient preferences, verified destinations, due time and quiet-hour policy.
Send and expose outcomes. Create an in-app item or invoke the external provider, then report suppression, pending, accepted, delivered, failed or unknown separately per channel.
Manage and recover delivery. Support preference changes, safe retries, verified callbacks, status lookup and inspection of exhausted or uncertain work.
03Non-functional requirements
These are illustrative interview assumptions, not product facts or measured benchmarks. Confirm them before choosing components, then validate the completed design under the stated workload. Here p95 means the 95th-percentile latency: 95% of measured requests take no longer than that value. Report errors and rejected work alongside latency; a fast failure is not a successful outcome.
Workload and quotas. Plan for 100 million external provider attempts/day and about 11,600 attempts/s at a tenfold burst. Retries consume the same provider quotas; these rates are not assumed provider entitlements.
Urgent-message timeliness. Target provider acceptance within thirty seconds for 99% of eligible, immediately due transactional deliveries. Measure from durable intent acceptance; unknown or failed provider calls count as misses. This is a service objective, not a guarantee of provider uptime or device delivery.
Durability and recovery. Accepted intents, delivery state and incoming provider evidence survive worker/API restarts through the transactional database. Preserve pending work through downstream outages, bound admission before storage fills and verify the database’s deployment failure policy separately.
Consent and privacy. A send authorization after a committed opt-out must refuse the disallowed delivery. A previously authorized external call may already be in flight; stop new authorization if current permission cannot be checked.
Duplicate and outcome semantics. Enforce one local inbox item per delivery. External retry safety depends on provider idempotency/lookup scope; preserve unknown outcomes and explicitly choose duplicate-versus-missing risk when neither is available.
04Complete the asynchronous flow with a database worker
Start with an authenticated API, a relational database and a worker polling indexed due-delivery rows. The order service reliably publishes ship-o81 from its own transaction using an outbox or equivalent retry mechanism. Our notification API cannot atomically commit with the separate order database, so repeated submission must be safe.
The API stores N44 and a planning task. The planner loads the approved template version and recipient preferences, creates stable channel delivery rows and records suppression reasons. D-app is completed by inserting a unique inbox item and updating its delivery record in one local transaction after checking current permission.
For D-email, the worker transaction checks current consent and destination, records authorization to send and releases its locks. It calls the provider using the delivery's stable external key where supported. It records provider acceptance or an unknown outcome, then later applies verified receipts. U7 can continue using the order page while this runs.
A worker restart finds the same durable rows. No broker is required to make this first version asynchronous or recoverable. Even one worker can crash after the provider sends a message but before the worker saves the reply; the local database cannot roll back that send.
The API acknowledges stored intent; channel workers later establish delivery facts.
Read each connection in order
- syncship-o81, recipient, templateOrder service → Notification API
- syncStore intent and planNotification API → Intent, preferences, deliveries
- syncAuthorize current due deliveryDue-delivery worker → Intent, preferences, deliveries
- syncStable delivery identityDue-delivery worker → Email / push / SMS provider
- syncUnique local itemDue-delivery worker → In-app inbox
05Count intents, deliveries and attempts separately
Assume 100 million provider attempts/day, including retries. That averages about 1,157 attempts/s; a tenfold burst is approximately 11,600/s. If the mix averages 1.1 attempts per external delivery and two external deliveries per intent, it represents about 90.9 million deliveries and 45.5 million intents/day. In-app insertions add database work without provider calls.
At 1 KB of attempt metadata, thirty days require roughly 3 TB before indexes and replicas. Rendered message bodies and destinations may be more sensitive than diagnostic metadata; retain only what is needed rather than multiplying full bodies across every attempt.
A 200 ms average provider call and 11,600 attempts/s imply about 2,320 calls in flight if quotas permit. This is a concurrency estimate, not permission to exceed a provider's contracted rate. A single sequential worker would manage only about five calls/s at that latency.
If a provider receiving 500 attempts/s fails for ten minutes, 300,000 attempts accumulate. After recovery, 1,000/s capacity minus 500/s new traffic leaves 500/s to drain backlog, requiring another ten minutes. Recovery capacity must exceed new arrivals; simply restoring normal throughput leaves the queue permanently behind.
06Persist the plan and the evidence
Interfaces
| Request or message | Contract |
|---|---|
POST /notifications with event, recipient, category and template version |
Returns accepted N44; duplicate matching events return the same intent. |
GET /notifications/N44 |
Reports queued, suppressed, accepted, delivered, failed or unknown per channel. |
PUT /users/me/preferences |
Changes versioned channel/category policy. |
Submit a notification intent
POST /notifications
Notification request values
event: ship-o81
recipient: U7
category: transactional order update
template version: the approved immutable version selected by the caller
Separate channel outcomes for N44
email / D-email: pending at the provider
in-app / D-app: committed locally
SMS: suppressed by preference
Stored records
| Record | Fields or identity | Purpose |
|---|---|---|
| Intent | tenant, event, user, category, payloadHash |
Stable business meaning and deduplication scope. |
| Delivery | intent, channel, destinationVersion, state, dueAt |
One logical recipient endpoint across attempts. |
| Attempt | delivery, number, providerKey, outcome |
Diagnostic record of each external invocation. |
| In-app item unique by delivery ID | Delivery ID (unique) | A locally enforceable no-duplicate inbox insertion. |
Templates are immutable versions with typed parameters. Callers may select authorized templates, not inject arbitrary destinations or executable markup. Reusing a business identity with different canonical parameters conflicts. Obtain verified email, phone and device registrations from trusted recipient records.
Before the first authorized send, freeze the exact destination and rendered parameters for that delivery. Changing an email address during recovery must not send different content to a new destination under the same provider key. A replacement is a deliberate new logical delivery with a policy for the older uncertain one.
Store incoming provider events in a durable inbox before acknowledging them; store outgoing dispatch work in an outbox with delivery changes. These records make both handoffs retryable after a crash.
07Define the exact opt-out boundary
Planning is not permanent permission. U7 may opt out after N44 is planned but before the worker sends it. Keep each recipient’s preferences and delivery records where one transaction can check and update them. The worker’s authorization transaction:
Locks the relevant preference and delivery records.
Checks current category/channel policy, due time and destination eligibility.
Either suppresses the delivery or records a send authorization.
An opt-out transaction locks the same preference record. If the opt-out commits first, the later authorization must refuse the send. If authorization commits first, an external call may already be in flight. Best-effort cancellation can help, but the service cannot promise to recall an email already accepted by a provider.
Release database locks before network calls. Holding them until delivery completes would block preference changes for an unpredictable time and still would not create a distributed transaction with the provider. Record the policy and destination versions used so operators can explain the decision.
For in-app items, authorization and unique item insertion can occur in one local transaction. That stronger local guarantee does not automatically extend to SMS, push or email. If the preference authority is unavailable, stop new authorizations rather than treating an old cached policy as current consent.
08Recover uncertainty without inventing another message
Suppose the email provider accepts D-email but its response is lost. The worker cannot tell whether nothing happened or whether the email is already on its way. Record unknown rather than definitive failure. If the provider supports idempotency, retry the same logical delivery key within its documented scope and retention. If it supports reliable lookup by client reference, query the original operation.
Attempt numbers identify individual calls for diagnosis. Using a new attempt number as a new provider deduplication key would allow every retry to create another message. A worker lease can prevent ordinary concurrent work but cannot stop a paused old process from later calling a provider that does not enforce our lease token.
When the provider has neither idempotency nor reliable lookup, an automatic exactly-once guarantee is unavailable. Choose a category-specific policy: a transactional update may tolerate a documented duplicate risk; marketing may prefer withholding another attempt until reviewed. Keep the uncertainty visible.
Switching providers does not solve this ambiguity because the second provider cannot deduplicate an effect at the first. Fail over after a known failure, or explicitly accept duplicate risk. Definitive invalid-address errors should stop or disable the matching destination version; temporary rate limits use bounded backoff and provider retry guidance.
The second attempt recovers D-email; it does not invent another logical message.
Read each connection in order
- syncAuthorize and record D-emailEmail worker → Delivery authority
- syncSend with stable delivery keyEmail worker → Provider
- blockedAcceptance reply lostProvider → Email worker
- syncRecord unknown outcomeEmail worker → Delivery authority
- syncLookup or same-key recoveryEmail worker → Provider
- syncApply verified original outcomeEmail worker → Delivery authority
09Schedule by provider capacity and recipient ownership
Add a broker when measurements show that buffering and separately scaled workers improve delivery of due work. Commit delivery rows and outbox records together; a relay publishes stable delivery IDs. Duplicate publication is acceptable because the worker rechecks the authoritative delivery state before acting. The broker schedules work and does not become the sole copy of consent or message meaning.
Group execution by provider/channel quota and urgency. Reserve capacity for transactional updates so a large campaign cannot occupy every provider slot. Use weighted scheduling and per-tenant limits rather than unlimited strict priority, which may starve lower-priority work. Unused reserved capacity can be borrowed under an explicit rule without removing the transactional reserve during a surge.
Recipient data partitions serve a different purpose. Keep a recipient's preferences, delivery decisions and inbox together so authorization remains local. Queue messages carry enough routing information to reach that owner. Before a new database owner authorizes sends, it must have the current preferences and complete state of deliveries already in flight.
A queue cannot guarantee a thirty-second deadline when backlog exceeds downstream capacity. Reject impossible new work before promising acceptance, defer marketing, and expose oldest eligible age. Adding workers helps only until provider quotas, reputation rules or account-specific limits become the bottleneck.
A product submits one stable notification intent. Planning and outbox work create per-channel deliveries; the broker carries delivery IDs to quota-controlled workers. Each worker reloads current policy and delivery state before a provider call. In-app items stay in the recipient’s database, and verified receipts update the same delivery history.
Read each connection in order
- syncStable intent IDProduct services → Notification API
- syncPersist intent and planNotification API → Recipient policy / delivery DB
- syncPlan and read outboxPlanner and outbox relay → Recipient policy / delivery DB
- asyncPublish delivery IDsPlanner and outbox relay → Channel-ready queues
- asyncSchedule delivery attemptsChannel-ready queues → Channel worker pools
- syncAuthorize / record outcomeChannel worker pools → Recipient policy / delivery DB
- syncSend or recover outcomeChannel worker pools → Email / push / SMS providers
- asyncVerified delivery receiptsEmail / push / SMS providers → Notification API
10Handle schedules, devices and receipts deliberately
Quiet hours depend on the recipient's named time zone, not the worker's local clock. Store the resolved scheduled instant and the rule version. Define what happens to a local reminder time skipped or repeated by daylight-saving changes, and what a subsequent user time-zone change affects. Marketing cannot bypass opt-out merely because it entered an urgent queue.
Push delivery targets app installations through provider registration tokens. A user may have several devices, and tokens can rotate. Bind registrations to authenticated users, version each destination and disable only the version associated with a definitive invalid-token response. A delayed rejection for an old token must not disable the replacement.
Verified callbacks can repeat or arrive out of order. Deduplicate provider event IDs where supplied and apply channel-specific facts without moving a delivered message backward to queued. Match the provider account and message reference to the intended delivery. An email bounce and an SMS device receipt have different meanings and should not be flattened into one generic success flag.
Inbox reads authorize the recipient and paginate by creation time plus stable item ID. A cursor orders results but is not automatically a frozen snapshot. Notification previews should not reveal content the user is no longer authorized to read through the underlying application.
11Operate recovery and consent as first-class behavior
Monitor intent admission, oldest eligible delivery age, unknown-outcome age, suppression reasons, provider rate-limit responses, bounce rates and retry amplification. Separate channels, tenants and urgency lanes so a healthy aggregate does not conceal an urgent-message backlog. Alert through an independent channel; this service cannot reliably report its own outage through itself.
Dead-letter handling means recording exhausted or invalid work for inspection, not deleting its history. Replaying a delivery preserves its logical identity and frozen parameters. A deployment rollback must not turn already accepted emails into new sends with new IDs merely because local status appears incomplete.
Test opt-out before and after authorization, provider acceptance followed by worker crash, repeated callbacks, destination rotation, daylight-saving scheduling and a ten-minute provider outage. Validate actual drain capacity, not just whether workers restart. Recovery queries also consume provider quota and need a budget.
The principal cost is attempts per channel under provider pricing, plus retained metadata and recovery work. Reduce wasteful retries only when the desired delivery outcome remains supported. Keep personal destinations and rendered bodies out of metric labels, broadly accessible logs and queue names; restrict template editing separately from sending permission.
12Check the design against its requirements
Use the numbered requirements to check the final design. FR refers to the functional list; NFR refers to the non-functional list. Performance rows specify tests still required, not achieved benchmark results.
| Requirement | Design mechanism | Verification and remaining limit |
|---|---|---|
| FR1–2; NFR3: accept and plan once | Durable intent identity, versioned templates, planning rows and outgoing work. | Repeat ship-o81, restart the planner and change a destination during recovery. Recover the same logical intent and frozen delivery parameters. |
| FR2,4; NFR4: honor current consent | Recipient-owned preference and delivery checks serialize opt-out against send authorization. | Race opt-out with authorization and make the preference authority unavailable. Refuse later unauthorized sends; do not claim already-sent messages can be recalled. |
| FR3–4; NFR5: report honest outcomes | Stable delivery identity, provider-supported recovery, verified callbacks and unique in-app insertion. | Lose a provider response and replay callbacks. Distinguish accepted from read, and never silently turn uncertainty into a new logical send. |
| NFR1–3: meet urgency under load | Reserved transactional capacity, provider quota limits, bounded admission and separate recovery budget. | Measure thirty-second success rate during a campaign and a ten-minute provider outage. The outage creates misses; validate net backlog-drain capacity instead of concealing them. |
13Rapid revision
Remember: Saved intent, provider acceptance and human reading are different outcomes. Recover the same delivery when a provider reply is lost.
| Concern | Complete mechanism |
|---|---|
| Admission | Check the sender and permitted template; save one notification per matching business request. |
| Planning | Create a delivery ID for each eligible channel; record why other channels were skipped. |
| Send permission | In a short transaction, recheck current recipient preferences and the allowed destination. |
| In-app completion | Atomically insert one item per delivery and mark it complete. |
| External completion | Reuse the provider key or look up the delivery; keep its status unknown until evidence resolves it. |
| Retries | Delay retries after temporary failures, stop invalid destinations, and keep records that recognize repeated attempts. |
| Scheduling | Honor named time zones, quiet hours and provider quotas; reserve urgent capacity. |
| Callbacks | Verify, persist, deduplicate and apply each channel’s status rules so an old callback cannot overwrite a later known outcome. |
| Recovery | Process faster than new work arrives; inspect old unresolved deliveries as well as queue length. |
Close with ship-o81: one accepted intent, email pending at the provider, in-app committed and SMS suppressed. That single example demonstrates the data model, asynchronous boundary, consent policy and honest partial status without requiring a universal delivered boolean.
Practise the interview questions
Say your answer aloud before opening the model answer. Then answer the follow-up and compare the reasoning.
Why separate intent, delivery and attempt?
Reveal a model answer
One business event can create several channel deliveries, and each delivery may require several calls. Separate identities preserve partial status and safe retries.
Interviewer follow-up
Which identity should a provider key normally represent?
Reveal the follow-up answer
The logical delivery, not each transport attempt.
What the answer must demonstrate: Keep business intent, channel delivery and transport attempt identities separate.
Does provider acceptance prove the user read the message?
Reveal a model answer
No. Admission, provider acceptance, device delivery and human reading are distinct facts with channel-specific evidence.
Interviewer follow-up
What should status return?
Reveal the follow-up answer
The known state and update time for each channel, including suppressed and unknown.
What the answer must demonstrate: Distinguish service acceptance, provider acceptance, device delivery and human reading.
The user opts out after planning. What prevents a later marketing send?
Reveal a model answer
The worker checks current preference in the same transaction that authorizes dispatch. An opt-out committed first suppresses the delivery.
Interviewer follow-up
Can it recall an already authorized email?
Reveal the follow-up answer
Not universally; that depends on provider cancellation and the declared in-flight boundary.
What the answer must demonstrate: Define the transaction that orders consent changes and send authorization.
The email call times out after possible acceptance. How do you retry?
Reveal a model answer
Recover the same delivery through supported idempotency or lookup. Without either, keep unknown and apply the category’s explicit duplicate-risk policy.
Interviewer follow-up
Can switching providers fix it?
Reveal the follow-up answer
No. A second provider cannot deduplicate an effect at the first.
What the answer must demonstrate: Use original provider identity or lookup and state the unsupported-provider limit.
Why can in-app notifications have a stronger uniqueness boundary?
Reveal a model answer
The item insertion and completed delivery state can share one local transaction with a unique delivery ID.
Interviewer follow-up
Does that guarantee email uniqueness too?
Reveal the follow-up answer
No. Email is an external effect outside that transaction.
What the answer must demonstrate: Identify the local transaction boundary that makes inbox insertion unique.
A campaign fills the queue. How do urgent order updates meet their target?
Reveal a model answer
Reserve provider capacity for transactional work and apply tenant fairness. More workers cannot exceed the same provider quota.
Interviewer follow-up
What if arrivals equal restored capacity?
Reveal the follow-up answer
The old backlog never drains; recovery needs spare capacity or reduced new work.
What the answer must demonstrate: Reserve urgency within actual provider quotas and calculate net backlog drain.
A rejection arrives for an old push token after refresh. What changes?
Reveal a model answer
Disable only the matching destination version, preserving the newer registration.
Interviewer follow-up
Can a retry silently use the new token under the old key?
Reveal the follow-up answer
No. Freeze the uncertain delivery destination; replacing it is a deliberate new delivery.
What the answer must demonstrate: Guard device invalidation by destination version and freeze uncertain send parameters.
How should duplicate or reordered callbacks behave?
Reveal a model answer
Verify and persist them, deduplicate available event identities and apply channel-specific facts without regressing known outcomes.
Interviewer follow-up
Why acknowledge after persistence?
Reveal the follow-up answer
A process crash must not lose the only evidence after telling the provider it was received.
What the answer must demonstrate: Persist verified provider facts and apply channel-specific idempotent transitions.
Blank-page exercise · 45 minutes
Build the answer yourself
Design shipment notifications over email and in-app, then race an opt-out with dispatch and lose the email provider’s response.
- Agree numbered functional and non-functional requirements, including channel outcomes, urgency and opt-out/duplicate behavior. Then distinguish a business intent, a channel delivery and a provider attempt.
- Explain opt-out ordering and unknown external outcomes without claiming universal exactly-once delivery.
- Scale due work, provider quotas and backlog recovery with visible per-channel status.
- Trace a timeout and a concurrent request using the actual durable records.
- Review the final architecture against every numbered requirement, including measured bottlenecks, targets still needing validation and remaining failure limits.
Check that each component and design decision follows from your requirements and workload.
Recall the key ideas
Answer from memory before opening each card. Explain why the choice works and what it costs. Revisit missed cards tomorrow.
Design a notification serviceN44 has a saved in-app item, but its email call timed out. What can status honestly report?Recall first, then reveal
The intent is durable, the in-app item is complete and email is unknown until provider evidence resolves it. Neither proves the user read the message.
Saved intent, provider acceptance and user reading need different evidence.
Return to lessonDesign a notification serviceDoes a queued notification still need permission before sending?Recall first, then reveal
Yes. The queue schedules work; the database checks current consent. The provider’s retry rules govern repeated external calls.
Queue work; authorize sends.
Return to lessonDesign a notification serviceWhen does a delivery backlog shrink?Recall first, then reveal
Only when completion capacity exceeds the rate of new deliveries arriving.
New work uses capacity too; only the spare part drains backlog.
Return to lessonFinal revision
Summary and interview notes
Save each requested notification once, create eligible channel deliveries, and recover uncertain provider results. Check current preferences before sending and keep capacity available for urgent messages.
Remember these points
- Separate business, channel and attempt identities.
- Check current consent in the transaction that authorizes sending.
- Keep a timed-out provider call unknown until evidence resolves it.
- Protect urgent traffic with actual provider capacity.
Interview tips
- Name the outcome measured by the delivery target.
- Trace ship-o81 through email, in-app and SMS, then race an opt-out with send authorization.
Important qualifications
- Traffic and latency figures are interview assumptions, not claims about a named company's deployment.
Continue after the core interview
Explore the advanced version
The advanced lesson keeps the full detailed design. Use these sections when you want to examine the stronger requirements and failure cases.
- Detailed send-authorization proof
Study worker pauses, takeover and external uncertainty beyond the core transaction.
- Provider recovery and failover
Category-specific duplicate risk depends on the exact provider capabilities.
- Device registration and immutable destinations
Large multi-device systems need fuller registration rotation and account-binding rules.
- Quiet hours and operational cost
Explore calendar transitions and provider economics after the complete main flow.
Technical references
- Twilio outbound status trackingOfficial example of channel-specific delivery callbacks and signature validation.
- Transactional outbox patternExplains atomic durable state and later event dispatch.
- Amazon SQS at-least-once deliveryDocuments duplicate delivery as a consumer design concern.
- FCM registration managementOfficial guidance on registration freshness, refreshing legacy tokens where used, and removing invalid or stale destinations. Provider-specific registration lifecycles belong in the push adapter.
Practice marks stay in this browser.