System designby Learnastra

System-design interview · Core interviews

Design a chat messaging service

By Anup Rai

Design durable chat by separating accepted messages, live delivery, per-device history recovery and advisory presence.

You will learn to

  • Trace one message through acceptance, delivery and offline recovery.
  • Explain stable send identities, conversation order and device cursors.
  • Scale sockets and storage without weakening durable acknowledgment.

Practice in this chapter

8 interview questions with model answers and follow-ups.

Go to interview practice

Useful foundations: Real-time communication: polling, long polling, SSE, and WebSocket · Message queues, event logs, delivery guarantees, and backpressure · Replication and durability

Workload and timing examples are interview assumptions.

01Define what the sender’s status actually promises

Nora sends “Train arrives at six” to Sam. Sam's phone is connected, but Sam's laptop is asleep. The service should save one message, deliver it promptly where possible and let the laptop recover the same history later. A network connection and a durable conversation record solve different parts of that problem.

Use three distinct statuses:

Status Meaning
Accepted The server committed the message under its storage durability policy.
Delivered A particular device acknowledged receiving it.
Read A client reported a read action; this is not proof of human attention.

An offline recipient does not make acceptance fail.

Support text, one-to-one conversations, bounded groups, multiple devices, history, receipts and advisory presence. Exclude attachments, editing and end-to-end key-management design from this exercise. Group membership must govern sending and reading, and the history available to new members needs an explicit policy. Here, members see messages from their join point while they remain eligible. Global order across unrelated conversations is unnecessary.

The sender may immediately see a pending local bubble. Only a committed server response can change it to accepted. Establishing these meanings early prevents a fast socket response from being mistaken for saved history.

Clarify whether success means saved, delivered to one device or delivered to every device, and agree on group size and online-delivery delay. The requirements below keep acceptance independent of whether the recipient is currently connected.

02Functional requirements

Agree on these supported actions before selecting components.

  1. Send and retain messages. Authenticated members send text in one-to-one conversations or groups, receive a stable accepted-message identity and read ordered history.

  2. Support multiple and offline devices. Deliver to connected eligible devices and allow other devices to retrieve missed history after reconnecting.

  3. Report distinct progress. Expose pending, accepted, device-delivered and client-reported read states, with advisory presence kept separate from those guarantees.

  4. Apply membership rules. Check send/read eligibility, including after a socket has connected. New members see history from their join point while they remain eligible.

03Non-functional requirements

Use these as illustrative interview assumptions to agree with the interviewer. Numerical targets require measurement; they are not claims about an existing product or a proven implementation. p95 (the 95th percentile) means 95% of measured operations finish within the stated time.

  1. Workload and group bound. Plan for approximately 1.16 million accepted messages/s at peak and 60 million simultaneous device connections. Choose a maximum of 100 members per group for this exercise; larger broadcasts require another fanout plan.

  2. Latency for reachable devices. Target regional p95 send-to-accepted latency of 200 ms and accepted-to-device delivery of one second for connected eligible devices under the admitted peak workload and tested healthy-network profile. Offline devices have no live-delivery deadline.

  3. Acknowledged-message durability. Accepted messages, their sequence and retry identity must survive a gateway restart or one database-node failure within the serving region. When that commit cannot be made safely, leave the send pending or return unavailable.

  4. Ordering and duplicate handling. Provide one committed order per conversation and recover repeated sends/deliveries using stable identities. Global order across conversations and exactly-once network transmission are not required.

  5. History and device progress. Use five years of retained message history for the stated storage estimate. Each device tracks its own contiguous delivered/read progress and receives an explicit reset boundary when its cursor predates retained history.

  6. Confidentiality and abuse control. Authorize each send and history/live delivery, protect connections and bound message size, send rate and recipient work. End-to-end key management remains outside this exercise.

04Save one message before trying to deliver it

Begin with one application, a relational database and a local map of authenticated device connections. Nora sends message identity send-71 for conversation c8. The application checks membership and commits message m901 with the next conversation sequence, 1042. It also records the delivery work in that transaction. Only after commit does it return accepted.

A dispatcher sends m901 to Sam's connected phone. The phone stores the message and reports delivery. The sleeping laptop receives nothing yet; when it reconnects, it requests messages after its last saved sequence and obtains m901 from history. A push notification can wake a mobile app, but the push provider is not the conversation database.

Use (conversationId, senderId, clientMessageId) as a unique request identity. Retrying send-71 with the same text returns m901 and sequence 1042. Changing the text under that identity is a conflict. Sending the same words intentionally twice under different IDs creates two messages.

WebSocket offers a persistent bidirectional channel; long polling is a workable simpler transport. Neither determines durability or deduplication. Those properties come from the database commit, message identity and client recovery behavior.

Design diagramSave once, then deliver or catch up

The application commits before acknowledging. Offline devices later read the same history.

Save once, then deliver or catch upThe application commits before acknowledging. Offline devices later read the same history. sender to app: Send stable message identity; app to db: Commit message and delivery task; app to sender: Accepted after commit; app to recipient: Deliver committed message; recipient to app: Fetch history / report progressSend stable message identityCommit message and deliverytaskAccepted aftercommitDeliver committed messageFetch history / report progressACTORSender deviceSERVICEChat applicationSTOREMessages,membership andoutboxACTORRecipient devicessyncasync
Read each connection in order
  1. syncSend stable message identitySender device → Chat application
  2. syncCommit message and delivery taskChat application → Messages, membership and outbox
  3. syncAccepted after commitChat application → Sender device
  4. asyncDeliver committed messageChat application → Recipient devices
  5. syncFetch history / report progressRecipient devices → Chat application

05Size stored messages and connected devices independently

Assume 500 million daily active users sending forty messages each: 20 billion messages/day, or about 231,481 messages/s. A fivefold peak is approximately 1.16 million/s. At 100 bytes of body text, that is 2 TB/day. A 300-byte stored envelope including identifiers and metadata is 6 TB/day, or 10.95 PB over five years before replicas and indexes.

Now estimate connections separately. If 10% of daily users are simultaneously online with 1.2 devices each, there are 60 million sockets. A benchmarked capacity of 20,000 connections per gateway implies 3,000 gateway-equivalents before spare capacity. That capacity figure is an assumption to test with realistic encryption, buffers and message traffic, not a universal server limit.

Additional work Example Implication
Heartbeats 60M / 30 seconds = 2M/s Idle connections still create work
Group delivery 100 members × two devices One stored message can require 200 live deliveries
Socket buffers 60M × 32 KB = 1.92 TB Small per-connection allocations become large fleet totals

Storage scales with committed envelopes and retention. Gateways scale with sockets and fanout. Presence and receipts add their own traffic. One average messages-per-second number cannot size all three.

06Give sends, history and receipts different contracts

A send identifies its conversation, the client's message operation and the text. The authenticated connection identifies its user and device; the client cannot impersonate another sender by changing a JSON field. An HTTP send endpoint can use the same identity and semantics as the socket operation.

send event payload

{
  "conversationId": "c8",
  "clientMessageId": "send-71",
  "body": "Train arrives at six"
}

accepted event payload, after durable commit

{
  "messageId": "m901",
  "sequence": 1042
}

delivered event payload from the device

{
  "through": 1042
}

A separate read event uses the same through field for client-reported read progress. These progress events apply to the relevant conversation and authenticated device.

Event or endpoint Meaning
send Request durable acceptance
accepted Saved message identity and order
GET /conversations/c8/messages?after=1040 Authorized ordered catch-up page
delivered This device applied the contiguous permitted history
read Separate client-reported read progress

A timeout does not establish whether a send committed. Keep the pending client operation and retry its original ID. Bound message size, send rate, group size and history-page length so one client cannot request unlimited work.

A conversation sequence orders committed history. Wall-clock timestamps remain useful for display but do not reliably order concurrent sends from different devices. If retained history no longer reaches an old cursor, return an explicit earliest-available position and reset behavior rather than making the device chase an unfillable gap forever.

07Keep shared history and device progress separate

The design keeps durable conversation records, per-device progress and temporary connection state distinct.

Record Key or stored information Purpose
Message (conversationId, sequence); message ID, sender, clientMessageId, body Stores one message at its committed conversation position.
Conversation Next sequence Allocates the next position in that conversation.
Membership Who may send or read Determines current access.
Outbox Message identity and delivery work Retains dispatch work in the message transaction.
DeviceCursor Device/conversation and delivered/read progress Records each device's own applied history and read reports.
Session directory User/device to current gateway connection Locates live connections; entries expire when heartbeats stop.

In one transaction:

  1. Check membership.
  2. Claim the unique send identity.
  3. Allocate the sequence and insert the message.
  4. Save the outbox delivery task and commit before returning accepted.

This keeps one coherent conversation order.

An Outbox row is the durable delivery task committed alongside the message. If the process crashes before dispatch, another worker can still find it. A relay may send twice, so consumers deduplicate by message identity. A separate DeviceCursor tracks delivered/read progress for each device and conversation; the phone's receipt must not advance the laptop's cursor.

History lookup is a range query within one conversation, so partition by conversation rather than hashing every message ID independently. A user-to-conversation index supports the inbox list, but it is a derived summary rather than the authoritative message body.

A session directory maps user/device to its current gateway connection. Entries expire when heartbeats stop. Losing that directory can interrupt live delivery without deleting history. Likewise an online indicator estimates recent connectivity; it does not prove receipt or permission.

08Recover a missing message without duplicating the display

Sam's phone has applied c8 through sequence 1040, but receives a live event for 1042. It must not immediately claim delivery through 1042: message 1041 may be missing. Instead it asks the history service for the ordered page after 1040, authorizes the request, and applies the missing entries.

The phone saves the messages and updated contiguous cursor together in its local database. Only then does it acknowledge delivery. If a receipt disappears, the server may deliver m901 again. The phone recognizes its identity, keeps one displayed message and repeats the receipt. This is recoverable repeated delivery, not a claim that packets travel exactly once.

Stored progress only moves forward: a delayed receipt for 1041 cannot overwrite a newer one for 1042. Read progress remains separate from delivery progress. Receipt aggregation also needs a product definition: “delivered to one device” differs from “delivered everywhere.”

Membership is checked for live delivery and history, not only when a socket first connects. Removing a member must not leave an old connection as indefinite permission to receive new content. Already released bytes may still arrive; stronger in-flight revocation needs a separately defined protocol.

Request traceA lost receipt leads to safe repeated delivery

The device stores one identity, so an uncertain delivery acknowledgment does not create a second bubble.

A lost receipt leads to safe repeated deliveryThe device stores one identity, so an uncertain delivery acknowledgment does not create a second bubble. server to phone: Deliver m901 at sequence 1042; phone to phone: Persist message and contiguous cursor; phone to server: Delivery receipt lost; server to phone: Repeat m901; phone to phone: Recognize saved identity; phone to server: Repeat delivered-through 1042PARTICIPANTDelivery servicePARTICIPANTRecipient phone1. Deliver m901 at sequence 10422. Persist message and contiguous cursor3. Delivery receipt lost4. Repeat m9015. Recognize saved identity6. Repeat delivered-through 1042syncblockedreturn
Read each connection in order
  1. syncDeliver m901 at sequence 1042Delivery service → Recipient phone
  2. syncPersist message and contiguous cursorRecipient phone → Recipient phone
  3. blockedDelivery receipt lostRecipient phone → Delivery service
  4. syncRepeat m901Delivery service → Recipient phone
  5. syncRecognize saved identityRecipient phone → Recipient phone
  6. returnRepeat delivered-through 1042Recipient phone → Delivery service

09Separate socket capacity from conversation writes

Move long-lived connections to gateway processes once a single application cannot handle them. Gateways authenticate connections and forward sends to the storage group responsible for that conversation. They do not keep a database connection open for every idle socket; storage uses a bounded request pool.

For the worked scaled design, distribute independent conversations across replicated database groups. Keep message identity, sequence, membership and outbox updates together for each conversation. This provides parallelism across conversations without inventing a global ordering service. One extremely busy group still has an ordering bottleneck and needs a measured product-specific extension.

Dispatchers read committed outbox tasks, locate recipient devices through the session directory and deliver via their gateways. They can batch group recipients by gateway to reduce repeated inter-server transfer. Limit fanout tasks and queue age; a huge group must not starve ordinary conversations.

Cache recent immutable message pages only where reuse exists, and move older history to a cheaper storage tier under an explicit latency contract. Subscribe to presence for relevant contacts rather than broadcasting every heartbeat to every user. These changes reduce work while retaining the same meaning of accepted, delivered and read.

For the accepted-message guarantee, each conversation group uses three database replicas in independent regional failure domains and acknowledges only after a durable majority commit. Failover must preserve the message, sequence, retry result and outbox together. A minority cannot acknowledge new messages; test this policy against the 200 ms acceptance target rather than assuming replication is free.

Design diagramConnection gateways and conversation storage scale independently

A gateway routes each send to its conversation service, which commits message identity, sequence, membership checks and outbox work together. Dispatchers find live devices and deliver through gateways. Offline or reconnected devices use the same service to retrieve saved history; the session directory never replaces that history.

Connection gateways and conversation storage scale independentlyA gateway routes each send to its conversation service, which commits message identity, sequence, membership checks and outbox work together. Dispatchers find live devices and deliver through gateways. Offline or reconnected devices use the same service to retrieve saved history; the session directory never replaces that history. device to gateway: Send / reconnect / fetch history; gateway to app: Route by conversation; app to db: Commit message / read history; db to dispatch: Relay committed outbox work; dispatch to directory: Locate device gateways; dispatch to gateway: Deliver committed messages; gateway to device: Message or change notification; gateway to directory: Register and renew connectionsSend / reconnect / fetch historyRoute by conversationCommit message / read historyRelay committed outbox workLocate device gatewaysDeliver committed messagesMessage or change notificationRegister and renewconnectionsACTORSender and recipientdevicesSERVICEConnection gatewaysSERVICEConversationservicesSTOREReplicatedconversation groupsSERVICEDelivery dispatchersSTOREExpiring sessiondirectorysyncasync
Read each connection in order
  1. syncSend / reconnect / fetch historySender and recipient devices → Connection gateways
  2. syncRoute by conversationConnection gateways → Conversation services
  3. syncCommit message / read historyConversation services → Replicated conversation groups
  4. asyncRelay committed outbox workReplicated conversation groups → Delivery dispatchers
  5. syncLocate device gatewaysDelivery dispatchers → Expiring session directory
  6. asyncDeliver committed messagesDelivery dispatchers → Connection gateways
  7. asyncMessage or change notificationConnection gateways → Sender and recipient devices
  8. syncRegister and renew connectionsConnection gateways → Expiring session directory

10Test the crash on each side of acknowledgment

If the application crashes before the message transaction commits, Nora has no accepted result and retries send-71. If it crashes after commit but before replying, the retry finds the saved message and returns the same sequence. A concurrent retry cannot create another message because the unique send identity is part of that transaction.

A gateway crash breaks sockets, not conversation history. Clients reconnect with randomized backoff, reauthenticate, resend unresolved operations under their original identities and fetch history after their saved cursors. Do not try to infer delivery from a stale session-directory entry.

If the storage group cannot safely commit, leave the client message pending or return a retryable error. A gateway must not acknowledge from its own memory merely to keep latency low. Replication and failover need to preserve acknowledged commits; backups cover different risks such as accidental deletion.

Accepted messages create delivery work that must remain recoverable. Limit new sends before that work exceeds what the durable queue can retain. Reduce nonessential presence updates first, and reserve capacity for history catch-up. A push-provider outage should delay wake-ups, not make previously accepted messages disappear.

11Measure the boundaries users can observe

Measure send-to-accepted latency, accepted-to-device delivery delay and client-reported read progress separately. A fast database can coexist with a broken dispatcher. Monitor oldest outbox work, duplicate retries, reconnect rate, gap-fetch frequency and especially large-group queue age.

Authorize every send and history read, cap message and group sizes, and apply spam controls by account and conversation. Avoid sensitive text in logs and lock-screen push previews. Transport encryption protects connections; end-to-end encryption also changes who holds keys, whether servers can search bodies and how a new device obtains history. Treat that as a real extension.

The main costs are retained message copies, gateways and delivery to every recipient device. Three copies of the illustrative 10.95 PB five-year envelope store require 32.85 PB before indexes and backups. Caching each user's entire history would duplicate that expense in memory; retain bounded recent pages instead.

Test lost sender responses, lost recipient receipts, sequence gaps, membership removal with an old socket and a reconnect wave after gateway failure. These cases demonstrate durable user behavior more directly than showing a successful WebSocket handshake.

12Check the design against the requirements

Use the agreed lists to check the finished design. The tests below still need to establish the targets; a proposed mechanism is not a measured result. FR refers to the numbered functional requirements above; NFR refers to the numbered non-functional requirements.

Requirement Design mechanism Validation and remaining limit
FR1,3 + NFR2–4: accepted send Conversation-local membership, identity, sequence, message and outbox transaction. Lose the commit reply and retry; return one message. Benchmark p95 acceptance separately from device delivery.
FR2 + NFR5: offline devices Durable history, per-device cursors and a local recovery journal. Reconnect a sleeping laptop and inject a sequence gap. Do not advance its cursor past an unapplied message.
NFR3: one-node loss Durable replicated conversation commits; gateways hold only connection state. Kill a database node after acceptance and a gateway before delivery. Acknowledged history must survive both tests.
FR4 + NFR6: current membership Authorization on sends, history and live delivery. Remove a member with an old socket. Already released bytes cannot be recalled, but later admission checks must deny access.
NFR1–2: bounded group fanout Separate gateway/storage capacity and bounded dispatch for groups of up to 100 members. Load-test multi-device fanout and reconnect waves. Connection-count arithmetic alone does not prove the delivery target.

13Rapid revision

Remember: Save before acknowledging; use history to recover delivery when a socket or receipt disappears.

Decision What it provides What it does not provide
Commit before reporting accepted Recoverable saved message Immediate delivery to an offline device
Stable client message ID Safe retry of one logical send Deduplication of intentionally repeated text
Conversation sequence One saved message order per conversation Global order across unrelated chats
Outbox dispatch Saved delivery work survives crashes Exactly-once network transmission
Device cursor advances without gaps Gap detection and reconnect recovery Proof that a human read the message
Gateway directory A route to currently connected devices Durable history or current membership by itself
Conversation partition Order each chat; process separate chats in parallel Unlimited throughput for one group

A complete interview answer follows send-71 from the sender's pending state through commit, delivery to one phone and later laptop catch-up. Then use the workload to separate gateway and storage scaling. End with the lost-response and lost-receipt cases: one tests the server's acceptance boundary, the other tests the device's durable progress. Broadcasting to a million-member channel and designing encryption keys are additional problems to negotiate.

Practise the interview questions

Say your answer aloud before opening the model answer. Then answer the follow-up and compare the reasoning.

Foundation · Question 1

What does accepted mean?

Reveal a model answer

The message has been committed under the storage durability policy, independently of recipient connectivity.

What the answer must demonstrate: Keeps durability, device receipt and reading as separate outcomes.

Foundation · Question 2

Does WebSocket make chat reliable?

Reveal a model answer

It provides a bidirectional transport, not durable message storage or retry semantics.

What the answer must demonstrate: Attributes reliability to storage and recovery rather than a transport name.

Applied · Question 3

Why deduplicate by identity rather than text?

Reveal a model answer

Identical text can be sent intentionally twice. The stable send identity distinguishes a retry from a new message.

What the answer must demonstrate: Separates intentional repeated content from repeated transmission.

Applied · Question 4

Why use a conversation sequence?

Reveal a model answer

Devices can disagree about network arrival and client clocks. One committed sequence provides their shared history order.

What the answer must demonstrate: Uses local ordering without unnecessary global coordination.

Applied · Question 5

What if sequence 1042 arrives after cursor 1040?

Reveal a model answer

Fetch the missing authorized history before advancing a contiguous delivered cursor.

What the answer must demonstrate: Detects gaps and persists device progress safely.

Applied · Question 6

What survives a gateway crash?

Reveal a model answer

Committed messages and device recovery identities survive in storage; clients reconnect and catch up.

What the answer must demonstrate: Distinguishes replaceable connections from authoritative message state.

Applied · Question 7

Why can groups cost more than their stored messages?

Reveal a model answer

One message may be delivered to many members and devices, creating fanout work beyond one database insert.

What the answer must demonstrate: Counts recipient amplification and bounds work.

Follow-up · Question 8

What changes with end-to-end encryption?

Reveal a model answer

Clients manage message-encryption keys, and servers store ciphertext rather than freely processing content.

What the answer must demonstrate: Recognizes that encryption changes product and recovery semantics.

Blank-page exercise · 45 minutes

Build the answer yourself

Design durable text chat with multiple devices, bounded groups and offline history. Trace one accepted message while a recipient phone is online and laptop is asleep.

  • Use 5 minutes to agree numbered functional and non-functional requirements: statuses, membership, 100-member groups, online latency, history and acknowledged-message durability.
  • Trace baseline send and reconnect in 8 minutes.
  • Estimate messages, sockets and group fanout in 7 minutes.
  • Specify identities, sequence and API records in 10 minutes.
  • Scale gateways/storage and test lost responses and receipts in 10 minutes.
  • Use 5 minutes to check the final design against the numbered FR/NFR lists, test accepted-message survival and online/offline delivery separately, and state encryption or broadcast exclusions.

Check that each component and design decision follows from your requirements and workload.

Recall the key ideas

Answer from memory before opening each card. Explain why the choice works and what it costs. Revisit missed cards tomorrow.

Design a chat messaging serviceThe server saves Nora’s message while Sam’s phone is online and laptop asleep. Which statuses can each device claim?Recall first, then reveal

The server can report accepted after commit. The phone reports delivered after saving the message and cursor; the laptop catches up later. Read is a separate client report, not proof of attention.

Save, deliver, report read.

Return to lesson
Design a chat messaging serviceA send times out. What must the client keep when retrying?Recall first, then reveal

Reuse the same client message ID and original payload so the server can return the first send’s result.

One intent, one identity.

Return to lesson
Design a chat messaging serviceHow does a device recover messages received while it was offline?Recall first, then reveal

Fetch ordered conversation history after its saved cursor, then advance the cursor only through messages it has durably saved without gaps.

Sockets notify; history recovers.

Return to lesson

Final revision

Summary and interview notes

Save messages before acknowledging them. Sockets and push help online devices receive them quickly; saved device cursors and ordered history recover missed delivery without skipping messages.

Remember these points

  • Acknowledge only after commit.
  • Keep retry identities independent of body equality.
  • Order within each conversation.
  • Track each device’s contiguous progress separately.

Interview tips

  • Describe the asleep device as well as the connected phone.
  • Measure acceptance and delivery lag independently.

Important qualifications

  • Complex leave/rejoin histories and strict in-flight revocation require deeper membership protocols.

Continue after the core interview

Explore the advanced version

The advanced lesson keeps the full detailed design. Use these sections when you want to examine the stronger requirements and failure cases.

Technical references

Practice marks stay in this browser.