System-design interview · Core interviews
Design a chat messaging service
Design durable message acceptance, ordered conversation history and device recovery; scale connection gateways separately from message storage and recipient delivery.
You will learn to
- Distinguish accepted, delivered, and read using one concrete message.
- Design durable per-conversation order and duplicate-safe retries.
- Scale connection routing without making online presence the source of truth.
Practice in this chapter
8 interview questions with model answers and follow-ups.
Go to interview practiceUseful foundations: Real-time communication: polling, long polling, SSE, and WebSocket · Message queues, event logs, delivery guarantees, and backpressure · Replication and durability
Workload and timing examples are interview assumptions.
Dotted concept links open the relevant explanation in a new tab.
01Problem and scope
Scope the design to durable text, one-to-one conversations, then bounded groups, with multiple devices and advisory presence. An offline recipient does not make a send fail: acceptance depends on storage, and the recipient catches up later. Mobile push is a wake-up hint. Acknowledging a message before durable storage would violate this contract, so the sender receives success only after commit.
Support text, history, multiple devices, delivered/read receipts, online indicators, bounded groups and optional mobile push. Exclude attachments, end-to-end encryption protocol design, message editing and globally ordered history across unrelated conversations from the core exercise. If encryption becomes mandatory, discuss key distribution, multi-device recovery and server search limitations separately. The design separates connection capacity, durable acceptance and recipient delivery because each can reach capacity or fail independently. The conversation owner is the storage leader responsible for ordering that conversation’s messages and membership changes; its replicas preserve committed history if the leader fails.
02Functional requirements
A device cursor records the conversation position through which that device has safely applied history without gaps. Its progress is monotonic: later reports can advance it but cannot move it backward. Presence has a different purpose: a renewable, expiring lease estimates whether a device is still connected; it says nothing about which messages the device stored.
- Send: Return accepted only after the chosen durable commit boundary.
- Receive live: Low-latency hint/message delivery with client deduplication.
- Fetch history: Ordered bounded pages with explicit membership/history boundary.
- Reconnect: Recover every retained authorized message after the device cursor.
- Report receipts: Monotonic device progress; read is distinct from delivered.
- Show presence: Lease-based estimate, never a requirement for durable send success.
- Change group membership: Membership change serialized with send authorization.
Client identity and ordered history
The sender's client creates stable clientMessageId=send-71 before sending. Repeating that identity with the same body returns the same accepted message; changing the body under it is a conflict. The recipient's devices independently resume their progress. An online event may arrive twice or out of order, but the displayed committed history uses the server's conversation sequence and one identity per message.
Membership, previews and pending sends
A removed group member cannot obtain newly unauthorized messages simply because a gateway still has an open socket. Decide history semantics explicitly: new members initially see messages from their join sequence onward; departed users cannot fetch new history after removal. Push previews avoid disclosing sensitive body text on a locked device by default. The sender may see a local pending bubble immediately, but only an accepted response gives it a committed sequence. A network timeout leaves the outcome unknown until retry resolves the same identity.
03Non-functional requirements
- Acceptance latency: Accepted-message p95 below 200 ms within the home region.
- Delivery latency: Online delivery p95 below 500 ms under normal load. Measure this separately: a fast commit can hide a slow delivery queue.
- Availability: 99.9% eligible send availability. If a conversation owner loses its majority, reject new sends or leave them pending on the client; never acknowledge from a minority to preserve an online dot.
- Durability: Accepted messages survive one storage-node or availability-zone loss through replicated authority.
- Regional recovery: Declare a separate asynchronous recovery objective and measured loss window. Three replicas in one region do not imply zero regional data loss.
- Retention: Keep history for an illustrative five years; changes affect storage and legal/product behavior. Scope each device cursor to a conversation and its permitted history range.
Message and receipt invariants
| Invariant | Required behavior |
|---|---|
| Stable message identity | One (conversationId,senderId,clientMessageId) produces one immutable message. |
| Conversation order | Committed sequences define one order; network delivery may arrive out of order. |
| Durable acceptance first | Delivery cannot precede durable acceptance. |
| Monotonic progress | Receipt cursors never decrease; clients fetch gaps before advancing the contiguous cursor. |
| Serialized membership | Group sends and membership changes share one authority, preventing stale-cache send authorization after removal. |
A read receipt is a client assertion, not proof of a human's attention. Acceptance, delivery and reading are separate observable outcomes.
04Capacity estimates
Workload assumptions and arithmetic
Use 500 million daily active users sending forty messages/day: 500M × 40 = 20B messages/day, or 20B / 86,400 = 231,481 writes/s. Fivefold peak is about 1.16 million writes/s. At 100 bytes of text, bodies total 2 TB/day and 3.65 PB over five years. A 300-byte stored envelope including IDs and metadata gives 6 TB/day and 10.95 PB over five years, before indexes, copies and retention cleanup.
Assume 10% of daily users are online simultaneously with 1.2 connected devices each: 500M × 0.10 × 1.2 = 60M connections. At a measured 20,000 active connections per gateway, the fleet needs 3,000 gateway-equivalents before failover reserve. That per-node figure must come from a representative TLS, heartbeat and messaging benchmark, not a universal limit. Thirty-second heartbeats produce 60M / 30 = 2M heartbeats/s even with no chat messages.
Worked estimates
| Work | Calculation | Why it matters |
|---|---|---|
| Mean text ingress | 2 TB / 86,400 ≈ 23.1 MB/s |
Bodies alone understate protocol traffic |
| Mean stored envelope ingress | 6 TB / 86,400 ≈ 69.4 MB/s |
Replication multiplies internal traffic |
| 100-member group, two devices/member | Up to 200 live sends/message | Recipient fanout differs from stored-message count |
| Illustrative socket memory | 60M × 32 KB = 1.92 TB fleet-wide |
Buffers dominate a tiny session-directory record |
Capacity implications and limits
A one-to-one read/write bandwidth ratio does not apply once groups, retries and multiple devices are included. Partition the append log for write throughput; scale gateways for sockets and delivery bandwidth; budget presence separately.
05APIs and contracts
Request and response example
The sender sends {"type":"send","conversationId":"c8","clientMessageId":"send-71","body":"Train arrives at six"} on an authenticated connection. The server replies {"type":"accepted","messageId":"m901","sequence":1042} only after commit. The same operation is available as an authenticated HTTP POST for retry/fallback. A body mismatch under send-71 returns 409 instead of silently replacing text.
Interface contracts
| Interface/event | Contract |
|---|---|
POST /conversations/c8/messages |
Accepted durable identity or explicit retryable rejection |
GET /conversations/c8/messages?after=1040&limit=100 |
Ascending committed history from the permitted range |
delivered {conversationId:c8,through:1042} |
This device has durably applied the contiguous visible stream |
read {conversationId:c8,through:1042} |
Monotonic user/device read report, separate from delivery |
GET /conversations?cursor=...&limit=50 |
User's conversation summaries with stable cursor |
presence.subscribe [u17,u31] |
Bounded subscription to relevant visible users |
Validation and response semantics
Connection handshake authenticates user and device. Each send/history operation still checks conversation membership at the authority. Payload size, per-sender message rate and group size are bounded; use 429 or a structured retryable error for overload. A reconnect includes device identity and per-conversation cursors, not a claim that every earlier push was received. History may include tombstones/control events needed to preserve cursor continuity while hiding removed content. Pagination uses sequence positions rather than wall-clock timestamps, which can tie or move backward.
06Data model and access patterns
The model needs both a shared message order and separate device progress. Conversation sequences order accepted history; client message IDs identify send retries; device cursors track delivery to each phone or laptop. Membership records govern access, and an outbox row records delivery work in the same commit as the message so a dispatcher can recover it after a crash.
| Record and fields | Responsibility / constraint |
|---|---|
Message(conversationId,sequence,messageId,senderId,clientMessageId,payloadHash,body,acceptedAt) |
Primary key (conversationId,sequence); unique (conversationId,senderId,clientMessageId). |
Conversation(conversationId,nextSequence,ownerEpoch) |
Allocates sequence inside the message transaction. |
Membership(conversationId,userId,role,currentState) |
Supplies current authorization. |
MembershipInterval(conversationId,userId,fromSequence,toSequence) |
Records every permitted history interval; its end is exclusive and an active interval has no end. |
Outbox(eventId,conversationId,sequence) |
Commits with the message. |
DeviceCursor(userId,deviceId,conversationId,deliveredThrough,readThrough) |
Stores monotonic device progress. |
Join, removal and rejoin are ordered control events in the same conversation log. One join timestamp cannot represent a member who leaves and rejoins: replacing it either leaks the absent interval or hides permitted earlier history.
The query WHERE conversationId='c8' AND sequence>1040 ORDER BY sequence LIMIT 100 is a range read. Partition by conversation, optionally using bounded time/sequence buckets for older history while retaining a conversation routing index. Hashing each message ID independently would scatter the exact range query we need. User-to-conversation summaries form a derived inbox index; they are not a second authority for the message body.
A session directory maps (userId,deviceId) to gateway and connection generation with a lease expiry. A newer generation supersedes a disconnected socket; delivery failures refresh directory state. The directory and presence caches are derived from active sessions and may be stale. Log durability never depends on them. Store only bounded recent pages in memory, such as the latest messages of visible conversations; archival history has different retrieval and redundancy economics. Both partitioned SQL and suitable wide-column systems can support the append/range pattern when their real consistency and throughput meet the contract.
07Basic working design
Conversation commit and local sockets
Start with one chat process, one SQL database and connected clients. The process holds a local map of authenticated device connections. The sender's send begins a transaction: check membership, look up send-71, allocate c8 sequence 1042, insert m901 and an outbox row, then commit. Only after commit return accepted. The dispatcher reads the outbox and writes to the recipient's socket if connected; otherwise the history remains available for catch-up.
Recipient persistence and retry recovery
The recipient's phone stores m901, advances its contiguous local cursor and reports deliveredThrough 1042. If the report disappears, it can repeat it. The recipient’s laptop later reads after 1040 and obtains 1041 and 1042. The sender does not need to retype the message because the recipient happened to be offline. A database row plus recovery query is more important than a perfect live notification path.
Long-poll baseline and its limits
Long polling can implement delivery initially: hold a request until new data or timeout, then reopen it. WebSocket provides a persistent bidirectional transport and may reduce repeated HTTP setup. Neither transport makes a message durable in the database or prevents a retried send from inserting a duplicate. The baseline already differentiates pending, accepted, delivered and read states, so later gateway fleets preserve those semantics instead of redefining “sent” whenever a process crashes.
The SQL commit precedes accepted; the connection map only helps low-latency delivery.
Read each connection in order
- sync1. Send send-71Sender client → Chat application and socket map
- sync2. Commit m901 / sequence 1042Chat application and socket map → SQL conversation log and outbox
- sync3. Accepted after commitChat application and socket map → Sender client
- sync4. Deliver or await catch-upChat application and socket map → Recipient devices
- sync5. Read history after cursorRecipient devices → Chat application and socket map
08Find the baseline flaws
| Bottleneck / counterexample | Evidence and design consequence |
|---|---|
| Connections and write throughput | One process cannot serve 60 million connections or 1.16 million peak writes/s. Even if it could accept enough file descriptors, memory, TLS processing, heartbeat work and outbound bandwidth would saturate. Increasing database connections to match every socket is especially harmful: most sockets are idle, while storage needs a bounded worker/connection pool. Queueing unbounded send tasks hides overload until memory collapses. |
| Acknowledgment before durable commit | The key correctness counterexample is acknowledgement before storage. At t0 the process receives send-71; at t1 it tells the sender “sent”; at t2 it crashes before the insert. The sender and the recipient can never reconstruct m901 from history. Persistence that occurs only after acknowledgment is incompatible with a promise that accepted messages are already durable. Asynchronous execution is fine inside an implementation as long as the user acknowledgement waits for the required commit. |
| Clock order and receipt gaps | Another counterexample uses timestamps as order. Two users send concurrently through different gateways; their clocks and event arrival order differ. If each client displays its local message first forever, devices disagree. The conversation owner must assign the committed sequence, and optimistic local bubbles may move when acceptance arrives. Finally, delivering to a socket does not remove the need for pending recovery: a connection may die after the kernel accepts bytes but before the client persists them. |
09Improve the design, step by step
Split connection gateways from conversation storage. When one process cannot manage the connection load, move sockets to gateways with bounded event-driven buffers and an expiring session directory. Storage workers consume a bounded number of concurrent requests, independent of idle connection count. This improves isolation and allows gateway replacement. It costs routing hops, directory staleness and reconnect logic. A single process remains simpler for small communities; a separate thread per socket is not required by the product.
Partition and replicate conversation authority. Write throughput and accepted-message durability trigger many logical conversation partitions, each with a replicated leader. Membership checks, sequence allocation, deduplication and outbox insertion remain one partition-local atomic operation. This distributes independent conversations but leaves a very hot group on one owner. A globally ordered log is rejected because unrelated conversations need no shared order; splitting one huge group's order would change the product semantics.
Decouple delivery with durable outbox dispatch. A slow recipient gateway or push provider must not delay durable acceptance. Commit m901 first; delivery workers then route notifications and update conversation summaries. The improvement is recoverable delivery without blocking acceptance on the recipient's device. The cost is queue lag and at-least-once delivery, meaning events may repeat. Exactly-once transport is not assumed; clients and storage deduplicate identities. Direct post-commit delivery may remain a low-latency fast path, but the durable outbox repairs missed attempts.
Bound history/presence work and introduce archive tiers. Five-year storage and millions of heartbeats trigger recent-page caches, time/sequence history buckets, colder storage and selective presence subscriptions. This reduces hot storage and unnecessary broadcasts. It costs archive latency, cache misses and advisory status. Erasure coding can reduce cold-history redundancy overhead but changes repair/read behavior; it is not a substitute for protecting current writes. Broadcasting every heartbeat to every friend is rejected because the amplification has little user value.
Erasure coding stores data fragments together with additional encoded fragments so the data can be reconstructed after the supported number of fragment losses. It can use less space than several full copies, but reconstruction and repair require extra work. That tradeoff may suit old history while newly accepted messages retain the chosen fast replicated commit path.
All steps preserve the same accepted boundary. A faster socket acknowledgement that drops durability would be a contract regression, not a performance improvement.
10Detailed architecture
Connection routing and conversation ownership
The edge balances new connections among healthy gateways; existing traffic stays on its authenticated socket. Gateways forward c8 sends through a conversation router that resolves the current owner and its ownership version, called an epoch. The owner checks membership and commits the ordered log, request deduplication and outbox to a replica group. A majority failure stops acceptance for that partition; an obsolete owner is rejected by storage-epoch checks.
Live delivery and push hints
Outbox dispatchers resolve each recipient's current devices through the session directory, then send to the appropriate gateways. Group fanout expands one log entry into bounded recipient tasks. An offline-device worker may send a privacy-conscious wake-up to the external mobile push provider, but provider acceptance is not a delivered/read receipt. Conversation summary indexes and device cursor stores support user views and recovery.
Authorized recent and archived history
History APIs route to the authoritative or sufficiently current log and may use recent immutable page caches. They check membership and allowed history bounds before returning entries. Presence uses a separate directory whose entries expire unless devices renew them. It combines rapid online/offline changes before notifying subscribed contacts, reducing status flicker and unnecessary updates. Its failure can make an online dot inaccurate without losing messages.
Acceptance boundary and connection leases
Synchronous send work ends when m901 is durably committed; live delivery, push, receipts and inbox summaries occur afterward. Receipts themselves become durable monotonic updates so device reconnects do not move progress backward. The diagram separates these steps so a successful send is not confused with successful device delivery.
Concrete storage and transport choices
A practical baseline uses PostgreSQL for the message, membership, request identity and outbox transaction, with separate gateway processes and a disposable leased routing cache. At the large illustrative load, move whole conversation partitions onto storage that actually supports the required atomic operation and safe leader changes. PostgreSQL synchronous replication, Cassandra-style replica counts and a custom Raft-backed log have different guarantees; a “quorum” label does not make them interchangeable. Verify committed-history preservation and current authorization reads in the chosen implementation.
Conversation owners serialize messages and membership. Gateways and external push deliver hints; durable history remains the recovery path.
Read each connection in order
- sync1. Connect and authenticateSender and recipient devices → Connection load balancer
- sync2. Route socketConnection load balancer → WebSocket / long-poll gateways
- sync3. Renew device session leaseWebSocket / long-poll gateways → Leased session and presence directory
- sync4. Send c8 / send-71WebSocket / long-poll gateways → Conversation owner router
- sync5. Route current owner epochConversation owner router → Conversation command and history service
- sync6. Atomic message + outboxConversation command and history service → Partitioned message log and outbox
- replication7. Replicate before acceptedPartitioned message log and outbox → Conversation replicas
- async8. Consume committed deliveryPartitioned message log and outbox → Delivery and group fanout workers
- sync9. Resolve live recipient devicesDelivery and group fanout workers → Leased session and presence directory
- async10. Route committed m901Delivery and group fanout workers → WebSocket / long-poll gateways
- sync11a. Deliver committed messageWebSocket / long-poll gateways → Sender and recipient devices
- sync11b. Device receipt / history requestSender and recipient devices → WebSocket / long-poll gateways
- sync11c. Route receipt / history requestWebSocket / long-poll gateways → Conversation owner router
- async12. Update derived inboxDelivery and group fanout workers → Device cursors and inbox index
- sync13. Monotonic receipt updateConversation command and history service → Device cursors and inbox index
- sync14. Authorized recent history readConversation command and history service → Recent history cache
- sync15. Range fetch after cursorConversation command and history service → Partitioned message log and outbox
- async16. Offline wake-up workDelivery and group fanout workers → Push notification worker
- async17. Send expiring hint without message textPush notification worker → Mobile push provider
- async18. Wake applicationMobile push provider → Sender and recipient devices
11Write path and acknowledgement
One conversation owner serializes membership, retry identity and sequence allocation in the acceptance transaction. Request send-71 becomes message m901 at sequence 1042; notification delivery starts only after that commit.
- The sender's client journals send-71 with its body before transmission and shows a pending bubble. It authenticates the connection and supplies conversation c8.
- The gateway enforces size/rate limits and forwards to c8's current owner. It does not invent a final sequence or accepted response.
- The owner begins a transaction, checks current membership and the unique sender/client-message identity. If already committed with matching payload, return its existing m901/1042 result.
- For a new request, allocate the next conversation sequence under the same owner lock, insert m901 and its deduplication identity, and append outbox event
deliver-c8-1042in the transaction. - Commit through the required replica quorum. Only now emit accepted 1042. A lost response is an unknown transport outcome; retrying send-71 returns the same message.
- The dispatcher reads the outbox, expands eligible recipients and resolves the recipient's phone gateway. It sends the committed message or a hint to fetch it, recording bounded retry state.
- If the recipient is offline, retain the log and schedule optional push. No client-visible send failure is inferred from advisory presence. When a delivery attempt fails, retry routing without re-inserting the message.
The conversation lock orders a membership removal against a send. If removal commits first, the send is rejected; if send commits first, its eligibility follows the defined membership boundary. This decision is local to c8, not a global transaction across every participant's inbox.
12Read and delivery path
Each recipient device maintains its own contiguous cursor. The following recovery path starts at sequence 1040 and handles message 1042 arriving before the intervening history has been applied.
- The recipient's phone receives committed c8 sequence 1042. If its durable cursor is 1040, it detects the gap and fetches from 1040 instead of claiming that everything through 1042 arrived.
- The history service authorizes the recipient and returns a bounded ordered page, including any necessary control/tombstone positions. It does not depend on whether the earlier gateway still exists.
- The phone atomically stores newly applied message identities and its new contiguous cursor in one local transaction, then sends deliveredThrough 1042. Advancing the cursor before saving the message could permanently skip it after a client crash. A repeated m901 event produces no second bubble.
- The server updates that device cursor with a maximum operation, so a delayed receipt for 1041 cannot move it backward. A read action generates a separate readThrough update constrained by the application's receipt policy.
- The recipient's laptop reconnects later with its own cursor 1040 and repeats the authorized history path. Phone delivery does not incorrectly advance the laptop's state.
- The sender's client receives updated receipt summaries asynchronously. It may display “delivered to a device” or another explicitly chosen aggregation, rather than implying every device or the human saw the text.
For group c9, the same log is stored once while delivery reaches multiple authorized members/devices. Large groups may receive lightweight wake-ups and fetch history in bounded pages. Presence subscriptions fetch an initial relevant snapshot and receive debounced changes; neither an online dot nor push-provider response is proof that a message was read.
An authorized page carries an opaque continuation through the log, including safe skip/control positions for intervals the device cannot read. Those positions reveal no hidden message body; current membership and stored history intervals still filter every returned message. If the cursor predates retained history, return an explicit history-expired/reset response with the earliest retained position instead of making the device fetch an unfillable gap forever. A reset acknowledges the retention limit; it does not claim the deleted history was delivered.
13Correctness deep dive
Identity and sequence authority
The storage transaction, not the transport, decides the message identity. Deduplication is scoped to (c8,sender17,send-71) and compares the payload hash. The sequence increment and insert share that transaction; a sequence reserved outside it would complicate contiguous recovery and failure handling.
| Actor/event | Atomic authority action | Observable result |
|---|---|---|
| the sender sends send-71 first time | Insert m901 at 1042 plus outbox | One committed message |
| Concurrent retry through another gateway | Unique identity detects m901 | Return same 1042; no second insert |
| Leader crashes before commit | Uncommitted state not accepted | Client retries safely |
| Leader crashes after commit but before response | New leader reads committed identity | Same accepted result |
| the recipient repeats delivery receipt | deliveredThrough = max(old,new) |
Progress cannot regress |
Lost receipt after device persistence
Stale gateway and cursor recovery
A stale G1 entry in the directory may receive another attempt. Its connection-generation check prevents directing data onto an unrelated reused session, while current recipient/membership checks protect disclosure. Failed routing is retried or left for cursor recovery. During a storage leadership change, the old leader's epoch must be fenced from commits; two leaders assigning sequence 1042 independently would violate the invariant. The replica protocol establishes that authority, while the gateway directory merely locates sockets.
Intentional duplicates versus retries
A timed-out client may have sent the same text twice intentionally under two different IDs; the service keeps both. Content equality is not a valid deduplication rule for chat.
Membership intervals at delivery
The accepted message is already durable. The recipient deduplicates sequence 1042 and resends a monotonic receipt after gateway failure.
Read each connection in order
- syncSend send-71Sender client → Conversation owner
- syncCommit m901 at 1042 plus outboxConversation owner → Conversation owner
- returnAccepted 1042Conversation owner → Sender client
- syncDispatch committed m901Conversation owner → Delivery gateway
- syncDeliver m901 / 1042Delivery gateway → Recipient device
- syncPersist once and advance cursorRecipient device → Recipient device
- blockedReceipt lost as gateway failsRecipient device → Delivery gateway
- syncRetry via new gatewayConversation owner → Delivery gateway
- syncRepeat m901 / 1042Delivery gateway → Recipient device
- syncDeduplicate; deliveredThrough 1042Recipient device → Conversation owner
14Failure and recovery
| Failure / trigger | User outcome, surviving state and recovery |
|---|---|
| Gateway crash | Thousands of sockets break. Clients reconnect with randomized backoff, reauthenticate, retry pending sends under existing identities and fetch after durable cursors. We do not try to transfer live TCP state between arbitrary hosts. The durable log survives, and the user may briefly see reconnecting. Duplicate events are expected and deduplicated. |
| Conversation-owner partition | A minority cannot safely accept new messages. The sender's client keeps send-71 pending and retries; it does not show accepted. Existing authorized history may be served under its consistency policy, but new membership decisions and sends require authority. After majority recovery, retry resolves whether the earlier attempt committed. Single-zone failover and region-wide disaster recovery remain distinct guarantees. |
| Delivery backlog at peak | Acceptance is about 1.16 million messages/s before group fanout. Bound queue age and per-group work; throttle abusive senders before acceptance and reduce nonessential presence notifications. If durable storage or outbox capacity is exhausted, reject new sends rather than accept an unbounded future delivery obligation. Protect history catch-up capacity so reconnects can drain the backlog. |
| Push-provider outage | Messages remain in history. Retry wake-ups within a bounded lifetime and let app reconnect perform catch-up. A missing push notification is not a lost message, and a successful provider request is not recipient delivery. Cold archive outages may temporarily affect old history while recent conversations work; surface the distinction instead of claiming all history has the same latency tier. |
15Operations, security, and cost
Separate acceptance, delivery and read metrics
Measure pending-to-accepted latency, accepted-to-device-delivered lag and delivered-to-read reports separately. Alert on outbox age, duplicate-send retries, owner failovers, replica lag, reconnect rate, gap-fetch frequency and hot-group queue length. Presence churn has its own budget. A high send-success rate can coexist with a broken delivery system, so acceptance alone cannot be the service dashboard.
Retention, sockets and group-fanout costs
The main costs are retained message copies, gateway resources for sockets, and delivery to each recipient device. Three copies of the 10.95 PB five-year envelope estimate need 32.85 PB before indexes and backups. If recent-page caching saves a history query but duplicates every user's entire five-year history in memory, it loses economically. Cache bounded visible conversations and measure reuse. Large-group delivery can be batched by gateway so one payload serves several local recipients, trading gateway CPU against inter-server bandwidth.
Message authorization and abuse controls
Authorize every send/read, bound group membership and message size, and rate-limit spam per account and conversation. Avoid putting sensitive text in logs and default push previews. Transport encryption does not equal end-to-end encryption; if the latter is required, the server stores ciphertext and the key-management design changes the product's recovery/search behavior.
Crash, reconnect and migration drills
Test crashes before/after the message commit and before/after client receipt persistence. Simulate stale gateway directory entries, leader fencing, group-member removal during send, and a large reconnect wave. Roll out storage migrations by copying a conversation partition, replaying its log, fencing old ownership and comparing range reads. Changing partition layout must not renumber committed messages or reset deduplication identities.
16Decision ledger and limitations
| Decision | Benefit | Cost / consequence | Reconsider when |
|---|---|---|---|
| Conversation-local sequence | Shared stable history across devices | One huge conversation has an ordering owner | Product accepts weaker ordering or partitioned threads |
| Commit before accepted | Survives declared storage failures | Replication latency on send path | Product explicitly accepts ephemeral messages |
| At-least-once delivery plus identities | Recoverable crashes and retries | Consumer deduplication and cursor logic | Never replace with a claim of exactly-once network |
| Separate socket and storage tiers | Independent connection/write scaling | Routing directory and extra hop | Small deployment benefits from one process |
| Advisory presence leases | Cheap bounded online estimate | Status may be stale | A stronger presence requirement justifies higher cost |
User-ID partitioning gives local user-history reads but can duplicate a conversation across participants and complicate one shared ordering authority. We choose conversation ownership plus a derived per-user conversation index. Time/sequence buckets bound old-history partitions. A wide-column log-oriented engine can fit append/range access, but rejecting SQL categorically is unjustified; benchmark the actual transaction, partition and storage requirements.
WebSocket and long polling are transport alternatives. WebSocket reduces repeated request setup and supports two-way events; long polling works through ordinary request infrastructure but reconnects frequently. Frequent short polling is simpler at tiny scale and wastes more empty work at high connection counts. None eliminates the durable offline log. Cold erasure-coded storage can save redundancy bytes but brings repair and read-latency tradeoffs; preserve fast replicated protection for new acknowledged messages.
17Interview closing
“I designed durable text chat with one-to-one conversations, groups, history, presence and multi-device recovery. A send is accepted only after the conversation owner commits its stable request identity, immutable message, ordered sequence and delivery outbox. Duplicate attempts recover that result. Every recipient device has an independent cursor, so live delivery can repeat or fail without changing committed history.
“The workload is about 231,000 average writes per second and sixty million assumed live connections, so socket gateways and storage partitions scale separately. Conversations own order and membership; gateways locate devices. I use bounded asynchronous delivery and push only as a wake-up mechanism. Presence is a lease-based hint rather than evidence that a message was read.
“I accept temporary send unavailability when a conversation lacks safe write authority. The remaining bottlenecks are very hot groups, recipient amplification and reconnect storms. I would next measure acceptance and delivery lag separately under a gateway-failure load test.”
If the interviewer requests a million-member broadcast group, avoid extending the small-group fanout loop blindly. Store one ordered channel log, send coalesced update hints, and have active subscribers fetch pages through caches. Re-estimate moderation, bandwidth and ordering requirements while preserving the accepted-message boundary.
Practise the interview questions
Say your answer aloud before opening the model answer. Then answer the follow-up and compare the reasoning.
What exactly does your send acknowledgment guarantee?
Reveal a model answer
It means the message and its delivery event are durably committed under the stated replica-failure policy. It does not mean the recipient is connected or has reported it as read. I expose delivered and read as separate later states.
Interviewer follow-up
Why not acknowledge immediately for low latency?
Reveal the follow-up answer
Then a gateway crash could erase a message shown as sent. I can show a local pending bubble immediately, but the accepted state waits for storage.
What the answer must demonstrate: The acknowledgment must name a failure guarantee.
The sender sends from two devices at the same instant. Which message comes first?
Reveal a model answer
The conversation owner assigns committed sequence numbers. Both devices reconcile pending bubbles to that order. I do not compare client wall-clock timestamps because ties and clock skew are normal.
Interviewer follow-up
Does every conversation share that sequencer?
Reveal the follow-up answer
No. Ownership is partitioned by conversation, so unrelated chats progress independently. A single unusually hot conversation has its own capacity limit.
What the answer must demonstrate: Scope the ordering guarantee.
A send request times out after submission. How do client and server prevent the retry from creating another message?
Reveal a model answer
The sender client retries the same clientMessageId within the same conversation. The conversation owner enforces uniqueness on (conversationId, senderId, clientMessageId) and returns the accepted message and sequence for the matching payload. Reusing that key with changed content is a conflict. Sender identity alone is not the key, and the gateway does not decide acceptance.
Interviewer follow-up
What if delivery itself is duplicated?
Reveal the follow-up answer
Recipients also deduplicate by message ID or sequence. API idempotency and consumer deduplication protect different boundaries.
What the answer must demonstrate: A random new retry ID defeats the guarantee.
How would you build the green online dot?
Reveal a model answer
A device renews a short lease through heartbeats. An expired lease means probably offline. I fetch initial status and subscribe for visible contacts, with a small delay to avoid flicker.
Interviewer follow-up
Can that status decide whether to accept a message?
Reveal the follow-up answer
No. Presence may be stale and offline messages are supported. Durable history determines eventual delivery.
What the answer must demonstrate: Presence is not a durable delivery test.
How does a group of 100 people change your design?
Reveal a model answer
I store one ordered conversation history, then fan out delivery events to member devices. Membership controls both sending and which history each member may read. The fanout queue absorbs short bursts.
Interviewer follow-up
What happens if a member leaves and later rejoins?
Reveal the follow-up answer
I record separate sequence-bounded membership intervals rather than overwrite one join time. Current membership authorizes the request, and the intervals determine which retained messages may be returned. The rejoining member does not automatically gain messages sent while absent; cursor responses explicitly skip hidden positions without returning their bodies.
What the answer must demonstrate: Account for devices and membership history, not only user count.
A gateway crashes after a recipient device persists a message but before the server saves its delivery receipt. What happens on reconnect?
Reveal a model answer
The recipient reconnects through another gateway and reports its last durably applied conversation sequence. The dispatcher may resend because it cannot know whether the earlier delivery completed. The device deduplicates by message ID or sequence, keeps one displayed message, and repeats its monotonic receipt. The committed server history never depends on that gateway surviving.
Interviewer follow-up
What would you monitor during a regional reconnect storm?
Reveal the follow-up answer
New connections, authentication load, catch-up read queries per second (QPS), queue age, and delivery lag. I would randomize reconnect delays to spread the burst and reserve capacity for durable acceptance.
What the answer must demonstrate: Recovery reads can exceed ordinary delivery traffic.
A group member is removed while their message is being sent. Which operation wins?
Reveal a model answer
The conversation owner serializes membership changes with send authorization. If removal commits first, the send is rejected. If the message commits first, its eligibility follows the preceding membership state and defined history rule. A gateway cache cannot make that final decision because it may be stale.
Interviewer follow-up
Does fanout need one transaction across every member inbox?
Reveal the follow-up answer
No. The log and outbox commit locally at the conversation owner. Recipient inbox updates and deliveries are recoverable derived work, with current access checks for history and media. I avoid a huge distributed transaction over every device. Live payload admission still checks the recipient’s current authorization. Removal cannot recall bytes already admitted before it committed.
What the answer must demonstrate: Identify one authority for the ordering decision, not independent cached checks.
What changes for a million-member broadcast channel?
Reveal a model answer
One stored message can imply millions of delivery attempts, so I keep the ordered channel log but coalesce wake-up notifications and let active subscribers fetch bounded pages through caches. I measure fanout and egress separately from message insert QPS. I would also revisit whether the channel truly needs interactive group semantics.
Interviewer follow-up
Would random message-ID sharding solve the hot channel?
Reveal the follow-up answer
It could distribute writes but destroy simple ordered range reads unless another ordering/index layer is introduced. If total order remains required, I first batch the owner path or change the product into partitioned threads rather than claim hashing alone solves it.
What the answer must demonstrate: Changing transport does not remove recipient amplification or ordering constraints.
Blank-page exercise · 45 minutes
Build the answer yourself
Design durable text messaging with groups and multiple devices. Separate acceptance, delivery and read receipts, size sockets and storage independently, then recover a gateway failure between device delivery and receipt persistence.
- Define accepted, delivered, and read.
- Calculate storage writes and concurrent socket capacity separately.
- Show sequence allocation and retry identity.
- Trace offline catch-up and duplicate delivery.
- Explain group fanout and advisory presence.
Check that each component and design decision follows from your requirements and workload.
Recall the key ideas
Answer from memory before opening each card. Explain why the choice works and what it costs. Revisit missed cards tomorrow.
Design a chat messaging serviceDoes accepted mean the recipient has read the message?Recall first, then reveal
No. Durable acceptance, device delivery, and user read state are separate events.
Saved → delivered → reported read.
Return to lessonDesign a chat messaging serviceThe same message is delivered twice. What prevents two bubbles?Recall first, then reveal
The client recognizes the stable message identity and tracks the last consecutive message it has saved with a conversation cursor that only moves forward. Repeated delivery therefore creates no second displayed message.
Retry transport; deduplicate effect.
Return to lessonDesign a chat messaging serviceCan an offline user receive a message?Recall first, then reveal
The service accepts it durably and their devices fetch it after reconnecting.
Presence is a hint; history is truth.
Return to lessonFinal revision
Summary and interview notes
Chat acceptance, device delivery and read reports are different events. The conversation’s storage group checks membership and saves the message, sequence, retry identity and outbox together. Each device uses its own cursor to recover missing messages from saved history.
Remember these points
- A WebSocket is a transport; accepted messages require the declared storage commit before acknowledgment.
- Conversation-local order scales across conversations, while a single hot conversation retains a sequencing limit.
- Persist messages and the contiguous device cursor atomically; retry delivery with stable identities.
- Membership intervals preserve leave/rejoin history boundaries, and current authorization still gates delivery and reads.
- Presence and push notifications are advisory; neither proves that a device stored a message or a person read it.
Interview tips
- Size sockets, heartbeat traffic, stored writes and recipient fanout separately.
- Walk through a crash after the device saves a message but before its receipt reaches the server.
- Define accepted, delivered and read before discussing protocol or database choices.
Important qualifications
- Removal cannot recall payload bytes already released to a socket; state the delivery admission boundary.
- Retention expiry requires an explicit cursor-reset response rather than an endless gap fetch.
- A replicated implementation must preserve committed conversation history during failover; replica counts alone are insufficient.
Technical references
- RFC 6455: The WebSocket ProtocolDefines the two-way TCP-based browser/server transport; it does not define chat durability.
- Amazon SQS at-least-once deliveryDocuments why consumers must tolerate repeated delivery.
- Transactional outbox patternExplains atomic persistence of business state and an event to be delivered.
Practice marks stay in this browser.