System-design interview · Extended interviews
Design a collaborative text editor
Build an online plain-text editor with immediate local typing, durable accepted operations and reconnect recovery, using a proven operational-transformation implementation.
You will learn to
- Explain why identified operations preserve concurrent edits better than whole-document replacement.
- Trace pending edits, accepted versions and reconnect without applying an edit twice.
- Scale documents and connections separately while preserving document authority and permissions.
Practice in this chapter
8 interview questions with model answers and follow-ups.
Go to interview practiceUseful foundations: Real-time communication: polling, long polling, SSE, and WebSocket · Quorums, consensus, leases, and fencing · Replication and durability
Workload and timing examples are interview assumptions.
Dotted concept links open the relevant explanation in a new tab.
01Choose online collaboration and define saved
Design plain-text documents shared by several users, with simultaneous edits, permissions, cursors, undo and recovery after short disconnections. Exclude rich formatting, embedded spreadsheets and months of offline editing from the first design. Those requirements change the operation model and retained metadata, not merely the number of servers.
Two users open document D7 containing cat at version 20. A inserts X at position 1 while B inserts Y at the same position. Both edits should survive under a defined ordering rule. Replacing the entire document with cXat and later cYat loses A's work even if each database write is individually atomic.
Distinguish immediate local display from saved state. A keystroke appears locally as pending; saved means the service has durably accepted that identified edit. A lost connection may leave a usable local draft without permission to claim it is saved. Presence is another category: a cursor is a temporary hint and need not survive a restart.
Assume documents up to 1 MB, at most 100 active participants and a thirty-minute automatic reconnect window. Target same-region acceptance below 150 ms p95 and remote display below 300 ms p95 on connected, responsive peers under admitted load, both measured from edit submission. Using a merge algorithm and WebSockets does not prove the service meets those latency targets.
Clarify whether the editor is plain text or rich text, how many users share one document, and how long disconnected clients must merge automatically. This answer supports plain text, up to 100 active participants per document and a selected thirty-minute automatic reconnect window; longer gaps preserve drafts for explicit resynchronization.
02Functional requirements
Open and share documents. Create or retrieve plain-text documents and assign current read/edit permissions.
Edit concurrently. Accept identified inserts, deletes and supported undo operations while preserving concurrent edits under the chosen transformation rules.
Collaborate live. Display local pending edits immediately and distribute accepted edits and temporary cursor/presence information to authorized participants.
Reconnect and recover. Retrieve saved content, replay missing accepted versions and reconcile pending operations after short disconnections without applying an edit twice.
03Non-functional requirements
These are illustrative interview assumptions, not product facts or measured benchmarks. Confirm them before choosing components, then validate the completed design under the stated workload. Here p95 means the 95th-percentile latency: 95% of measured requests take no longer than that value. Report errors and rejected work alongside latency; a fast failure is not a successful outcome.
Document and fleet size. Limit a document to 1 MB and 100 active participants. Plan for one million connected users and 100,000 incoming edits/s across documents; benchmark a single hot document separately.
Latency. Target same-region durable edit acceptance below 150 ms p95 and display on connected, responsive peers below 300 ms p95, measured from submission under admitted load. Slow or disconnected peers use bounded buffers and replay; immediate local display is not a save acknowledgment.
Convergence and ordering. All clients applying the same accepted history must converge under the supported editing protocol. Give every accepted edit one stable identity and version; whole-document last-write replacement cannot satisfy this requirement.
Saved-state durability and reconnect. Accepted edits and operation identities survive application/coordinator restarts. Retain the transformation and retry history needed for thirty minutes of automatic reconnect; older drafts require explicit resynchronization. Database disaster tolerance depends on the selected replication/backup policy, not socket reconnection.
Authorization. Check current permission when accepting edits and serving snapshots or replay. Reject stale coordinators at storage and reject edits after revocation takes effect; previously downloaded text cannot be recalled.
04Complete one edit with a coordinator and durable log
Start with one application serving document APIs and WebSockets, backed by a transactional database. For each active document, a coordinator holds the accepted content in memory and uses a proven operational-transformation implementation. Operational transformation adjusts an incoming edit to account for concurrent edits accepted since its base version. It requires compatible rules on both client and server.
Follow A17 through the edit path:
A opens D7 at version 20, displays X locally and submits operation A17 with its base version.
The coordinator checks current edit permission and whether A17 already exists.
It transforms the operation against intervening accepted history, then transactionally appends the accepted edit and advances the document version.
Only after durable commit does it acknowledge A17 and broadcast the result.
Clients reconcile the accepted operation with their pending local operations. A's own acknowledgment removes its pending entry; it must not insert X a second time. A server restart rebuilds accepted content from a snapshot and subsequent operations. Lost broadcasts are repaired by replaying missing versions.
This baseline already needs a real editing library, persistent operations and a reconnect protocol. WebSockets only provide a low-latency connection; they neither merge conflicting edits nor recover a message that vanished before reaching a peer.
Clients render pending edits; the coordinator acknowledges only durable accepted history.
Read each connection in order
- syncA17 based on version 20Client A: accepted + pending text → Document coordinator / OT library
- syncB9 based on version 20Client B: accepted + pending text → Document coordinator / OT library
- syncAuthorize and append accepted editDocument coordinator / OT library → Document, grants, operations, snapshots
- asyncAcceptance and remote editsDocument coordinator / OT library → Client A: accepted + pending text
- asyncAcceptance and remote editsDocument coordinator / OT library → Client B: accepted + pending text
05Walk through two concurrent insertions
Positions in the example are zero-based: position 1 lies between c and a. A17 and B9 both describe edits based on cat at version 20. Suppose the selected library's tie-break rule places A's same-position insert first.
| Accepted step | Operation and resulting text |
|---|---|
| Version 20 | Initial text is cat. |
| Version 21 | Accept A17: insert X at position 1, producing cXat. |
| Transform B9 | A inserted before B's target, so adjust B's insertion to position 2. |
| Version 22 | Apply transformed B9, producing cXYat. |
B had already displayed cYat locally. Its client transforms incoming A17 and its own pending edit consistently, preserving both letters and converging to cXYat. The rule makes clients agree on the text. It cannot know which word the users intended.
Deletes, overlapping ranges and undo need their own correct rules. Undo should reverse the selected edit using the library’s rules, preserving other users’ work where specified, rather than upload an old entire document. Choose one documented text representation for all clients. Byte offsets, UTF-16 code units, Unicode scalar values and visible grapheme clusters are not interchangeable; use the library's agreed units and conversion rules.
A worked insert example explains the mechanism but is not a complete algorithm implementation. Evaluate a maintained OT stack and its persistence adapter rather than treating this table as sufficient production code.
The shown tie-break orders A before B; the full library handles other edit types.
Read each connection in order
- syncA17: insert X at 1, base 20Client A → Coordinator
- syncB9: insert Y at 1, base 20Client B → Coordinator
- returnA17 accepted at 21: cXatCoordinator → Client A
- syncTransform B9 to position 2; persist 22Coordinator → Coordinator
- asyncB9 at 22: cXYatCoordinator → Client A
- asyncReconcile A17 and B9: cXYatCoordinator → Client B
06Size active typing and peer delivery separately
Assume one million connected users, five percent currently typing and two operations/s per typist. That yields 100,000 incoming edits/s. At 200 bytes per operation, payload ingress is 20 MB/s before framing, indexes and replication. If sustained for a full day, the operation payload is about 1.728 TB; actual duty cycles must be measured rather than assumed constant.
Ten other participants receiving each edit create one million peer deliveries/s. Connection memory can also be significant: one million connections at an illustrative 32 KB each consume about 32 GB before process and encryption overhead. Moving sockets and fanout to gateways may matter before optimizing the storage engine.
A single document with 100 typists at two edits/s has 200 ordered edits/s and about 19,800 peer deliveries/s. Sharding other documents does not split this document's edit order. Measure transformation CPU and fanout separately before claiming a global shard count solves the hotspot.
Snapshots trade storage work for shorter recovery. A 1 MB snapshot every 1,000 edits on that hot document means a snapshot every five seconds, or 200 KB/s for that document alone. Use a replay-byte target and minimum interval rather than an unexamined count-only rule. Bound per-client queued bytes to prevent slow viewers consuming unlimited memory.
07Name operation identity, base and accepted version
Interfaces
| Request or message | Contract |
|---|---|
GET /documents/D7 |
Returns authorized content, current accepted version and supported reconnect boundary. |
Edit message: operationId, baseVersion, operation |
Names one local edit and the accepted state it was based on. |
Acceptance message: operationId, version, canonicalOperation |
Confirms the durable order and lets the client reconcile pending state. |
These JSON examples show the worked insert and its acceptance. The selected editing library defines the actual operation encoding.
A17: client edit based on version 20
{
"operationId": "A17",
"baseVersion": 20,
"operation": {
"type": "insert",
"position": 1,
"text": "X"
}
}
A17: acceptance after durable commit
{
"operationId": "A17",
"version": 21,
"canonicalOperation": {
"type": "insert",
"position": 1,
"text": "X"
}
}
Stored records
| Record | Fields or identity | Purpose |
|---|---|---|
| Document | id, headVersion, ownerGeneration |
Current accepted head and coordinator authority. |
| Operation | document, version, actor, operationId, payloadHash |
Ordered replay and duplicate recognition. |
| Snapshot | document, version, content |
Complete accepted content at one exact boundary. |
| Grant | document, user, role |
Current read/edit permission. |
Scope operation identity to the document and authenticated actor. A retry with identical content returns its original acceptance; the same ID with different content is an error. Check for an already accepted operation before rejecting its now-old base, because a lost acknowledgment does not turn a retry into a new unsupported edit.
The client keeps a pending-operation buffer, ideally persisted locally under a stated draft policy. The server retains enough accepted history for its advertised reconnect interval. Presence messages have session sequence and expiry but do not advance durable document versions. Merely knowing an actor ID or document URL does not grant access.
08Recover the gap between saved history and live messages
On open, return a snapshot and a fixed accepted upper boundary, such as version 22. If the snapshot is version 20, replay operations 21 and 22 to reconstruct that boundary. Newer edits can arrive through a subsequent replay or live subscription; they must not silently alter which operations the initial response promised.
Subscribe with the last applied version. The server must replay or buffer edits accepted between the initial read and live registration. Otherwise version 23 could be lost in the handoff. A client receiving version 24 while expecting 23 requests the gap before applying later positional operations.
After a brief disconnect, first obtain missing accepted history, then reconcile and resubmit pending edits with their original identities using the editing library's protocol. The server returns saved acceptances for duplicates. Arrival order between acknowledgment and broadcast should not cause the sender to apply its own insertion twice.
If a client predates retained transformation history, preserve its local draft and require an explicit resynchronization/merge workflow against current content. A snapshot of today's text does not reconstruct every old position shift. Advancing the supported history boundary is therefore a product decision, not a side effect of deleting old logs to save disk.
09Separate document ordering from sockets and fanout
Move WebSocket connections to gateways when connection count and outgoing traffic overwhelm the first process. Gateways authenticate sessions, enforce buffer limits and forward edits to the document owner. They broadcast accepted versions to subscribers but cannot declare an edit saved just because they received it.
Partition ownership by document. Different documents use independent coordinators and storage partitions, while D7 retains one accepted order. A routing directory chooses the coordinator; the durable append path remains the source of authority. Workers receiving D7’s edits at random would still need to agree on their order and transform them against that same history.
A replacement coordinator reconstructs D7 from durable content before accepting edits. Give ownership an increasing generation and check it at log append alongside the expected head version. If an old coordinator resumes after replacement, its stale generation is rejected by storage. Routing new clients elsewhere is insufficient because the old process may still hold open sockets and prepared writes.
If a transform was computed against head 21 but the append finds head 22, reload the intervening edit and recompute rather than appending the old transformed position blindly. Use the chosen library/adapter's supported concurrency controls and test the ownership boundary rather than implementing a second, inconsistent edit protocol around it.
An edit crosses a gateway to the current document coordinator, which transforms it against accepted history and commits through the fenced append path. Only committed versions return to subscribers. Snapshot workers compact recoverable document state in the database; a replacement coordinator rebuilds from that state before accepting edits.
Read each connection in order
- syncBase version and operation IDEditing clients → WebSocket gateways
- syncFind current ownerWebSocket gateways → Document-owner directory
- syncForward authorized editWebSocket gateways → Document coordinators
- syncFenced append and new headDocument coordinators → Operation / snapshot DB
- asyncCommitted accepted versionDocument coordinators → WebSocket gateways
- asyncAcknowledge and broadcastWebSocket gateways → Editing clients
- syncRead history / save snapshotSnapshot workers → Operation / snapshot DB
10Keep recovery short without hiding snapshot assumptions
Create snapshots from a precise committed version so the saved content and advertised version agree. A simple baseline stores snapshot content and its version transactionally in the same database. The 1 MB document bound makes this a reasonable starting choice, avoiding a separate object-publication protocol in the first design.
After a snapshot at version 22, recovery starts there and replays operations after 22. Do not delete transformation history still required for supported pending edits merely because the visible text has a newer snapshot. A snapshot may restore the saved text without retaining the earlier edits needed to transform a reconnecting client’s pending work.
At larger historical volume, immutable snapshot objects plus database references can reduce database storage, but publishing, retaining and deleting those objects safely adds another protocol. Keep that extension explicit rather than compressing object lifetime guarantees into “use object storage.”
Preserve accepted operation and metadata transactions across application restarts; configure and test database replication for any additionally agreed storage-node or zone failure tolerance. Presence, gateway connections and in-memory content caches can be rebuilt; the log and operation identities cannot be treated as disposable. Backups cover accidental deletion or corruption replicated to every serving copy. A restore test should reconstruct exact accepted content and retry behavior, not merely find a snapshot file with the right name.
11Preserve drafts, permissions and accepted history
| Failure | Behavior |
|---|---|
| Append commits, acknowledgment disappears | Retry the same operation and return its accepted version without inserting again. |
| Gateway disconnects | Client reconnects with last applied version; durable edits remain available for replay. |
| Document authority loses safe write access | Local typing can remain pending, but save acceptance pauses. |
| Slow receiver accumulates output | Bound its queue and require replay-based reconnect; coalesce or discard old cursor updates. |
| Edit permission is revoked | Subsequent authorized append decisions reject edits; preserve rejected local text as a draft. |
Check permission at edit acceptance, not only at socket establishment. Read permission also applies to snapshots and replay. Revocation cannot erase bytes a user already downloaded, so distinguish preventing new access from recalling past disclosure.
Test concurrent insert/delete/undo histories under different delivery orders and compare final accepted content. Add coordinator failure, duplicate operation IDs, lost acknowledgments and old-client reconnect. Before upgrading the editing library, replay representative histories through old and new versions and verify protocol compatibility.
Measure pending age, accepted-edit latency, remote-delivery delay, transform failures, replay bytes and buffer evictions. A responsive local cursor can hide a saving outage. Keep document text out of broadly accessible logs; IDs, versions and timing usually suffice for diagnostics.
12Check the design against its requirements
Use the numbered requirements to check the final design. FR refers to the functional list; NFR refers to the non-functional list. Performance rows specify tests still required, not achieved benchmark results.
| Requirement | Design mechanism | Verification and remaining limit |
|---|---|---|
| FR1–2; NFR3,5: authorized concurrent editing | Current document grants, a proven OT library and one accepted order per document. | Run concurrent insert/delete/undo histories and revoke permission during a session. Clients converge to the accepted result without unauthorized new edits. |
| FR3; NFR1–2: responsive collaboration | Separate gateway fanout from document transformation; render pending locally and acknowledge only after commit. | Load-test fleet traffic and the 100-participant hot document against both p95 targets. A fast local cursor cannot substitute for measured save latency. |
| FR4; NFR4: reconnect without duplication | Snapshot plus versioned replay, a gap-free subscription handoff and retained operation identities. | Drop A17’s acknowledgment and reconnect inside thirty minutes. After the advertised boundary, preserve the draft and require explicit resynchronization rather than silently losing text. |
| NFR3–5: recover document authority | Storage validates coordinator generation and expected document head; reconstruct from committed history. | Pause an old coordinator through replacement and crash after append. Its stale write must fail and the acknowledged edit must replay once; test the configured database durability separately. |
13Rapid revision
Remember: The merge algorithm combines edits; the durable log preserves them. A socket does neither by itself.
| Concern | Complete design |
|---|---|
| Concurrent edits | Send each edit with its own ID, rather than replacing the whole document. |
| Merge model | Use proven operational transformation (OT) to adjust edits for changes the server accepted in between. |
| Local responsiveness | Render pending edits immediately and retain their identities for recovery. |
| Saved state | Acknowledge only after the accepted operation is durable. |
| Retry | Return the saved result for an already accepted operation; never insert it twice. |
| Reconnect | Load a snapshot and later edits; join live updates without gaps, then resolve the client’s pending edits. |
| Scaling | Add gateways for connections; each document’s coordinator maintains that document’s edit order. |
| Failover | When appending, storage checks the current owner token and expected latest document version. |
| History | State how long automatic reconnect works; preserve drafts when the edit history needed to merge them is gone. |
Close by showing cat → cXat → cXYat, then lose A17's acknowledgment. Explain which component combines edits and which preserves the result. A CRDT is an alternative worth evaluating for extensive offline collaboration, with its own operation metadata and cleanup requirements; it changes the editing model rather than simply adding a component to OT.
Practise the interview questions
Say your answer aloud before opening the model answer. Then answer the follow-up and compare the reasoning.
Why not save the whole document after each edit?
Reveal a model answer
Two users can replace the same base with different complete copies, causing the later save to erase independent work. Operations preserve the intent the merge algorithm needs.
Interviewer follow-up
Does a row lock recover the overwritten edit?
Reveal the follow-up answer
It orders replacements but does not reconstruct the edit lost from the later copy.
What the answer must demonstrate: Explain lost independent edits under whole-document replacement.
How do X and Y inserted at position 1 both survive?
Reveal a model answer
A defined tie-break accepts one first and transforms the other position around it. Both clients reconcile pending and accepted edits under the same rules.
Interviewer follow-up
Does the example implement a full editor?
Reveal the follow-up answer
No. Deletes, undo, Unicode and overlapping operations require a proven complete algorithm.
What the answer must demonstrate: Trace both client reconciliation and deterministic server transformation without generalizing an insert demo.
When can the interface show saved?
Reveal a model answer
After the authoritative operation log durably accepts the edit, not merely after local rendering or gateway receipt.
Interviewer follow-up
Must every peer acknowledge first?
Reveal the follow-up answer
No. One offline peer must not block everyone else from saving.
What the answer must demonstrate: Identify durable acceptance separately from local display and peer delivery.
A17 committed but its reply vanished. How is a duplicate X avoided?
Reveal a model answer
Retry A17 with the same authenticated identity and payload. The owner returns the existing accepted version, and the client removes its pending entry.
Interviewer follow-up
What if its base is now old?
Reveal the follow-up answer
Check the stored duplicate outcome before treating it as a new edit requiring old transformation history.
What the answer must demonstrate: Reuse operation identity and remove the matching pending edit without applying it twice.
How can an edit disappear between opening and subscribing?
Reveal a model answer
An operation may commit after the initial snapshot/read but before live subscription. Replay or buffer from the last applied version to cover that gap.
Interviewer follow-up
What if version 24 arrives before 23?
Reveal the follow-up answer
Fetch the missing history before applying later positional edits.
What the answer must demonstrate: Close the snapshot-to-subscription gap and repair missing versions before positional application.
Why must storage reject a stale document coordinator?
Reveal a model answer
The old process may resume with open sockets after a replacement takes over. Checking its ownership generation at append prevents a second accepted history.
Interviewer follow-up
Why check the head too?
Reveal the follow-up answer
A transform computed against old content must be recomputed after intervening operations.
What the answer must demonstrate: Check ownership generation and expected head at the protected append boundary.
Can a snapshot replace all operation history immediately?
Reveal a model answer
It can reconstruct current text but may not preserve the context needed to transform supported old pending edits. Retention follows the reconnect contract.
Interviewer follow-up
What happens beyond that contract?
Reveal the follow-up answer
Preserve the local draft and require an explicit current-snapshot merge workflow.
What the answer must demonstrate: Retain required transform context or explicitly preserve drafts for resynchronization.
Why treat cursors differently from document edits?
Reveal a model answer
Cursors are ephemeral hints that may expire or be coalesced. Accepted edits are durable content requiring replay and uniqueness.
Interviewer follow-up
Does revocation remove already downloaded text?
Reveal the follow-up answer
No. It prevents subsequent authorized access and acceptance, not recall of existing copies.
What the answer must demonstrate: Separate ephemeral presence, durable edits and the limits of access revocation.
Blank-page exercise · 45 minutes
Build the answer yourself
Design an online editor where two users insert into cat concurrently, then lose one acknowledgment and reconnect through a new server.
- Agree numbered functional and non-functional requirements, including text model, save latency and the automatic reconnect window. Then explain why identified operations preserve concurrent edits better than whole-document replacement.
- Trace pending edits, accepted versions and reconnect without applying an edit twice.
- Scale documents and connections separately while preserving document authority and permissions.
- Trace a timeout and a concurrent request using the actual durable records.
- Review the final architecture against every numbered requirement, including measured bottlenecks, targets still needing validation and remaining failure limits.
Check that each component and design decision follows from your requirements and workload.
Recall the key ideas
Answer from memory before opening each card. Explain why the choice works and what it costs. Revisit missed cards tomorrow.
Design a collaborative text editorA and B insert X and Y at position 1 in cat. If X is accepted first, where does Y go?Recall first, then reveal
The OT rule moves Y to position 2, yielding cXYat. Both clients apply the same rule to accepted and pending edits.
An earlier insert shifts later positions: transform the edit, not the whole document.
Return to lessonDesign a collaborative text editorHow can a client recover an accepted edit whose broadcast was lost?Recall first, then reveal
Read the saved operation log; operation IDs let the client recognize edits it has already seen.
A lost broadcast is recoverable because the accepted edit is saved.
Return to lessonDesign a collaborative text editorWhy might a current snapshot be insufficient to merge an old pending edit?Recall first, then reveal
The transformation algorithm may still need the intervening edit history to adjust that pending edit.
Current text is not all history.
Return to lessonFinal revision
Summary and interview notes
Show typing immediately, combine concurrent text edits using proven operational transformation, and save accepted operations durably. On reconnect, recover missed edits and resolve pending ones without inserting them twice.
Remember these points
- Represent edits as identified operations.
- Keep pending display distinct from durable save.
- Recover missing versions and deduplicate retries.
- Scale connection gateways separately while keeping one accepted edit order per document.
Interview tips
- Separate local pending text, saved edits and peer delivery.
- Work through cat → cXat → cXYat, then lose A17’s acknowledgment and reconnect.
Important qualifications
- Traffic and latency figures are interview assumptions, not claims about a named company's deployment.
Continue after the core interview
Explore the advanced version
The advanced lesson keeps the full detailed design. Use these sections when you want to examine the stronger requirements and failure cases.
- Detailed merge and coordinator proof
Study expected-head checks and stale-owner interleavings after the basic operation model.
- Object-backed snapshot publication
Large snapshot storage introduces separate publication, reader and reclamation rules.
- Long-offline collaboration with CRDTs
Months of disconnected editing may justify a different merge model and metadata contract.
- Hot-document limits and alternatives
Splitting one document requires independently mergeable semantics, not arbitrary operation hashing.
Technical references
- Yjs shared typesOfficial examples of collaborative shared data types and transactions.
- Yjs document updatesDocuments update exchange, state vectors, and merge behavior for a concrete CRDT implementation.
- RFC 6455: WebSocketDefines the bidirectional transport used for interactive edit and presence events.
- ShareDB documentationOfficial operational-transformation backend documentation; evaluate the supported text type and persistence adapter rather than inferring custom authority guarantees.
- CRDT definitions and glossaryStandard convergence property for conflict-free replicated data types; sequence text is one application, not the general definition.
Practice marks stay in this browser.