System designby Learnastra

System-design interview · Core interviews

Design a file synchronization service

By Anup Rai

Build file synchronization around immutable revisions, explicit conflicts and durable device catch-up, then add chunk reuse and notification gateways.

You will learn to

  • Explain a complete save and catch-up without losing offline edits.
  • Connect stable identity, expected revisions, manifests and cursors.
  • Justify chunking and workspace partitioning while preserving recovery.

Practice in this chapter

8 interview questions with model answers and follow-ups.

Go to interview practice

Useful foundations: Storage engines and data models · Transaction isolation · Message queues, event logs, delivery guarantees, and backpressure

Workload and timing examples are interview assumptions.

01Synchronize saved revisions without losing offline edits

Two laptops share budget.xlsx. Both start from revision 12, then edit while one is offline. The service must eventually distribute saved work without allowing the last upload to erase a change its author never saw. File synchronization therefore needs more than byte copying: it needs stable file identity, saved revisions, change discovery and a conflict policy.

Support files and folders, uploads, downloads, rename, deletion, shared workspaces, offline edits and retained versions. A workspace is a shared set of files and membership rules. Keep its metadata changes within one transaction domain for this interview. A stable file ID survives a rename; the pathname tells us where that file currently appears.

Use conflict copies for concurrent binary edits. Automatically merging arbitrary spreadsheets or images is outside scope. Distinguish saved locally, upload pending, committed on the server and applied on another device. A successful server commit does not mean an offline laptop has already received the file. Limit files to 1 GiB and declare a history-retention window; indefinitely disconnected clients cannot depend on a log kept for only a month.

Clarify whether concurrent offline edits may overwrite each other, how long history must remain recoverable, and what delay is acceptable for an online second device. This design preserves conflicts and separates a change notification from completion of byte transfer.

02Functional requirements

Agree on these supported actions before selecting components.

  1. Manage a shared file namespace. Create files and folders, rename and delete them, and authorize access through workspace membership. File identity remains stable through a rename.

  2. Synchronize across devices. Upload and download saved revisions, let offline devices catch up, and show local, pending, server-committed and locally-applied states distinctly.

  3. Preserve concurrent edits. Retain both users’ work when an offline edit conflicts with a newer saved version, and make the conflict visible instead of silently overwriting either edit.

  4. Recover and restore. Resume interrupted operations, retrieve retained versions and resynchronize safely when a device cursor is older than the change history.

03Non-functional requirements

Use these as illustrative interview assumptions to agree with the interviewer. Numerical targets require measurement; they are not claims about an existing product or a proven implementation. p95 (the 95th percentile) means 95% of measured operations finish within the stated time.

  1. Workload and file limits. Use 100 billion current files, approximately 10 PB of current logical bytes and a peak of about 28,935 metadata commits/s. Limit each file to 1 GiB; retained revisions add capacity beyond those current-file totals.

  2. Metadata latency and discovery. Target regional p95 metadata responses within 200 ms and p95 committed-change discovery within five seconds for connected devices during normal operation. Completing a large download depends on file size and bandwidth and is not covered by the five-second target.

  3. Revision integrity and durability. A committed revision must name complete protected chunks and survive a process restart or one storage-node failure within the region. Concurrent commits cannot silently lose either author’s accepted work.

  4. Bounded retained history. Choose 30 days of prior revisions and change-log history for this exercise. Devices beyond that history window must preserve local pending edits and recover from a consistent snapshot instead of skipping directly to the newest cursor.

  5. Workspace isolation. Authorize upload, commit, revision reads and downloads. Chunk reuse must not disclose or grant access to another workspace’s private content.

  6. Recovery under overload. Bound transfer concurrency and spread reconnect retries. If metadata cannot safely commit, clients retain pending local work and must not report synchronized success.

04Commit a whole file, then let another device catch up

Start with one metadata database, durable file storage and a client that periodically asks for changes. Laptop A uploads an immutable copy of its saved file. After verifying the bytes, the server checks that A edited revision 12 and the current file is still revision 12. In one metadata transaction it records revision 13, makes it current and appends a change entry.

Laptop B asks for changes after its last saved position. It learns that file f42 now has revision 13, downloads that immutable version to a temporary location, verifies it and replaces its local copy. It records progress so a restart can resume. If B has unsent edits, it preserves them instead of blindly replacing the file.

The local client must capture a stable byte version before uploading. Reading a live file while another application rewrites it can mix two saves. A staging copy or suitable filesystem snapshot makes the uploaded bytes correspond to one captured revision.

This whole-file baseline is slower for large repeated edits, but it already establishes the essential rules. Uploading bytes does not make them current; the metadata commit does. Polling may delay discovery, but missed network notifications cannot lose a committed revision.

Design diagramWhole-file synchronization before chunk optimization

Bytes are verified before metadata commit; devices use saved history to discover the new revision.

Whole-file synchronization before chunk optimizationBytes are verified before metadata commit; devices use saved history to discover the new revision. a to api: Upload / commit / catch up; api to bytes: Store or retrieve revision bytes; api to meta: Check base and commit change; api to a: Revision outcome or changesUpload / commit / catch upStore or retrieve revision bytesCheck base and commitchangeRevision outcome or changesACTOREditing deviceSERVICESync applicationSTOREFiles, revisions andchangesSTOREImmutable filestoragesyncmedia
Read each connection in order
  1. syncUpload / commit / catch upEditing device → Sync application
  2. mediaStore or retrieve revision bytesSync application → Immutable file storage
  3. syncCheck base and commit changeSync application → Files, revisions and changes
  4. syncRevision outcome or changesSync application → Editing device

05Count changed bytes, metadata and online devices separately

Assume 500 million accounts, 200 files per account and 100 KB average file size. That is 100 billion current files and 10 PB of logical current content. At an illustrative 1 KB of file metadata, metadata alone reaches 100 TB before indexes and replicas. Old revisions add storage beyond the current-file total.

For 100 million daily active users making five committed changes each, average commits are 500M / 86,400 ≈ 5,787/s. A fivefold peak is about 28,935/s. If a change uploads 200 KB on average, mean changed-byte ingress is approximately 1.16 GB/s. Downloads depend on how many other devices actually receive the change.

Workload Estimate Design implication
Maximum 4 MiB chunks per file 1 GiB / 4 MiB = 256 Manifest checks can be bounded
Ten million devices polling each minute About 166,667 polls/s Many requests can find no changes
One million reconnects in one minute About 16,667 sessions/s Recovery needs admission and randomized retry

The byte estimates motivate independent object storage; the metadata estimate motivates partitioning. Connection work motivates notification hints. These are different pressures. Do not claim a fixed deduplication saving: small files may change entirely, and compressed formats can change many bytes after a seemingly small edit.

06Name the base revision and the recovery cursor

Begin an upload with the file identity, expected base revision and a client-generated request identity.

Start-upload information for the running example

File ID:                f42
Expected base revision: 12
Request ID:             edit-77

A retry keeps edit-77 and the same payload. The server returns an upload session; completing byte transfer is distinct from committing a new current revision.

Operation Contract
Start upload for f42, base 12 Reserve one resumable attempt
Upload content or chunks Store verified immutable bytes
Commit edit-77 Check permission and expected revision; return saved outcome
Get changes after cursor 880 Return ordered committed changes
Get revision 13 Return its manifest and authorized download route
Rename or delete f42 Update its stable identity's metadata and change history

A conflict returns the preserved conflicting revision or conflict-file identity. Repeating the same operation returns that outcome, rather than creating another conflict copy. A restore creates a new revision from retained history; it does not rewrite the past.

A cursor identifies how far a device has applied committed changes. If its position is older than retained history, return a clear resynchronization response. The client obtains a consistent file snapshot and its matching log position, preserves local pending work, and resumes after that boundary. Jumping straight to the newest cursor would miss deletions and renames.

07Separate identity, revision and byte layout

The metadata records separate the current directory entry, saved content and the history other devices must apply.

Record Fields Purpose
File workspaceId, fileId, parentId, name, currentRevision Describes a directory entry and its current version.
Revision fileId, revision, manifest, author, createdAt Identifies saved content; the manifest lists immutable chunks in order.
Change workspaceId, sequence, fileId, action, revision Records what other devices must learn.

Membership and request-result records participate in the relevant metadata decisions.

Folder browsing needs (workspaceId, parentId, name). Device catch-up needs (workspaceId, sequence). File lookup uses stable file ID. A name-based shard key would complicate renames; hashing every file independently would complicate workspace-wide ordering and atomic folder changes. Partition independent workspaces first so their local checks remain understandable.

A delete leaves a tombstone: a retained deletion record that an offline device can observe. Merely removing the row would make absence indistinguishable from a change the client never learned about. Retained old revisions still need their chunks, so cleanup must check more than the current revision.

The client has its own durable local journal for pending uploads, accepted revisions and in-progress downloads. It is part of the system, not an optional cache. A restart must not forget work just because the server is reliable.

08Preserve both edits when the expected revision changed

Laptop A and laptop B both edit revision 12. B's metadata transaction checks currentRevision=12 and commits revision 13. A arrives later with the same expected base. Its check now fails because currentRevision is 13. The server preserves A's content as a conflict copy and records that outcome without replacing B's accepted head.

The compare and update must happen atomically. A separate read followed by an unconditional write allows both clients to pass the check and overwrite each other. Client timestamps do not solve the problem: clocks differ, and being later does not establish an intention to discard an unseen edit.

A conflict copy is a real saved object. Its bytes, metadata and change event must survive just like a normal revision, so other devices can discover it and the user can resolve it. Recheck write permission at publication; a long-running upload must not bypass a membership revocation by calling its output a conflict.

Within one workspace, commit each revision update and its change-log position together, in order. A device must not advance beyond an earlier change that can still commit later. This is why a counter allocated before commit is not automatically a safe synchronization cursor.

Request traceTwo offline edits preserve both versions

The second commit observes a different current revision and records a conflict instead of replacing unseen work.

Two offline edits preserve both versionsThe second commit observes a different current revision and records a conflict instead of replacing unseen work. a to server: Read revision 12; b to server: Read revision 12; b to server: Commit edit based on 12; server to b: Revision 13 committed; a to server: Commit different edit based on 12; server to a: Current is 13: saved conflict copyPARTICIPANTLaptop APARTICIPANTLaptop BPARTICIPANTWorkspacemetadata1. Read revision 122. Read revision 123. Commit edit based on 124. Revision 13 committed5. Commit different edit based on 126. Current is 13: saved conflict copysyncreturn
Read each connection in order
  1. syncRead revision 12Laptop A → Workspace metadata
  2. syncRead revision 12Laptop B → Workspace metadata
  3. syncCommit edit based on 12Laptop B → Workspace metadata
  4. returnRevision 13 committedWorkspace metadata → Laptop B
  5. syncCommit different edit based on 12Laptop A → Workspace metadata
  6. returnCurrent is 13: saved conflict copyWorkspace metadata → Laptop A

09Transfer changed chunks after the basic revision rule works

Suppose a 9 MiB file is split into 4 MiB, 4 MiB and 1 MiB chunks. Only the middle chunk changes.

Revision Ordered chunk references Bytes uploaded for this version
12 A, B, C The initial 9 MiB file.
13 A, D, C Only the new 4 MiB chunk D; reuse A and C.

This transfers 4 MiB instead of 9 MiB for that upload, saving about 56% in this particular example.

Chunk identities include a verified content checksum and size. Scope reuse to content the workspace is authorized to access; a global “does this hash exist?” endpoint can reveal another customer's private content. The manifest is committed only after every referenced chunk exists and is protected from deletion.

Protect chunks needed by an active upload, then transfer that protection to saved-revision references when commit succeeds. Cleanup first marks an unreferenced chunk as unavailable for future publication before deleting its bytes. A stale scan finding no reference is insufficient if an upload can start using the chunk afterward.

Fixed-size chunks are straightforward and provide bounded retry units. An insertion near the start may shift later boundaries, reducing reuse. Content-defined chunking chooses boundaries from content patterns and can improve that workload, at additional CPU and implementation cost. Patches inside changed chunks are another later optimization; neither is necessary to explain the first correct sync service.

10Use notifications to prompt durable catch-up

Millions of periodic empty polls create work even when nothing changes. A notification gateway can keep a persistent connection or long poll and send a small “workspace changed” hint. The device still reads the durable change log using its cursor. A lost hint delays discovery until reconnect or another poll; it does not erase history.

Separate bulk byte transfer from small metadata operations, with independent concurrency and bandwidth limits. A slow 1 GiB upload should not prevent another user from renaming a folder. Cache popular chunks and manifests only when measured reuse justifies their memory cost; a large chunk displaces far more cache space than one metadata row.

Distribute workspaces across replicated metadata groups. This scales independent workspaces while keeping their conflict checks, directory changes and change history within one transaction boundary. One unusually large workspace remains a potential bottleneck; splitting it requires a new ordering and transaction design, not simply a different hash function.

Gateways do not own revisions. They can be replaced and clients reconnect from durable cursors. Spread reconnects with randomized delays and limit concurrent snapshot recovery and downloads during a fleet restart. More connections and faster notifications improve responsiveness only if metadata and byte services can handle the recovery traffic.

Implement the one-node-loss requirement with durable majority commits across three metadata replicas per workspace group in independent regional failure domains, plus chunk storage whose acknowledged writes survive one storage-node loss. Retain the selected 30-day revision/change window and benchmark metadata and change-discovery targets independently of bulk-transfer duration.

Design diagramHints wake devices; durable metadata tells them what changed

A device uploads missing chunks. The metadata API verifies and protects every manifest chunk before committing against the expected revision. Another device receives a change hint and uses its saved cursor to read committed history. Independent transfer workers move authorized bytes; workspace metadata groups retain the conflict and ordering decisions.

Hints wake devices; durable metadata tells them what changedA device uploads missing chunks. The metadata API verifies and protects every manifest chunk before committing against the expected revision. Another device receives a change hint and uses its saved cursor to read committed history. Independent transfer workers move authorized bytes; workspace metadata groups retain the conflict and ordering decisions. device to api: Commit / read changes / get manifest; api to meta: Workspace-local transactions; device to bytes: Authorized chunk upload / download; bytes to objects: Store or fetch verified chunks; meta to gateway: Relay committed-change hints; gateway to device: Workspace changed; fetch history; api to objects: Verify and protect manifest chunksCommit / read changes / getmanifestWorkspace-local transactionsAuthorized chunk upload /downloadStore or fetch verified chunksRelay committed-change hintsWorkspace changed; fetchhistoryVerify and protect manifestchunksACTORSync clients + localjournalsSERVICEMetadata APISTOREReplicatedworkspace groupsSERVICEBulk transfer workersSTOREImmutable chunkstorageSERVICENotificationgatewayssyncmediaasync
Read each connection in order
  1. syncCommit / read changes / get manifestSync clients + local journals → Metadata API
  2. syncWorkspace-local transactionsMetadata API → Replicated workspace groups
  3. mediaAuthorized chunk upload / downloadSync clients + local journals → Bulk transfer workers
  4. mediaStore or fetch verified chunksBulk transfer workers → Immutable chunk storage
  5. asyncRelay committed-change hintsReplicated workspace groups → Notification gateways
  6. asyncWorkspace changed; fetch historyNotification gateways → Sync clients + local journals
  7. syncVerify and protect manifest chunksMetadata API → Immutable chunk storage

11Recover local and server interruptions separately

If a client crashes after uploading chunks but before commit, its local journal and server upload session identify the pending operation. A retry may finish it while the session remains valid. If commit succeeded but its response was lost, edit-77 retrieves the original result and must not create revision 14 accidentally.

A receiving device may crash after replacing the local file but before recording its new cursor. Its journal must recognize the completed replacement and finish the metadata update on restart. Conversely, advancing the cursor before retaining the bytes or a durable recovery plan can permanently skip a change. The local filesystem and the client's metadata database do not automatically share a transaction.

When the server cannot safely accept metadata writes, the client may continue editing locally with pending status. It should not claim synchronization. After recovery, the actual base revision determines whether a conflict exists. An already downloaded file cannot be made secret again by server-side revocation.

Backups must restore metadata, retained revisions and their referenced chunks together. Test an old-version restore after newer revisions and deletion. A healthy current file does not prove that historical chunks still exist.

12Measure synchronized work, not just successful uploads

Measure metadata commit latency, time for connected devices to catch up, conflict rate, transferred bytes per revision and time spent pending. Missing chunks in an accepted revision are an integrity incident; a delayed notification is a different problem. Track old upload sessions and cleanup backlog so temporary bytes do not grow without bound.

Authorize upload, final commit and downloads. Scope transfer credentials to the intended object and operation. Keep filenames and file contents out of broad telemetry. Count retained revisions toward quotas, otherwise repeated overwrites can consume large historical storage while the visible current file stays small.

Cost combines current bytes, version history, replicas, transfers, chunk-reference metadata and client/server CPU. Chunking is worthwhile when saved bytes exceed its lookup and bookkeeping cost. Test a tiny-file workload as well as large files with localized edits.

Exercise concurrent offline edits, a rename during synchronization, an expired history cursor, cleanup racing an upload and mass reconnection. These scenarios test the user's promise that work survives and converges, rather than merely demonstrating that files can cross the network.

13Check the design against the requirements

Use the agreed lists to check the finished design. The tests below still need to establish the targets; a proposed mechanism is not a measured result. FR refers to the numbered functional requirements above; NFR refers to the numbered non-functional requirements.

Requirement Design mechanism Validation and remaining limit
FR1–3: namespace and concurrent edits Stable file IDs, workspace-local transactions and expected-revision checks. Race two edits based on revision 12 and rename a file; retain both edits without changing identity.
FR2 + NFR2: online synchronization Durable changes and cursors, with notification hints. Drop hints and reconnect a device; benchmark discovery separately from metadata latency and whole-file download time.
FR4 + NFR4: old devices and restore Thirty-day history plus consistent snapshot/log-position recovery and local journals. Reconnect beyond retention with unsent edits and restore an old version. Count retained revision bytes in sizing.
NFR3: committed bytes survive Protect chunks before manifest commit; replicated metadata and byte storage. Fail one storage node after a commit; no accepted manifest may reference lost or prematurely collected chunks.
NFR5–6: safe shared use Workspace-scoped authorization, bounded transfer work and pending status during unsafe writes. Revoke membership mid-upload and run a reconnect wave. A fast hint cannot substitute for authorized durable catch-up.

14Rapid revision

Remember: Compare the saved base before replacing the current revision; an unseen edit must become a conflict, not disappear.

Concept Purpose Mistake to avoid
Stable file ID Recognize the file after a rename Treating the path as immutable identity
Expected base revision Detect an unseen concurrent edit Last upload wins regardless of intent
Immutable revision manifest List chunks needed to reconstruct the revision Publishing before verifying chunks and protecting them from cleanup
Chunk reuse Transfer less when content repeats Claiming every format benefits equally
Change log and device cursor Recover missed updates and deletions Treating a socket event as durable progress
Local journal Resume interrupted local work Saving progress before local changes survive a crash
Workspace partition Transact related workspace metadata together Splitting records without a way to commit related changes

A strong summary follows revision 12 into two offline edits, explains which metadata transaction wins and where the other edit is saved, then shows how a second device catches up. Chunking and notifications are efficiency improvements around that model. Automatic document merging and extremely large shared workspaces require further semantics; they are not solved merely by adding more transfer servers.

Practise the interview questions

Say your answer aloud before opening the model answer. Then answer the follow-up and compare the reasoning.

Foundation · Question 1

Why separate file ID from path?

Reveal a model answer

A rename changes the directory entry, not which file and history the devices are tracking.

What the answer must demonstrate: Distinguishes stable identity from mutable directory placement.

Applied · Question 2

What happens when two devices edit revision 12?

Reveal a model answer

The first valid commit changes the head. The second expected-revision check fails and preserves its work as a conflict.

What the answer must demonstrate: Uses an atomic expected-version check and preserves the losing edit.

Applied · Question 3

What does chunking save in the 9 MiB example?

Reveal a model answer

Changing only the middle 4 MiB allows the other 5 MiB to be reused.

What the answer must demonstrate: Calculates changed bytes without promising universal deduplication.

Applied · Question 4

Why is a notification not enough for synchronization?

Reveal a model answer

Notifications can disappear while a device is offline. A durable change log and cursor recover what was missed.

What the answer must demonstrate: Treats cursor recovery as durable state rather than transport behavior.

Applied · Question 5

When does an upload become the current file?

Reveal a model answer

The server verifies that chunks exist and cannot be deleted, checks current permission and the expected base revision, then commits the new revision in metadata.

What the answer must demonstrate: Separates byte upload from publication and request replay.

Applied · Question 6

Why keep old-revision references during cleanup?

Reveal a model answer

Historical versions still need their chunks for restoration. Checking only which chunks the current revision uses would lose older versions.

What the answer must demonstrate: Counts retained history and active-upload protection.

Applied · Question 7

What can fail on the receiving device?

Reveal a model answer

A crash can occur between file replacement and local cursor persistence, so a recovery journal must connect them.

What the answer must demonstrate: Recognizes the local filesystem/metadata crash boundary.

Follow-up · Question 8

Can the service automatically merge spreadsheets?

Reveal a model answer

Only with format-specific semantics and a defined conflict policy; a generic byte synchronizer cannot infer user intent.

What the answer must demonstrate: Avoids inventing merge semantics for arbitrary binary formats.

Blank-page exercise · 45 minutes

Build the answer yourself

Design shared file synchronization for multiple devices, offline edits and retained history. Use a 9 MiB file edited concurrently from revision 12 as the running example.

  • Use 5 minutes to agree numbered functional and non-functional requirements for offline edits, conflict copies, sharing, 30-day history, discovery latency and durability.
  • Trace whole-file save and catch-up in 8 minutes.
  • Estimate bytes, metadata and reconnect load in 7 minutes.
  • Define revisions, manifests, APIs and cursors in 10 minutes.
  • Add chunking and explain conflict, cleanup and crash recovery in 10 minutes.
  • Use 5 minutes to review the final design against the numbered FR/NFR lists, including expired cursors, protected chunks, one-node loss and the limit on download latency.

Check that each component and design decision follows from your requirements and workload.

Recall the key ideas

Answer from memory before opening each card. Explain why the choice works and what it costs. Revisit missed cards tomorrow.

Design a file synchronization serviceBoth laptops edit revision 12; one commits revision 13. What happens when the other uploads its edit?Recall first, then reveal

The atomic expected-revision check fails. Save the second edit as a conflict copy and record that outcome so retrying does not create another copy.

Compare before replacing.

Return to lesson
Design a file synchronization serviceA device misses a socket notification. How does it find the update later?Recall first, then reveal

Read the committed change log after the device’s saved cursor. The notification only prompts an earlier check.

Hints wake; history recovers.

Return to lesson
Design a file synchronization serviceWhen may cleanup delete a stored chunk?Recall first, then reveal

When no retained revision references it and no still-valid upload can publish a revision that uses it.

No references, no publication right.

Return to lesson

Final revision

Summary and interview notes

Compare an edit’s base revision with the current revision to save offline work as a new revision or conflict copy. Reuse chunks to reduce transfers; replay logged changes to recover missed updates.

Remember these points

  • Stable IDs survive path changes.
  • Expected revisions detect unseen edits.
  • Committed changes support every device’s recovery.
  • Historical versions and pending uploads protect chunks.

Interview tips

  • Show both devices’ base revisions before discussing conflicts.
  • Separate server commit from other-device completion.

Important qualifications

  • Fine-grained concurrent folder operations and split-workspace authority need more detailed protocols.

Continue after the core interview

Explore the advanced version

The advanced lesson keeps the full detailed design. Use these sections when you want to examine the stronger requirements and failure cases.

Technical references

Practice marks stay in this browser.