Concept lesson · Foundations
Multi-region architecture and disaster recovery
Start here
Definition
Multi-region architecture deploys a service across geographically separate regions. Disaster recovery is the planned restoration of usable service and data after a major disruption; the recovery point objective (RPO) specifies the targeted data-loss window and the recovery time objective (RTO) specifies the targeted restoration time.
Why it matters: A regional outage, accidental deletion, or failed dependency can affect every local replica. Recovery requires knowing which saved changes survived, ensuring only the designated replacement can accept writes, and providing enough capacity to serve users.
Compare two intervals: how far the recovered data lags behind the disruption, and how long users wait for service to return.
Read the diagram step by step
- West has O16 at 12:00:00. East acknowledges O17 at 12:00:04 but it has not replicated. Connectivity fails at 12:00:05.
- The safe replica is five seconds behind disruption, and O17 may be lost despite its acknowledgement.
- Service is validated at 12:07:05: measured recovery takes seven minutes. RPO and RTO are objectives against these separate kinds of loss.
- Fence the old writer before promotion, and reconcile a possibly successful external payment before retrying it.
Worked example
East acknowledges O17 at 12:00:04, fails at 12:00:05, and West has only data through 12:00:00. Restoring West can miss O17; restoring service at 12:07:05 takes 7 minutes.
Key takeaways
You will learn to
- Distinguish high availability from disaster recovery.
- Calculate recovery objectives from a concrete outage.
- Explain safe promotion, routing, and failback.
Practice in this chapter
8 interview questions with model answers and follow-ups.
Go to interview practiceUseful foundations: Replication and durability · Quorums, consensus, leases, and fencing
Workload and timing examples are interview assumptions.
Dotted concept links open the relevant explanation in a new tab.
01Multi-region architecture, high availability, and disaster recovery
Multi-region architecture runs a service across geographically separate deployment regions. Disaster recovery is the planned restoration of usable service and data after a major disruption. A region is a geographical deployment area whose infrastructure can share risks such as a regional network failure; an availability zone is a separate failure domain within a region under the provider’s isolation model. Putting servers in two locations does not provide regional recovery if both still depend on the same regional database, credential service, or network.
High availability keeps the service operating through expected component failures. Disaster recovery restores a useful service after a larger disruption. Backups preserve earlier recoverable states. These capabilities overlap, but a replica that immediately copies an accidental deletion is not a substitute for a backup that can restore yesterday's data.
Specify allowed data loss and recovery time before choosing a regional topology. An East-primary/West-asynchronous-replica example illustrates the tradeoff: an acknowledged order O17 can be absent from West when East fails. Whether that loss is acceptable, and whether writes may pause during recovery, determines the required coordination and cost.
02Recovery point objective (RPO) and recovery time objective (RTO)
The recovery point objective, RPO, is the target maximum amount of data loss measured as a time window. An RPO of 30 seconds means the recovery plan targets a recoverable state no more than 30 seconds behind the disruption. The recovery time objective, RTO, is the target time to restore the agreed service after disruption. Neither is a guarantee merely because it appears in a diagram.
The timestamps are concrete; distances on this timeline are not to scale. The gaps shown are actual outcomes, to compare with the objectives.
Remember: Look backward for the recovery point; forward for service recovery.
Read the diagram
- Measure the history gap before the disruption and the recovery duration after it.
- Recoverable data stops at 11:59:40; disruption is at 12:00:00: a 20-second gap.
- Service is usable at 12:05:00: five minutes of recovery.
Try from memoryWhich gap would a 10-second RPO fail to meet?
The 20-second history gap from 11:59:40 to 12:00:00. The five-minute service recovery is compared with RTO instead.
| Objective | Example | What must support it |
|---|---|---|
| RPO | At most 30 seconds of accepted changes lost | Replication or recoverable logs within that bound, plus measurement |
| RTO | Ordering restored within 10 minutes | Detection, safe promotion, routing, capacity, and validation within that budget |
| Restore correctness | Existing payments reconciled | Durable external IDs and recovery procedures |
03Active-passive, active-active, and write ownership
A regional topology defines where the service runs and which regions may serve each operation. Compare write ownership separately from replication timing: a region may serve reads while another owns writes, and a write may wait for remote durability before success. These choices determine both normal latency and what remains possible after a region is lost.
| Topology | Write behavior | Benefit | Cost or limit |
|---|---|---|---|
| Primary with asynchronous standby | East writes; West catches up later | Simple normal ownership and lower write coordination cost | Acknowledged changes may be missing after regional loss |
| Cross-region synchronous commit | Success waits for the required remote durable state | Can protect acknowledged writes against the named regional failure | Network latency and possible refusal during partitions |
| Multiple serving regions, one home writer per key | Each tenant/key has a defined write owner | Geographic service without arbitrary concurrent conflict | Remote writes and ownership-transfer work remain |
| Concurrent writers with defined merge semantics | Regions independently accept mergeable operations | More local write availability for suitable data | Not safe for arbitrary inventory, money, or ownership changes |
Start with a primary region and a standby. East owns writes. West receives the ordered change stream. Reads may use West only under a stated staleness policy. This is easier to reason about than allowing both regions to update the same inventory row independently.
Active-active means more than two copies of a web server. If both regions accept writes, specify ownership or conflict handling. Assigning each tenant a home region gives one authority per tenant. Globally coordinating a row can preserve stricter guarantees but adds cross-region latency. Accepting concurrent updates and merging them requires business-compatible semantics; “last timestamp wins” can silently erase an order or inventory reservation.
Read replicas, immutable assets, and regional caches can reduce geographic read latency without making all writes multi-primary. Choose the narrowest distributed-write requirement the product actually needs.
A different design puts one voting, data-bearing replica in each of three regions and commits through a proven majority protocol. Every acknowledged write is durable in two regions. After any one region is lost, the two survivors can elect according to the protocol and recover the committed history; a lagging survivor cannot simply ignore the protocol's election restrictions. This is a constructed quorum example, not a claim that every three-region product uses this layout. It costs cross-region commit latency and still depends on surviving network and service capacity.
04Regional failover: detection, fencing, promotion, and routing
Assume East acknowledged O16 at 12:00:00 and West durably applied it. East acknowledged O17 at 12:00:04, but its log entry has not reached West. Connectivity fails at 12:00:05.
- At 12:00:10 monitoring detects failure. It cannot infer whether East is dead or merely unreachable from West.
- A promotion procedure establishes that the old writer cannot continue accepted writes under the ownership protocol. A fencing epoch is an increasing ownership-generation number. Resources that check the current epoch can reject an old writer’s operations; changing a number without an enforcing resource does not stop the old process.
- West is promoted from its last safe durable position. In this example O17 may be absent, despite its prior acknowledgement. The observed loss window is five seconds; the missing record was accepted one second before disruption.
- Routing moves eligible traffic. DNS caches, connection pools, and clients may keep using old endpoints, so routing changes alone do not fence the old writer.
- The team validates order creation and payment reconciliation before declaring recovery complete. If that happens at 12:07:05, service recovery took seven minutes.
The client retries O17 using its original operation identity. A payment might have succeeded outside the lost database state. The recovery path queries the payment attempt or reconciles provider events rather than charging blindly. The write and payment contracts must survive the disaster plan together.
The payment recovery identity must also survive. Store the original client operation ID and provider attempt/resource reference in recoverable state, or ensure the provider can recover the mapping from a durable business reference. If both the mapping and the acknowledged order are lost, the client retry alone does not prove whether a charge exists. Hold new charging attempts while reconciliation reconstructs that fact.
- 1 → 2replicate durable logEast: writer epoch 7 → West: asynchronous standby
- 3 → 1may not yet exist in WestO17 acknowledged in East → East: writer epoch 7
- 2 → 4last safe recovery positionWest: asynchronous standby → Verify East cannot write; then promote
- 4 → 5only after old writer is stoppedVerify East cannot write; then promote → West: writer epoch 8
- 5 → 6restore useful serviceWest: writer epoch 8 → Validate and reconcile payments
05Standby capacity and restore-time estimates
A warm standby has some running resources and scales up during recovery. A hot standby keeps more capacity ready. Backup-and-restore starts from stored snapshots/logs and generally has more work on the recovery path. These are cost and recovery-time choices, not universal time guarantees.
Suppose peak traffic is 10,000 requests/s and West is provisioned for 2,000. Promotion without a capacity plan creates a second outage. Reserve or validate capacity, warm critical caches carefully, and use admission control while recovering. Include database connections, queue throughput, key management, identity providers, configuration, and secrets distribution in the dependency inventory.
For backup transfer alone, restoring 6 TB over a sustained 1 GB/s path takes approximately 6,000 seconds, or 100 minutes, before replay, indexing, startup, and validation. That cannot support a ten-minute RTO without another recovery mechanism. Use measured restore throughput, not a network-interface headline rate.
06Backups, point-in-time recovery, and restore validation
Point-in-time recovery restores a backup and replays retained changes only up to a selected moment. Choosing a point before a destructive update can recover data that live replicas have already deleted. The backup, required log history and decryption keys must all be available for that selected point.
Replication can faithfully copy corruption, deletion, or an application bug. Preserve point-in-time recovery logs and backups under access and retention policies that reduce correlated loss. Test restoration into an isolated environment, validate application-level invariants, and measure the entire process.
Retention has a business and security cost. Keep enough history to detect and recover from plausible mistakes while applying deletion and regulatory obligations deliberately. A disaster-recovery copy remains sensitive production data.
07Failback and disaster-recovery exercises
When East returns, it may have different data from West. Keep West in charge of new writes. Rebuild or reconcile East from West, verify replication, then plan the transfer back. Choose a clear switch point and prevent the former writer from continuing afterward. Old clients and running jobs must be rejected if they use an obsolete ownership version.
Run exercises that fail a database, sever regional connectivity, remove a dependency, and restore a backup. Record detection time, last recoverable write, promotion time, routing convergence, and usable capacity. The interview answer becomes credible when it identifies which promise the exercise validates and what would prevent declaring success.
Practise the interview questions
Say your answer aloud before opening the model answer. Then answer the follow-up and compare the reasoning.
What are RPO and RTO in disaster recovery?
Reveal a model answer
RPO, recovery point objective, is the targeted maximum data-loss window. RTO, recovery time objective, is the targeted time to restore the agreed service. If the last recoverable state is 5 seconds before a disruption and service returns 7 minutes later, those are separate data-loss and restoration measurements to compare with the objectives.
Interviewer follow-up
Does observing 5 seconds of replica lag guarantee a 5-second RPO?
Reveal the follow-up answer
No. It is an observation under one condition. The design needs a survival and recovery mechanism that supports the target under its stated failure assumptions; lag may grow during a worse outage.
What the answer must demonstrate: Keep the two objectives separate and distinguish targets from measured guarantees.
Can asynchronous regional replication promise zero loss of acknowledged writes?
Reveal a model answer
“Not by itself. East can acknowledge a write and fail before it reaches West. To survive that regional loss without losing acknowledged writes, acknowledgement must require a durable copy or quorum outside East, within the stated failure model.”
Interviewer follow-up
What is the price?
Reveal the follow-up answer
Commit must wait for surviving remote durable state under a safe protocol, adding latency and possible refusal during partitions. Also inspect voting placement: two of three voters in one region can acknowledge a majority that disappears with that region.
What the answer must demonstrate: Place the acknowledgement boundary.
The East primary stops responding to West. Why is that alone insufficient to promote West safely?
Reveal a model answer
A failed health check cannot prove that East stopped writing. In the asynchronous two-region design, I require a verified stop or removal of its write capability before promotion; if that is impossible, writes remain paused. Alternatively, a proven quorum protocol prevents the isolated minority from committing. A new epoch stored only in West is not sufficient fencing.
Interviewer follow-up
Does changing DNS solve split brain?
Reveal the follow-up answer
No. Cached DNS, existing connections, and background workers can still reach or execute the old writer.
What the answer must demonstrate: Separate routing from write authority.
Would active-active remove all regional outages?
Reveal a model answer
“It can improve continuity for some operations, but shared dependencies and write conflicts remain. I would state whether each key has one home writer, uses global coordination, or permits a defined merge. Inventory cannot simply merge arbitrary decrements without a rule.”
Interviewer follow-up
Where is multi-region serving straightforward?
Reveal the follow-up answer
Immutable assets and sufficiently stale-tolerant reads can be served regionally without concurrent mutable ownership.
What the answer must demonstrate: Describe per-record semantics.
Can a nightly backup meet a 30-second RPO?
Reveal a model answer
“A snapshot alone cannot: it can leave almost a day of changes absent. Continuous recoverable logs or another replication mechanism may narrow that gap. I also need to test restore and replay time against the RTO.”
Interviewer follow-up
How long does transferring 6 TB at 1 GB/s take?
Reveal the follow-up answer
About 6,000 seconds, or 100 minutes, before other recovery work, using decimal units.
What the answer must demonstrate: Check both freshness and duration.
The standby has one fifth of peak capacity. Is failover ready?
Reveal a model answer
“Only if the recovery contract allows bounded degradation and the remaining capacity or scaling is verified. I would test databases and dependencies too, prioritize essential operations, and limit admission rather than overload the new primary.”
Interviewer follow-up
Why might cold caches hurt?
Reveal the follow-up answer
Their miss storm shifts traffic to the recovering database precisely when it has the least headroom.
What the answer must demonstrate: Capacity is part of recovery.
Why keep backups when there are three replicas?
Reveal a model answer
“Replicas can copy an accidental deletion or corruption. Backups and point-in-time recovery preserve earlier states under a separate protection policy. I would regularly restore and validate business records, not only check the backup job status.”
Interviewer follow-up
What if encryption keys are unavailable?
Reveal the follow-up answer
The bytes may be intact but unusable; key recovery is a dependency in the restore exercise.
What the answer must demonstrate: Replication is not historical recovery.
An old primary region recovers after failover. Why should writes not immediately be routed back?
Reveal a model answer
“West has accepted new writes, so keep it in charge. Bring East up to date or rebuild it from West, verify the data, then switch writers through a controlled handover. Prevent the former writer from continuing. If the regions have conflicting histories, resolve them before switching.”
Interviewer follow-up
How do you measure success?
Reveal the follow-up answer
Run a real read/write/reconciliation check and confirm the agreed capacity, data state, and SLO, rather than checking only that processes are up.
What the answer must demonstrate: Failback is a controlled state transition.
Blank-page exercise · 20 minutes
Build the answer yourself
Design recovery for an order service with a 30-second RPO and ten-minute RTO. Then change the requirement to no loss of acknowledged orders.
- Place each acknowledgement and durable copy.
- Show the isolated old writer and its fencing mechanism.
- Budget detection, promotion, routing, and validation time.
- Include capacity, payments, backups, and failback.
Check that each component and design decision follows from your requirements and workload.
Recall the key ideas
Answer from memory before opening each card. Explain why the choice works and what it costs. Revisit missed cards tomorrow.
Multi-region architecture and disaster recoveryRPO vs RTORecall first, then reveal
RPO: how far back data may go. RTO: how long recovery may take.
Point = data; time = service.
Return to lessonMulti-region architecture and disaster recoveryWest is ready to take over. Why not redirect traffic immediately?Recall first, then reveal
East may still accept writes. First prevent the old writer from committing, then promote West, route traffic and validate recovery.
Stop old writes → enable new writer → verify.
Return to lessonMulti-region architecture and disaster recoveryBackup vs replicaRecall first, then reveal
A replica follows changes; a protected backup preserves an earlier recovery point.
Copies need history.
Return to lessonFinal revision
Summary and interview notes
Disaster recovery is a tested procedure for restoring an agreed service from a surviving data point. Choose RPO and RTO first, then align acknowledgment, replica placement, write authority, capacity, external-effect recovery and failback with those objectives.
Remember these points
- RPO is the target data-loss window; RTO is the target restoration time, and observed lag is neither promise by itself.
- Asynchronous replication can lose acknowledged writes; zero-loss acknowledgment must depend on state surviving the named failure.
- Replica and voter placement matter: a majority concentrated in one region does not survive that region’s loss.
- Changing routes does not stop the old writer. Before promoting another, enforce exclusive write ownership or verify that the old writer has stopped.
- Backups protect historical recovery points, while replicas can quickly copy corruption and deletion.
Interview tips
- Mark every acknowledgment and durable copy on the failover trace.
- Challenge the design with a partition where the old primary remains alive, not only a clean power-off.
- Budget detection, authority transfer, capacity, routing and validation; calculate restore bytes divided by measured throughput.
Important qualifications
- Six decimal TB at one GB/s needs about 100 minutes for transfer alone.
- Payment identity and encryption-key recovery must survive the disaster along with primary business records.
- Before moving back to the recovered region, rebuild or reconcile its data from the region currently accepting writes.
Technical references
- AWS disaster recovery strategiesBackup/restore, standby, and regional recovery strategies; actual objectives require measurement.
- Google SRE: addressing cascading failuresCapacity loss, overload, and recovery interactions.
- Ongaro and Ousterhout: RaftMajority commitment, leader completeness and election restrictions used in the separate three-region quorum example.
Practice marks stay in this browser.