System designby Learnastra

System-design interview · Extended interviews

Design a payment system and ledger

By Anup Rai

Design a payment workflow that can recover a lost processor response, record each financial operation once in a balanced journal and prevent concurrent refunds from exceeding captured funds.

You will learn to

  • Represent payment workflow states separately from immutable accounting entries.
  • Resolve ambiguous processor outcomes using stable operation identity and reconciliation.
  • Explain balances, refunds, concurrency, and auditability using concrete amounts.

Practice in this chapter

8 interview questions with model answers and follow-ups.

Go to interview practice

Useful foundations: Databases, data models, and ACID transactions · Message queues, event logs, delivery guarantees, and backpressure · Quorums, consensus, leases, and fencing

Workload and timing examples are interview assumptions.

01Problem and scope

A merchant payment platform coordinates authorization, capture, status and refunds through an external processor while maintaining an internal accounting record. Each payment operation must have one logical effect, each journal must balance within its currency, and refunds must stay within captured funds. A timeout is an unknown external outcome. A USD 25.00 capture is the example. Merchant checkout, stored-value transfers and processor accounting are distinct products and should not be conflated.

Candidate: “I will design merchant checkout through one external processor, plus the internal payment history and ledger. Does a timeout permit us to show pending while we find the result?” Interviewer: “Yes, but a retry must not charge the customer again, and merchants need partial refunds.” The service can show pending while it recovers the result instead of guessing whether the charge succeeded.

A payment intent records the workflow: what the customer wants and which steps have completed. A ledger records financial movements. In double-entry bookkeeping, each journal has debit and credit entries whose totals balance within its currency. Workflow state may change as new evidence arrives; a posted journal is corrected with a new journal, not erased.

We support authorization, one full capture per payment, status and full or partial refunds. Partial or incremental capture is a separate extension requiring its own reserved-capacity model. Authorization reserves spending capacity under the provider’s contract; capture requests the financial movement. We exclude lending, foreign-exchange conversion and a complete dispute platform. The account names below form a simplified platform example, not a claim about a particular company’s accounting system. Every volume, latency and retention value is an explicit interview assumption.

02Functional requirements

  1. Create a payment intent. A merchant creates a payment intent for order o81 and a server-validated price of USD 25.00 (2500 minor units). A retry with the same identity returns the same pay81, even if it reaches another application instance.
  2. Collect provider authorization. The customer supplies a processor-issued payment token and completes any required customer action through the provider-supported flow. Our service never stores raw card numbers in its own intent rows.
  3. Authorize and capture. The service requests authorization and capture, separately when required. The merchant sees captured only after authoritative processor evidence has been durably applied locally. Accepted-for-processing is a distinct state.
  4. Refund captured funds. An authorized merchant agent requests a full or partial refund. The total of successful refunds plus unresolved reserved refunds cannot exceed the captured refundable amount.
  5. Read status and history. The customer and the merchant retrieve status and paginated history. An unknown processor outcome appears as processing or reconciliation-required, with a stable resource to check.
  6. Reconcile financial facts. Operators compare processor and settlement facts with internal journals and resolve discrepancies through auditable actions.

Acceptance boundaries

A browser redirect or a client message saying “payment succeeded” is not authoritative evidence. Likewise, sending an email is not part of the financial commit: it follows a durable outcome event. A customer action, a processor result and a local database commit prove different things; the API reports each separately.

03Non-functional requirements

  1. Intent latency. Target p95 below 200 ms, excluding customer/provider steps; reject before acceptance when durable capacity is unavailable.
  2. Status latency and availability. Target p95 below 150 ms and 99.95% eligible-request availability. Return known pending state rather than fabricated success.
  3. Capture completion. Target p95 within 3 seconds when the provider is healthy. External uncertainty can remain pending beyond this objective.
  4. Durability. Acknowledged intents and postings survive one zone loss through durable quorum in the selected regional design.
  5. Journal retention. Use seven-year illustrative retention; confirm the actual business policy before deployment.
  6. Isolation and authorization. Enforce customer authorization and merchant isolation on every status, refund and export path.

Workload and money representation

Assume ten million new payment intents/day and a twentyfold peak over the daily average, with merchants across several currencies. Each payment uses one currency and integer minor units; not every currency has two decimal places. The amount comes from a verified order/quote, not an editable browser total.

Financial invariants

A financial invariant is a rule every committed operation must preserve, even if doing so means declining work during an outage. Here the rules protect accounting integrity and prevent the same captured funds from being promised to overlapping refunds.

Invariant Rule when targets conflict
One journal per financial operation; balanced by currency Stop unsafe posting rather than improve an uptime number
Captured ≥ successful + unresolved reserved refunds Keep the reservation while the outcome is unknown

Regional disaster boundary

A total regional loss is separate from one-zone durability. Zero regional data loss requires the service to wait for the necessary durable copies outside the region before acknowledging writes, with the corresponding latency and availability costs. Preserve processor references for reconciliation, test restore, and describe the recovery gap.

04Capacity estimates

Assume ten million new payment intents/day and a twentyfold peak over the daily average.

Estimate Arithmetic Boundary
New intents Ten million ÷ 86,400 ≈ 116/s; peak ≈ 2,315/s Validate the burst multiplier
Status reads Four/intent ≈ 463/s average and 9,260/s coincident peak Polling policy can invalidate this ratio
Retained records Ten records × 1 KB/intent = 100 GB/day Includes workflow facts and journals
Long-term storage 36.5 TB/year × seven years ≈ 255.5 TB Before indexes, replicas and backups
Three live copies 255.5 TB × 3 ≈ 766.5 TB Archival access is significant even at moderate QPS
Local transaction phases Intent/authorization command + authorization outcome + separate capture claim + capture posting: 4 × 2,315 ≈ 9,260 transactions/s Before refunds, callbacks and reconciliation
Outage backlog 2,000 unresolved intents/s × five minutes = 600,000 Already accepted work needs recovery
Net recovery 3,000 completions/s − 2,000 new/s = 1,000/s Another ten minutes, if the provider permits that rate
Provider metadata traffic Two provider calls/intent × 2 KB/call × 2,315 intents/s ≈ 9.26 MB/s Authorization plus capture requests/responses under this payload assumption; before retries, status queries, TLS and protocol overhead

Check the assumptions

A page polling ten times per second breaks the four-reads-per-intent estimate; backoff and push notifications affect capacity. Do not multiply the retained-record estimate again as though journal entries were omitted.

Provision the transaction and recovery paths

Benchmark the real transactional workload and merchant partitioning plan before concluding that one database cannot work. Reserve reconciliation capacity. Dependency limits and financial correctness dominate raw provider bandwidth, and backlog drain requires actual spare completion capacity.

These rates count new logical payments, not every HTTP or provider retry. The four-phase local estimate follows this chapter’s separate authorization and capture APIs. Automatically claiming capture in the authorization-outcome transaction could combine two phases, but that is a different flow and must be stated explicitly.

05APIs and contracts

A payment identifies the purchase workflow; an operation identifies one authorization, capture, cancellation or refund within it. Keeping those identities separate lets a caller retry one step or inspect its unresolved outcome without creating another purchase.

Interface Example Result and failure meaning
Create POST /payments, key checkout-81, {orderId:o81, amountMinor:2500, currency:USD, paymentToken:tok_demo} 201 with pay81 and state; same key/payload returns original resource
Authorize POST /payments/pay81/authorizations, key authorize-81 202 plus an authorization operation; report any provider-required customer action
Cancel authorization POST /payments/pay81/authorization/cancel, key void-81 Accepted cancellation intent or conflict with an already claimed capture; provider confirmation may remain pending
Capture POST /payments/pay81/captures, key capture-81, {amountMinor:2500} 202 plus operation op81 while external work proceeds
Read GET /payments/pay81 Known state, version, captured/refund totals and pending operation IDs
Refund POST /payments/pay81/refunds, key refund-81-a, {amountMinor:500} Refund rf81 if capacity is reserved; 409 if the requested amount is unavailable
History GET /payments?after=(createdAt,paymentId)&limit=50 Tenant-scoped deterministic cursor page
Provider event Signed event evt902 referencing processor capture ch81 Acknowledge after durable inbox acceptance, not before

Scope keys to the authenticated merchant and operation type. Store a canonical payload fingerprint; changing amount under an existing key is a conflict, not a new attempt. A 202 never means captured. A provider timeout returns a durable pending resource, while validation or permission errors do not create processor work.

The capture API describes our service’s operation. The adapter maps it to the chosen provider’s documented object/state flow, and stores the provider resource ID as soon as known. Provider idempotency retention is finite and implementation-specific. After its safe retry window, an old local key alone cannot make a repeated remote call safe. Status queries and reconciliation replace blind retries.

The capture command locks the payment and claims its single capture slot in the same transaction as inserting the operation and outbox. A different idempotency key cannot create a second capture for that payment: return the existing capture operation or a conflict. The slot remains claimed while its outcome is unknown. This product permits the validated full amount once; supporting partial captures would require atomically reserving the remaining authorized amount, just as refunds reserve captured capacity.

Authorization, capture and cancellation

These are illustrative service states, not literal names shared by every processor. This exercise uses a payment method that supports separate authorization and capture. The adapter verifies provider evidence before advancing financial state.

State or transition Meaning and guard
Authorization pending / customer action required Persist a stable authorization operation before calling the provider. Complete any required customer action through the provider-supported flow; a timeout remains unknown.
Authorized Record provider authorization ID, currency, authorized amount and the provider-supplied capture deadline. Authorization reserves payment capacity; it is not captured revenue or a capture journal.
Capture pending / unknown Lock the payment; require known eligible authorization, the full requested amount within authorized capacity, a valid deadline and no accepted cancellation. Claim the one capture slot and persist its outbox intent together.
Captured Apply verified capture evidence through the existing unique balanced-journal posting transaction. Settlement remains separate.
Authorization expired or canceled A confirmed expiration or void releases the unused authorization; do not capture it. A new authorization requires a deliberate new operation and current customer/provider eligibility.

Capture and cancellation lock the same payment row, so the database decides which claim comes first. Once capture may be in flight, cancellation cannot report a successful void merely from local intent: reconcile that operation, then cancel a still-unused authorization or refund a confirmed capture as appropriate. Read the authoritative database clock after acquiring the row lock and leave a provider-dependent processing margin before the capture deadline. A valid local check does not stop authorization expiring before the remote call completes; a definitive provider rejection or unknown response still follows the established outcome-recovery path. Do not hard-code a universal seven-day hold: methods and networks have different validity rules. Stripe authorization and capture.

06Data model and access patterns

Payment and operation records track workflow progress. Journals and their entry lines record financial movements; a projection is a derived view, such as a balance total, maintained for convenient reads. The journal remains the accounting source from which that view can be checked or rebuilt.

The inbox stores received provider events durably before applying them. The outbox stores commands or notifications in the same transaction as the local state change that requires them, so a later worker can publish them without losing the handoff.

Record and key Important fields Ownership and query
Payment (merchantId,paymentId) orderId, currency, amount, state, version, authorizationRef, authorizedAmount, captureBefore, captureOperationId, cancellationOperationId Merchant’s payment authority; status by ID
Operation unique (merchantId,type,key) payloadHash, paymentId, providerKey, providerRef, state Retry lookup and pending-work scan
CaptureBalance (merchantId,captureId) captured, refunded, reserved Locked for each refund capacity decision
Journal unique (merchantId,operationId,postingKind) currency, immutable posting time, evidenceRef Prevents one operation posting twice
Entry (journalId,lineNo) accountId, debitMinor, creditMinor Append-only audit and balance reconstruction
Inbox unique (provider,eventId) verified body reference, processing state Duplicate notification intake
Outbox (merchantId,eventId) aggregate version, payload, published state Recoverable downstream publication

Initially, partition by merchant so its payment and journal records can commit in one transaction. A directory maps merchant ranges to database shards. Large merchants may eventually need an explicit subledger design; casually hashing individual entry IDs would scatter a balanced journal across independent commits.

Indexes support merchant/payment lookup, (state,nextAttemptAt,operationId) recovery scans and (merchantId,createdAt,paymentId) history. A balance projection is updated with journal posting or reconstructed from entries. It is not allowed to override the journal. Direct entry writes are denied to ordinary services; an authorized posting routine validates currency and totals and commits the entire journal. A row-level CHECK alone cannot enforce an arbitrary multi-row journal sum.

07Basic working design

A useful baseline is a payment API, one transactional database and a small background worker calling one processor. The API creates pay81, its authorization operation and an outbox command in one transaction, then returns accepted. After verified authorization and any customer action, a separate short transaction validates the authorized amount/deadline and claims the capture slot as op81 with its own outbox command. The worker scans pending commands directly; an external broker is unnecessary at this size. The processor lies outside the database transaction.

In this example, processor receivable records money owed to the platform by the processor, and merchant payable records money the platform owes the merchant. Processor settlement transfers funds owed to the platform; paying the merchant is a separate movement that reduces merchant payable. These later movements need their own journals. These account roles explain the two sides of the capture journal below.

The API first replies after the payment intent commits. If the service dies after that commit, another worker can find the pending operation. Authorization and capture use distinct stable identities; an authorization alone creates no capture journal. When the processor later reports ch81 captured for USD 25.00 (2500 minor units), a local transaction records that fact and journal j81. In the simplified platform account model, j81 debits processor receivable by 2500 and credits merchant payable by 2500. The two totals match within USD. Settlement later changes receivable/cash/fee accounts through additional journals.

Status reads use this same database and return its current known state. The single region simplifies authority: there is one place to decide whether op81 was applied. The baseline already has durable request identity because a low-traffic service can still lose a response. It deliberately does not yet have many shards, a cache or multiple providers. We benchmark the real posting transaction before deciding which scaling change solves an observed limit.

architecture · baselineBaseline: local commits around provider calls

Intent and posting are local transactions. The processor’s effect is a separate external fact.

Baseline: local commits around provider callsIntent and posting are local transactions. The processor’s effect is a separate external fact. client to api: 1. Create / authorize checkout-81; api to db: 2. Commit payment + operation + outbox; db to worker: 3. Read pending operation; worker to provider: 4. Authorize with stable operation key; worker to db: 5. Commit authorization outcome; worker to db: 6. Commit capture claim and outbox; worker to provider: 7. Capture with stable cap-op81; worker to db: 8. Commit capture and balanced j811. Create / authorizecheckout-812. Commit payment +operation + outbox3. Read pending operation4. Authorize with stableoperation key5. Commit authorizationoutcome6. Commit capture claim andoutbox7. Capture with stablecap-op818. Commit capture andbalanced j81ACTORCheckout clientsSERVICEPayment APISTOREPayment DB andjournalWORKERProvider workerEXTERNALExternal processorsyncasync
Read each connection in order
  1. sync1. Create / authorize checkout-81Checkout clients → Payment API
  2. sync2. Commit payment + operation + outboxPayment API → Payment DB and journal
  3. async3. Read pending operationPayment DB and journal → Provider worker
  4. sync4. Authorize with stable operation keyProvider worker → External processor
  5. sync5. Commit authorization outcomeProvider worker → Payment DB and journal
  6. sync6. Commit capture claim and outboxProvider worker → Payment DB and journal
  7. sync7. Capture with stable cap-op81Provider worker → External processor
  8. sync8. Commit capture and balanced j81Provider worker → Payment DB and journal

08Find the baseline flaws

Suppose the baseline database sustains an assumed 2,500 local transactions/s at the target tail latency. The estimated 9,260 peak phases/s exceed that measured capacity before recovery work is added. Increasing API replicas does not fix the shared posting bottleneck. Status reads at 9,260/s may compete for the same I/O and buffer pool, so history reads that tolerate lag should run separately from decisions that require current balances.

Now examine a correctness counterexample. Worker W1 sends capture op81. The processor commits ch81, but its response disappears. W1 marks the request failed and a replacement worker invents op82. Both captures can succeed. No amount of database replication repairs this duplicated external effect. Preserve op81, record uncertainty, and reconcile the same remote operation instead.

A second race concerns refunds. Agents A and B both read captured=2500 and refunded=0, then each request 2000. If the read and reservation are separate, both may send externally, totaling 4000. A unique refund ID does not help because these are two different valid IDs. Both decisions must lock or atomically compare the same captured-capacity record.

These failures drive different changes: distribute independent merchant workloads for capacity; use stable identities and a protected refund balance for correctness. A queue alone prevents neither duplicate charges nor excessive refunds.

09Improve the design, step by step

First, protect and parallelize provider work. The trigger is the 600,000-item outage backlog and provider concurrency limits. A transactional outbox relay publishes operation IDs to a durable queue; bounded workers claim work and use the already-stored provider key. A separate recovery pool handles old unknown outcomes. The API can quickly save accepted work, while provider workers scale separately. The cost is queue storage, extra handoffs and pending states. Duplicate delivery is the new risk, handled by operation state and stable provider identity. Direct database polling is the rejected alternative only once scans or scheduling fairness become expensive; it remains simpler at modest load.

Second, split independent merchant authorities. The trigger is a measured 2,500-transactions/s shard versus roughly 9,260/s peak workload. Six comparable shards offer 15,000 transactions/s of assumed measured capacity, about 62% utilization before the omitted recovery and refund work. This is an initial sizing candidate; benchmark merchant skew and failure reserve rather than dividing blindly. Merchant routing sends pay81 and its journal to one shard. A synchronous replica set protects each shard against the stated zone failure. Each shard holds a smaller active dataset and posts independently; operators must manage routing, migrations and more database groups. A stale directory is a new risk, so owners validate routing epochs. A larger single database is a reasonable alternative when it meets the target with less operational work. Cross-merchant financial transfers remain outside this partition-local scope until a deliberate transaction design is added.

Third, isolate reads without weakening decisions. Status-history and reporting traffic now competes with posting. Serve explicitly stale-tolerant history from read replicas or a derived reporting view, while immediate pay81 status after submission carries a minimum committed version or goes to its authority. A refund checks current balances in the primary database transaction; it never trusts the reporting copy. The benefit is predictable posting capacity; the cost is replication traffic and read routing. Replica lag is the new risk. Keeping every read authoritative is preferable when the read volume fits or the product requires current answers everywhere.

Fourth, close the evidence gap. Missing callbacks and lost responses motivate a durable webhook inbox and scheduled settlement reconciler. Both feed the same outcome applier, which validates facts and invokes the unique journal-posting transaction. This recovers effects the request path missed. Costs include provider queries, settlement ingestion and explicit unresolved cases. Badly matched imported facts can create incorrect postings, so match processor reference, merchant, amount, currency and operation kind. Relying only on signed webhooks is simpler but cannot independently detect a missed event or accounting mismatch; retain that alternative only when a weaker recovery contract is acceptable.

Each change addresses a measured bottleneck or a specific failure. None makes a remote call part of our SQL transaction. Their success is measured by backlog drain, posting latency and reconciled outcomes, not the number of new components.

10Detailed architecture

Merchant-owned transactions

The public payment API authenticates the merchant, validates the order/quote, and routes by merchant through a versioned shard directory. Each shard owns intents, operations, capture capacity, journals, entries and local inbox/outbox state. Its replicas provide the selected durability protocol; the diagram’s replication edge is not a second independent writer.

External execution and evidence

An outbox relay exports committed commands into an operation queue. Provider workers consume operation IDs and read/claim their durable state at the owner. The queue is a wake-up and scheduling mechanism, not the sole record of what money should move. Workers call the external processor with stable keys, and submit verified outcomes to the posting service. A webhook receiver verifies the provider signature, persists an inbox record, then returns promptly; applying the callback to payment and ledger state happens asynchronously.

One posting authority

Both callback processing and settlement reconciliation use the same merchant database and posting routine. They do not append independent journals into separate databases. A merchant-event relay publishes captured/refunded facts after the journal commit. Order fulfillment consumes these idempotently; email and analytics cannot delay the financial transaction.

Read paths and implementation

History replicas answer only reads whose freshness contract permits them. Immediate status reads and commands that change available funds reach the authoritative database. Region-local synchronous work is therefore short: authorize the request, route, execute one database transaction and respond. Workers handle provider calls, callbacks, reconciliation and merchant notifications later. Saved operation IDs let them resume after a crash.

A practical starting stack is PostgreSQL for the merchant-local transaction, a provider SDK for documented authentication and request semantics, and a bounded worker using the database outbox. Add a broker for measured scheduling or isolation needs. With PostgreSQL synchronous replication, explicitly select the durability policy and synchronous standbys across the intended failure domains; ordinary asynchronous replication does not satisfy the zone-loss acknowledgment claim by itself. Neither the SDK nor a unique SQL key makes a remote processor part of the local transaction.

architecture · finalFinal: merchant authority and recoverable evidence

All outcome paths meet at the merchant’s posting authority; queues and processor calls remain outside its transaction.

Final: merchant authority and recoverable evidenceAll outcome paths meet at the merchant’s posting authority; queues and processor calls remain outside its transaction. client to api: 1. Pay, status or refund; api to directory: 2. Resolve merchant owner + epoch; api to db: 3. Commit intent / read current state; db to replicas: Durable replication before ack; db to relay: 4. Read committed outbox; relay to queue: 5. Publish operation ID; queue to worker: 6. Deliver bounded work; worker to db: Claim / load stable provider key; worker to provider: 7. Call or query same operation; worker to applier: 8. Submit verified outcome; provider to webhook: Signed event evt902; webhook to db: Persist inbox before 2xx; webhook to applier: Wake recoverable inbox processing; reconcile to provider: Read settlement / operation evidence; reconcile to applier: Submit matched missing facts; applier to db: 9. Atomic unique journal + outcome; relay to consumer: 10. Committed captured/refunded event; api to replicas: Permitted stale history reads1. Pay, status or refund2. Resolve merchant owner +epoch3. Commit intent / read currentstateDurable replication before ack4. Read committed outbox5. Publish operationID6. Deliver bounded workClaim / load stable provider key7. Call or query same operation8. Submit verified outcomeSigned event evt902Persist inbox before 2xxWake recoverable inboxprocessingRead settlement / operationevidenceSubmit matched missing facts9. Atomic unique journal +outcome10. Committedcaptured/refunded eventPermitted stale history readsACTORCheckout / merchantclientsG1SERVICEAuthenticatedpayment APIG1STOREVersioned merchantdirectoryG1STOREMerchant shard:operations / journalG2STOREDurable shardreplicasG2WORKEROutboxcommand/event relayG2QUEUEBounded operationqueueG3WORKERProvider workersG3EXTERNALExternal processorG4SERVICESignature check +durable inboxG3SERVICEOutcome applier /posting routineG2WORKERSettlement reconcilerG3SERVICEMerchant orderconsumersG3syncreplicationasyncG1 Authenticated application boundaryG2 Merchant ownership / local transactionsG3 Recoverable asynchronous processingG4 Independent processor authority
Read each connection in order
  1. sync1. Pay, status or refundCheckout / merchant clients → Authenticated payment API
  2. sync2. Resolve merchant owner + epochAuthenticated payment API → Versioned merchant directory
  3. sync3. Commit intent / read current stateAuthenticated payment API → Merchant shard: operations / journal
  4. replicationDurable replication before ackMerchant shard: operations / journal → Durable shard replicas
  5. async4. Read committed outboxMerchant shard: operations / journal → Outbox command/event relay
  6. async5. Publish operation IDOutbox command/event relay → Bounded operation queue
  7. async6. Deliver bounded workBounded operation queue → Provider workers
  8. syncClaim / load stable provider keyProvider workers → Merchant shard: operations / journal
  9. sync7. Call or query same operationProvider workers → External processor
  10. async8. Submit verified outcomeProvider workers → Outcome applier / posting routine
  11. asyncSigned event evt902External processor → Signature check + durable inbox
  12. syncPersist inbox before 2xxSignature check + durable inbox → Merchant shard: operations / journal
  13. asyncWake recoverable inbox processingSignature check + durable inbox → Outcome applier / posting routine
  14. syncRead settlement / operation evidenceSettlement reconciler → External processor
  15. asyncSubmit matched missing factsSettlement reconciler → Outcome applier / posting routine
  16. sync9. Atomic unique journal + outcomeOutcome applier / posting routine → Merchant shard: operations / journal
  17. async10. Committed captured/refunded eventOutbox command/event relay → Merchant order consumers
  18. syncPermitted stale history readsAuthenticated payment API → Durable shard replicas

11Write path and acknowledgement

Local intent acceptance and external money movement are separate commit boundaries. Persist operation identity before the processor call and retain uncertainty until evidence resolves it.

  1. The customer submits checkout-81. The API derives merchant M7 from the verified checkout context, validates the order amount and token reference, and routes to M7’s current shard.
  2. One transaction inserts the unique create identity, pay81 and the initial outgoing intent. Repeated keys compare the canonical payload and return the saved resource. A commit failure returns no acceptance promise.
  3. The authorization operation persists its identity, resolves any required customer action and records verified authorization amount, currency and capture deadline. A capture command then locks pay81, checks that authorization is eligible, unexpired and not being canceled, claims its only capture slot and records op81 with immutable amountMinor=2500, currency=USD and provider attempt key cap-op81. A unique per-payment capture claim rejects a second command even if its idempotency key differs. The outbox insertion shares this commit.
  4. A worker claims the operation, checks that it still requires a remote call, and invokes the provider outside database locks. A delivery retry reuses cap-op81; it does not manufacture another payment attempt.
  5. A definitive capture result identifies ch81. A timeout changes the local state to unknown and schedules status recovery. The provider may already have succeeded, so the refund capacity is not inferred from an HTTP error.
  6. The outcome applier validates merchant, currency, amount and allowed state, locks op81 and posts its unique capture journal together with captured status, balance projection and an order-update outbox event.
  7. Only that successful local commit allows a captured response or event. If its acknowledgment is lost, retrying the same outcome finds the journal and returns the stored result. The worker acknowledges its queue item after a recoverable local state has been saved.

A successful remote effect with failed local posting is repaired by replaying the outcome, not by calling capture again. The diagram’s lost response interval is exactly where that distinction matters.

12Read and delivery path

Payment status exposes the authoritative known outcome. Derived history or caches cannot turn an unresolved processor call into a definitive failure.

  1. The customer requests pay81 under authenticated checkout ownership. The API verifies access before exposing status; knowing a payment ID is insufficient.
  2. Routing uses M7’s current shard. For a status read immediately following submission, use the authoritative owner or a replica proven to have applied at least the acknowledged version. A stale replica must not turn a recorded capture back into “not found.”
  3. The service loads the payment and its operation summary. It returns captured only after verified processor evidence has committed locally, otherwise a precise pending, action-required, failed or reconciliation state. These values describe known facts, not an attempt to guess the provider’s current state from elapsed time.
  4. A refund screen may display an explanatory total, but the subsequent refund command rechecks capacity atomically. A displayed available amount is never a reservation.
  5. Merchant history uses (createdAt,paymentId) keyset pagination, capped at fifty records. Stale-tolerant history can use a read replica and report its freshness. A cursor is a position, not a fixed snapshot unless an explicit versioned export contract says so.

Cache static payment-method metadata if useful, but do not put mutable refund capacity in an unversioned shared cache. Large exports run against an auditable snapshot or a declared reporting cutoff so scanning seven years of history does not consume the database resources needed for live posting.

13Correctness deep dive

A unique webhook event ID suppresses repeated delivery of that event. It does not alone prevent two different notifications about the same capture from posting twice. The financial operation has its own uniqueness key: (M7, op81, capture). The posting routine serializes all outcomes for op81 and either returns its existing journal or commits the new one.

applyCapture(op81, fact ch81):
  begin transaction
  lock Operation(op81)
  verify fact merchant, currency, amount and provider reference
  if Journal(M7, op81, capture) exists: return existing result
  if fact conflicts: abort; record durable review case; return
  construct debit receivable 2500 / credit payable 2500 in USD minor units
  assert sum(debits) == sum(credits) and account/currency validity
  insert unique Journal and all Entries
  update Payment, CaptureBalance and balance projection
  insert captured event in Outbox
  commit

The review case is recorded after an aborted conflicting attempt, without posting money. Restrict direct writes so every posting uses this transaction. Worker W1 and callback applier W2 may reach it together. One holds the operation lock; the other waits and then observes the existing journal. If W1 aborts, it leaves no half-journal. If W1 commits and loses its reply, W2 returns j81. The unique constraint remains a second guard against an implementation race.

Refund requests lock the same capture balance before reserving funds:

This protects the limit even though provider calls occur outside the lock. A definitive failed refund can release its reservation transactionally. A refund confirmation appends reversing movements; it never edits j81. Partial refunds retain the remaining capacity, and reconciliation resolves unknown outcomes before capacity is reused.

sequence · retry-raceLost capture reply, then two local outcome paths

Remote success is recovered through op81. Both appliers reach the same operation lock and unique journal.

Lost capture reply, then two local outcome pathsRemote success is recovered through op81. Both appliers reach the same operation lock and unique journal. worker to provider: Capture USD 25.00, key cap-op81; provider to provider: Commit capture ch81; provider to worker: Capture reply lost; provider to applier: Signed event: ch81 captured; applier to db: Persist verified inbox event; commit; applier to db: Apply inbox: lock op81; post j81; db to applier: Commit captured + journal + outbox; worker to provider: Query/retry original cap-op81; provider to worker: Existing capture ch81; worker to db: Apply same outcome under op81 lock; db to worker: Existing j81; no second postingPARTICIPANTProvider workerPARTICIPANTProcessorPARTICIPANTCallback applierPARTICIPANTMerchant DB1. Capture USD 25.00, keycap-op812. Commit capture ch813. Capture reply lost4. Signed event: ch81captured5. Persist verified inboxevent; commit6. Apply inbox: lock op81;post j817. Commit captured + journal+ outbox8. Query/retry originalcap-op819. Existing capture ch8110. Apply same outcome under op81 lock11. Existing j81; no second postingsyncblockedreturn
Read each connection in order
  1. syncCapture USD 25.00, key cap-op81Provider worker → Processor
  2. syncCommit capture ch81Processor → Processor
  3. blockedCapture reply lostProcessor → Provider worker
  4. syncSigned event: ch81 capturedProcessor → Callback applier
  5. syncPersist verified inbox event; commitCallback applier → Merchant DB
  6. syncApply inbox: lock op81; post j81Callback applier → Merchant DB
  7. returnCommit captured + journal + outboxMerchant DB → Callback applier
  8. syncQuery/retry original cap-op81Provider worker → Processor
  9. returnExisting capture ch81Processor → Provider worker
  10. syncApply same outcome under op81 lockProvider worker → Merchant DB
  11. returnExisting j81; no second postingMerchant DB → Provider worker

14Failure and recovery

Failures can occur before the processor acts, after it acts but before we learn the result, or after our local posting commits. The recovery action depends on which boundary was crossed; treating every failure as permission to issue another capture would duplicate effects.

Failure or race Required response and boundary
Capture completed; worker crashed At 10:00:00 the processor captures ch81; at 10:00:01 the worker loses its reply and crashes. The durable op81 still has cap-op81 and a recoverable state. A replacement first queries or safely retries under the provider contract. A callback can independently supply the fact. The customer sees pending until the posting transaction commits; we do not falsely promise completion within three seconds during this incident.
Posting authority partition During a database-shard partition, only a side authorized by the replication protocol may accept posting. An isolated API cannot use an old read replica to approve a refund. It returns unavailable or pending for that operation. If the processor completed before local majority access disappeared, the external fact survives there and can be reconciled once our owner recovers. The service temporarily refuses new work rather than risk an incorrect journal or excessive refund.
Provider overload During provider overload, queue age rises. Per-provider concurrency and retry budgets limit calls, while admission control caps new accepted work according to the business’s allowed pending horizon. Already accepted operations are not silently dropped. Status reads remain separately provisioned; retries use backoff and jitter. Recovery follows the ten-minute drain estimate only if actual provider headroom exists.
Regional disaster A regional disaster can lose locally acknowledged asynchronous replication state under the stated regional limitation. Restore and reconciliation must compare processor effects before retrying old payment commands. If the product demands zero such loss, change cross-region acknowledgment and measure the extra latency rather than leaving the disaster promise ambiguous.

15Operations, security, and cost

Use processor tokens, least-privileged merchant roles and a separate audited refund permission. Verify webhook signatures over the provider-required payload representation, keep replay/deduplication records, and restrict outgoing provider credentials to adapters. Logs contain operation IDs and safe state transitions, not payment tokens or raw credentials. The merchant boundary must also appear in history queries, cache keys, settlement matching and operator tools.

Measure p95 intent latency and successful status availability, but alert first on unknown-operation age and unmatched amounts by currency. A five-dollar discrepancy is not equivalent to five failed HTTP calls. Track duplicate posting suppression, rejected refund reservations, provider throttling and reconciliation backlog. The posting routine’s balance assertion should reject a malformed journal and create an actionable audit signal.

The retention estimate gives a useful cost calculation without inventing cloud prices. Keeping one year hot and six years in a verified archive reduces hot logical record payload from about 255.5 TB to 36.5 TB, roughly sevenfold, before different index and replication policies. It introduces archive retrieval delay and restore complexity; archiving must keep the request identities and history needed to safely process current commands available. Processor operation fees and reconciliation calls may dominate storage savings, so report them as separate cost terms.

Roll out a new posting schema with old/new writer compatibility, shadow reconciliation and a small merchant cohort. Test crashes after remote success, after local commit and before queue acknowledgment; send reordered callbacks; race two refunds; restore a backup and replay captured events. The restored service must preserve operation/journal uniqueness. Passing a balance-total check alone cannot prove every real processor effect was recorded exactly once.

16Decision ledger and limitations

The design uses a local transaction to protect journal integrity and a separate recovery protocol to learn external processor outcomes. The decisions below explain what each mechanism protects, what it costs and when its assumptions would need to change.

Chosen mechanism Benefit Cost and consequence Change trigger
Durable intent plus asynchronous provider worker Recovers an accepted request across process loss Pending user state and queue operations Synchronous UX still waits on the same durable workflow, not a separate unsafe path
Merchant-local journal transaction Balanced posting and refund checks meet at one authority Hot merchant and cross-merchant operations remain hard Benchmark skew; design subledgers or distributed transactions explicitly
Stable provider identity plus reconciliation Resolves lost replies without inventing new effects Retention windows and provider evidence handling Provider lacks reliable lookup/dedupe: change contract or retain manual resolution
Immutable journals with balance projections Rebuildable audit trail and fast reads More records, explicit reversals and projection checks Archive historical records under a tested retrieval policy
Authoritative decisions, stale-tolerant history Protects financial writes from read traffic Read routing and visible lag More reads require current state: add proven fresh capacity

We have not designed cross-currency transfers, chargeback adjudication or unrestricted multi-region multi-primary posting. Each would require new account rules and coordination, beyond simply adding a status value.

17Interview closing

“I start with a durable payment intent and a stable identity for each external operation. The API commits intent before acceptance. Provider workers run outside the database transaction; callbacks and reconciliation return evidence to one outcome applier.

“Each merchant routes to a posting authority where a unique operation and a balanced journal commit together. Concurrent outcomes cannot create a second journal. Refund requests reserve capacity under the capture’s lock before calling the processor, so unknown refunds retain their claim. Status can remain pending while we learn an external outcome; an error response never proves no money moved.

“The assumed peak leads me to benchmark about nine thousand local transaction phases per second before callbacks, recovery and refunds, isolate status reads and split independent merchant authorities only when needed. My costs are pending states, extra records, provider recovery work and shard operations. The next measurements are the heaviest merchant’s posting latency and the provider-limited backlog drain rate.”

Interviewer: “Now require zero loss after a region disappears.” Candidate: “I change the acknowledgment boundary to durable survival outside that region and validate recovery of operation identity as well as journals. I would price cross-region latency and reduced write availability during partitions. Merely placing an asynchronous copy in another region would not meet the new promise.”

Practise the interview questions

Say your answer aloud before opening the model answer. Then answer the follow-up and compare the reasoning.

Foundation · Question 1

Why not keep a paid flag and current balance?

Reveal a model answer

“They tell me today’s state but not how it arose or how to correct it. An immutable balanced journal records each movement; current balances can be derived and checked against that history.”

What the answer must demonstrate: Explain auditability and the balancing invariant.

Applied · Question 2

A payment request times out after submission. Should the client create a new payment identity?

Reveal a model answer

“No. The capture might already have succeeded. The client retries the same checkout identity, while the service retries or queries the same processor operation and returns its known state.”

What the answer must demonstrate: Timeout does not prove no money moved.

Applied · Question 3

Can you put charging and ledger posting in one transaction?

Reveal a model answer

“I can atomically update my own database, but an external processor does not participate in that SQL transaction. I record intent, perform the call, and apply the outcome with stable identity and reconciliation.”

What the answer must demonstrate: Identify the external boundary.

Follow-up · Question 4

Two agents simultaneously refund USD 20 from a USD 25 capture. What happens?

Reveal a model answer

“They must transactionally reserve refundable capacity against the same capture. Only one can reserve USD 20; the other sees insufficient remaining capacity. Each accepted refund has its own stable operation identity.”

What the answer must demonstrate: Concurrent financial limits need atomic enforcement.

Foundation · Question 5

How do you process a duplicate success webhook?

Reveal a model answer

“Verify the callback, record its unique provider event ID, and apply the capture outcome under a uniqueness constraint on its ledger posting. Repetition returns success without another journal.”

What the answer must demonstrate: Do not trust callback order or authenticity by default.

Follow-up · Question 6

What does reconciliation add if your webhooks are reliable?

Reveal a model answer

“It independently compares internal operations and balances with processor and settlement records. It catches missing events, amount mismatches, fees, and operational mistakes that the normal callback path can miss.”

What the answer must demonstrate: Reconciliation is a correctness path, not only a dashboard.

Applied · Question 7

A processor has a 600,000-operation backlog, 3,000/s completion capacity and 2,000/s continuing arrivals. Why is its drain time not 200 seconds?

Reveal a model answer

“Because 2,000 new operations per second still consume capacity. Net drain is 1,000/s, so the backlog needs about 600 seconds if provider capacity remains available. I reserve recovery headroom and bound intake instead of equating worker throughput with backlog reduction.”

What the answer must demonstrate: Subtract continuing arrivals and identify the external limit.

Follow-up · Question 8

Two different provider events describe the same capture. How do you prevent two journals?

Reveal a model answer

“Event-ID deduplication alone is insufficient because the IDs differ. Both events resolve to the same stored capture operation, whose row is locked during application, and the journal has a unique merchant/operation/posting-kind key. One transaction inserts the balanced entries and outcome; the other observes that existing journal.”

What the answer must demonstrate: Locate the atomic enforcement, not just a generic deduplication claim.

Blank-page exercise · 45 minutes

Build the answer yourself

Run a 45-minute interview for a merchant payment service supporting one USD 25 capture and partial refunds. Produce the baseline, break it, evolve it, then prove capture posting and concurrent refund safety.

  • Clarify pending versus captured and the regional durability contract.
  • Calculate local transaction load, retained bytes and provider-limited backlog drain.
  • Draw baseline and final authority boundaries; trace create, capture and status.
  • Prove duplicate outcomes cannot repost and two refunds cannot overspend captured capacity.
  • Recover a lost provider response, test restoring the ledger, and explain the chosen design's storage, latency and operating costs.

Check that each component and design decision follows from your requirements and workload.

Recall the key ideas

Answer from memory before opening each card. Explain why the choice works and what it costs. Revisit missed cards tomorrow.

Design a payment system and ledgerCan SQL roll back an external capture?Recall first, then reveal

No. Track the uncertain result and reconcile or issue a separate refund operation.

A local rollback does not undo a processor capture.

Return to lesson
Design a payment system and ledgerHow is a refund recorded?Recall first, then reveal

Append new balanced entries that reverse the relevant amount; preserve the original journal.

Append correcting entries; preserve the original journal.

Return to lesson
Design a payment system and ledgerWhat does checkout-81 identify?Recall first, then reveal

One logical payment request with a fixed payload, reused across client retries.

Same checkout, same operation.

Return to lesson

Final revision

Summary and interview notes

A payment platform preserves stable external operation identity and applies verified outcomes through a merchant-local posting authority. Balanced immutable journals preserve accounting history; one capture claim prevents a second capture; refund reservations prevent overspending captured funds; reconciliation compares those local records with processor operations and settlement records.

Remember these points

  • Claim one full capture per payment atomically; a new request key must not bypass the business limit.
  • After a processor timeout, keep the operation unresolved: the charge may already have occurred.
  • Apply operation outcome, journal entries, balance changes and outbox event in one local transaction.
  • Keep unknown refund amounts reserved until definitive evidence resolves them.
  • Matching debit and credit totals do not prove that the journal names the right merchant, amount or processor result.

Interview tips

  • Name the local transaction and remote-effect boundary before drawing workers or queues.
  • Test two distinct capture commands, two distinct refunds and two different callbacks for the same operation.
  • Estimate local transaction phases and net reconciliation drain, not only payment API QPS.

Important qualifications

  • This example supports one full capture, one currency per payment and merchant-local posting; cross-merchant transfers and incremental capture require extensions.
  • Provider retry windows are finite; keys retained locally do not grant unlimited safe remote replay.
  • Retention and recovery objectives are interview assumptions to agree with the business.

Technical references

Practice marks stay in this browser.