System-design interview · Extended interviews
Design a payment system and ledger
Coordinate a merchant payment with an external processor, post balanced immutable journals and protect refund capacity while uncertain outcomes are reconciled.
You will learn to
- Trace durable payment intent, provider execution and one local ledger posting.
- Explain balanced journals, operation uniqueness and concurrent refund reservations.
- Recover unknown processor outcomes without issuing a new charge or trusting stale balances.
Practice in this chapter
9 interview questions with model answers and follow-ups.
Go to interview practiceUseful foundations: Databases, data models, and ACID transactions · Message queues, event logs, delivery guarantees, and backpressure · Quorums, consensus, leases, and fencing
Workload and timing examples are interview assumptions.
Dotted concept links open the relevant explanation in a new tab.
01Separate purchase workflow from accounting
Design merchant checkout through one external processor, with authorization, one full capture, status and full or partial refunds. A payment intent tracks the purchase workflow. A ledger records financial movements in immutable journals. Exclude foreign-exchange conversion, lending, disputes and unrestricted transfers between merchants; those change the account and coordination model.
Use payment P81 for USD 25.00, represented as 2500 minor units. Amount and currency come from a validated order, not an editable browser total. Use integer minor units and currency metadata; not every currency has two decimal places. Store a provider-issued token rather than raw card details.
Authorization reserves spending capacity under the processor contract. Capture requests the financial movement. Settlement later transfers funds according to the processor arrangement, and paying the merchant is another movement. These are distinct events, not interchangeable success labels.
The customer may see pending while an external result is unknown. Retrying one operation must not create another logical charge, and posted journals must balance within each currency. For this design, acknowledged operations and postings must survive one zone failure under synchronous regional replication and safe failover. Regional disaster recovery is a separate promise.
Clarify whether the prompt means merchant card payments or transfers of existing wallet balances, which operations the processor supports, and what “success” means to the caller. The main design covers authorization, capture and refunds with durable pending states; the separate wallet variant changes the account transaction.
02Functional requirements
Create and authorize a payment. Validate merchant, order, amount and currency, create a recoverable payment intent and initiate authorization through the processor.
Capture and report status. Request one full capture for an eligible payment and expose authoritative local state, including pending or unknown external outcomes.
Refund safely. Support full or partial refunds against remaining captured capacity, with recoverable operation identities and visible progress.
Record and reconcile money movements. Post immutable balanced journals, compare processor and settlement evidence with local records and repair missing postings through the same validated routine.
03Non-functional requirements
These are illustrative interview assumptions, not product facts or measured benchmarks. Confirm them before choosing components, then validate the completed design under the stated workload. Here p95 means the 95th-percentile latency: 95% of measured requests take no longer than that value. Report errors and rejected work alongside latency; a fast failure is not a successful outcome.
Workload. Plan for ten million intents/day, about 2,315 new intents/s at the assumed peak and 9,260 status reads/s. Size authorization, capture, callbacks and refunds as separate local transaction work.
Local response and evidence freshness. Target 300 ms p95 for durable local command acceptance and authoritative status reads under admitted load. After verified processor evidence is durably received, target its local application within five seconds for 99% of outcomes. Neither target promises processor authorization, capture or settlement completion within that time.
Financial consistency. Permit one full capture claim per payment, keep each currency’s journal balanced and never promise refunds beyond captured amount minus successful refunds and unresolved reservations. Timeout does not release uncertain financial capacity.
Durability and safe availability. Require acknowledged operations and postings to survive one zone failure under synchronous regional replication and safe failover. When posting authority is unavailable, keep outcomes pending and refuse financial decisions from stale replicas; regional loss is a separate recovery contract.
Security and external retry limits. Authenticate merchant ownership and refund/correction authority, verify processor evidence and keep raw card data with the provider. Preserve operation identity; after the provider’s deduplication window expires, reconcile instead of blindly resubmitting.
04Trace the full payment with one database and adapter
Begin with a payment API, one transactional database and a bounded worker using one processor adapter. The API authenticates merchant M7, validates the order and stores P81 with a stable authorization operation and outgoing-work record. It returns accepted only after that state is durable. The worker can find the operation after a process restart.
After verified authorization and any required customer action, a capture command locks P81, checks amount, currency and capture eligibility, and claims its single capture slot. It stores operation C81 and its immutable provider key before the network call. The call runs outside database locks.
When verified processor evidence establishes the capture, one local posting transaction records that fact, creates its journal, updates payment status and saves a merchant notification event. Only then does the service report captured locally. A provider success followed by a failed local transaction requires replaying the evidence, not capturing again.
The worker can initially poll indexed durable work; a broker is optional. Webhooks and a reconciliation worker offer additional ways to recover the same outcome. They all call the same posting routine. Separate paths must not create independent journals for the same financial operation.
The external processor and local database have separate commit boundaries.
Read each connection in order
- syncCreate / capture / refundMerchant / customer → Payment API
- syncDurable operation and workPayment API → Payments, operations, journals
- asyncPending operation IDPayments, operations, journals → Processor adapter
- syncStable provider operationProcessor adapter → External processor
- asyncVerified outcomeExternal processor → Outcome applier
- syncAtomic journal and statusOutcome applier → Payments, operations, journals
05Estimate transaction phases and long-lived records
Assume ten million new payment intents/day: about 116/s on average. A twentyfold peak is approximately 2,315 new intents/s. If clients make four status reads per intent, corresponding peak reads are about 9,260/s; aggressive polling can easily invalidate that estimate, so use bounded polling or notifications followed by status reads.
Separate authorization and capture can require four local transaction phases: initial authorization command, authorization outcome, capture claim and capture posting. At the assumed peak, that is about 9,260 transactions/s before refunds, callbacks and recovery. Benchmark this actual workload before deciding whether a single database needs sharding.
Ten retained records of 1 KB per intent produce 100 GB/day, or 36.5 TB/year. Seven years is about 255.5 TB logically before replicas and indexes. Historical records may move to a verified archive, but the request IDs, saved outcomes and recovery records needed by current operations must remain available.
At 2,000 unresolved operations/s, a five-minute processor incident adds 600,000 operations. If recovery completes 3,000/s while 2,000/s new work continues, net drain is 1,000/s and takes another ten minutes. This requires actual processor headroom; more worker threads cannot manufacture a higher downstream quota.
06Give each financial operation its own stable identity
Interfaces
| Request or message | Contract |
|---|---|
POST /payments with merchant-scoped key |
Creates or recovers P81 for the same validated purchase. |
POST /payments/P81/capture |
Claims the one full capture operation or returns the existing result. |
POST /payments/P81/refunds |
Reserves refundable capacity and returns a stable refund operation. |
GET /payments/P81 |
Reports known payment state and unresolved operation IDs. |
Request a partial refund
POST /payments/P81/refunds
Send the refund with the merchant- and operation-scoped key described below.
Refund amount in integer minor units
{
"amountMinor": 500
}
Stored records
| Record | Fields or identity | Purpose |
|---|---|---|
| Payment | Payment identity | Amount, currency, authorization reference, capture deadline and lifecycle. |
| Operation | Merchant, operation type and scoped request key | Type, scoped key, payload fingerprint, provider identity, state and evidence reference. |
| Journal and entries | Financial operation and posting kind | One immutable financial posting and all its balanced account lines. |
| Capture balance | Captured payment | Captured amount, successful refunds and unresolved refund reservations. |
Keys are scoped to the authenticated merchant and operation type. Reusing a key with a different amount conflicts. This deduplicates the same command; a separate per-payment capture claim prevents two different keys from creating two full captures.
An inbox durably receives verified provider events; an outbox records work and notifications with local state changes. Index unresolved operations by next-attempt time and merchant history by creation time plus payment ID. Status immediately after a write uses the authority or a proven caught-up replica; a reporting view cannot approve refunds.
07Explain the balanced journal before the posting algorithm
In double-entry accounting, each journal has debit and credit lines whose totals are equal within one currency. For this simplified platform model, processor receivable represents money the processor owes the platform; merchant payable represents money the platform owes the merchant. For the USD 25 capture, the journal uses integer minor units:
| Account | Debit (USD minor units) | Credit (USD minor units) |
|---|---|---|
| Processor receivable | 2500 | 0 |
| Merchant payable | 0 | 2500 |
This records the claim and obligation, not settlement into a bank account.
The posting routine uses one transaction:
Lock operation C81 and validate the evidence's merchant, provider reference, kind, currency and amount.
Check whether the unique journal for that operation and posting kind already exists. If so, return its result.
Otherwise validate all account lines and totals, then commit the journal, entries, payment state, balance projection and outgoing event together.
No observer should see half a journal.
Restrict direct entry writes so every writer uses the posting transaction. A row-level check on individual entries cannot by itself prove that an arbitrary multi-row journal balances. A balance projection is a convenient total derived from the journal; it can be checked or rebuilt and is not an independent source allowed to override history.
Correct errors by adding an authorized reversing or correcting journal with an audit link, not by erasing the original. Internal balance alone is insufficient: a perfectly balanced journal for the wrong merchant or amount is still wrong.
08Make timeout recovery part of the capture flow
The processor may complete C81 and lose its reply. Keep C81 unknown and use its original provider identity for supported idempotent recovery or lookup. A new operation C82 could charge again. Local uniqueness does not force an external processor to remember a key beyond its documented retention window; after that window, reconcile rather than retry blindly.
Verified webhooks are evidence, not commands to post without checks. Authenticate the provider payload, persist the incoming event, and match it to the stored operation. Distinct provider events can describe the same capture, so webhook event deduplication and journal-operation uniqueness solve different problems. Out-of-order facts must not turn a captured payment back into pending or declined.
Claim capture and authorization cancellation through the same payment row. Require a known eligible authorization, matching full amount and currency, a valid capture deadline and no incompatible cancellation. Read the authoritative database clock after acquiring locks. A local eligibility check does not prevent the authorization expiring during the remote call; the resulting rejection or uncertainty still follows the recovery path.
Once capture may be in flight, cancellation cannot simply declare a successful void. Resolve the capture, then cancel an unused authorization or refund a confirmed charge according to the observed outcome. Never treat a browser redirect as processor evidence.
09Reserve refundable capacity before calling the processor
Suppose the captured amount is 2500 and no refund has completed. Two agents each request a refund of 2000 using different valid keys. Deduplicating keys does not help because the requests really are different. Both agents must lock the same capture-balance record before deciding how much to refund.
Agent A locks that record and computes available capacity as captured minus successful refunds minus unresolved reservations. Initially it is 2500, so A reserves 2000 and stores its refund operation plus outgoing work in one transaction. Agent B then sees only 500 available and is rejected before making any provider call. The external calls occur after those short transactions release their locks.
| Outcome of A’s refund | Capacity and accounting transition |
|---|---|
| Succeeded | Move 2000 from reserved to refunded and append the balanced refund journal. |
| Definitively rejected | Release the reservation transactionally. |
| Unknown | Keep the reservation while reconciliation determines the result. |
Releasing it merely because a timeout occurred would let B promise the same funds again.
Partial refunds repeat this capacity decision against the remaining amount. Refund journals are new entries linked to the original capture, never edits to its history. A screen may display available balance, but the command must recheck it atomically because another refund can win after the screen was loaded.
The first reservation reduces capacity before either external refund can complete.
Read each connection in order
- syncReserve 2000 from captured 2500Refund API A → Capture authority
- returnReservation committed; 500 remainsCapture authority → Refund API A
- syncRequest another 2000Refund API B → Capture authority
- returnReject insufficient remaining capacityCapture authority → Refund API B
- syncExecute stored refund identityRefund API A → Processor
- syncUnknown keeps reservation; success postsRefund API A → Capture authority
10Scale merchant-local decisions and isolate processor work
Add a queue when pending-work scans or provider scheduling need independent scaling. Publish operation IDs from the outbox, then let workers load current state and invoke the stored provider operation. Queue delivery can repeat; it is a wake-up mechanism, not permission to invent a fresh capture.
Bound concurrency by processor and merchant. Reserve recovery capacity for old unknown operations so a stream of new payments does not leave them unresolved forever. Limit new acceptance when the wait for pending work would become unacceptable, while preserving already accepted operations and current status reads.
When measured posting capacity is insufficient, partition independent merchants. Keep a merchant’s payment, journal, refund balance and local inbox/outbox where they can commit in one transaction. Hashing journal lines across unrelated shards would prevent one local transaction from committing the balanced journal. A very large merchant may need a carefully designed subledger later rather than arbitrary extra hash buckets.
Move stale-tolerant reporting and history to replicas or derived views, with an explicit freshness contract. Immediate payment status and every financial capacity decision remain authoritative. Static payment-method metadata is a safer cache target than mutable refund capacity. Cross-merchant transfers and multi-region concurrent posting remain separate extensions that require new coordination rules.
The merchant API records payment intent before a relay schedules its stored operations. Processor workers and verified callbacks recover external outcomes, then use one posting routine for payment state and the balanced journal. Independent merchants can occupy different database groups; each merchant’s financial transaction boundary remains intact.
Read each connection in order
- syncStable payment commandMerchant application → Payment API and callbacks
- syncIntent / posting transactionPayment API and callbacks → Merchant payment / journal DB
- syncRead stored operation IDsOutbox relay → Merchant payment / journal DB
- asyncPublish operation IDsOutbox relay → Operation queues
- asyncSchedule bounded attemptsOperation queues → Processor / recovery workers
- syncLoad state / post evidenceProcessor / recovery workers → Merchant payment / journal DB
- syncCall or reconcile same keyProcessor / recovery workers → Payment processor
- asyncVerified payment evidencePayment processor → Payment API and callbacks
11Compare internal history with independent processor facts
Reconciliation compares provider operations and settlement records with local payments and journals. Match by provider account, operation reference, merchant, currency and amount. A missing local capture should go through the same outcome applier as a webhook; it should not trigger another capture call. An unmatched or conflicting fact becomes a durable review case with evidence.
Settlement is distinct from capture. Processor transfers, fees and merchant payouts require their own appropriate journals in the account model. The capture journal must balance, but recording it does not complete those later financial movements.
If the posting database is unavailable after the processor succeeds, retain pending local status and recover evidence when authority returns. Do not approve refunds from a stale replica to improve apparent availability. Synchronous replica placement and safe failover must actually meet the stated one-zone durability promise.
For regional recovery, state whether acknowledgments included a durable copy outside the region. If not, a regional loss may require restoring history and reconciling processor effects before replaying old commands. A backup that restores balances but loses operation identities can reopen duplicate-charge risk. Test restored uniqueness, evidence matching and unresolved operations as well as accounting totals.
12Observe money and unresolved age, not only HTTP success
Monitor unknown-operation age, unmatched amounts by currency, duplicate journal suppression, refund-capacity rejections, provider throttling and reconciliation backlog. Request latency matters, but ten healthy API instances do not make an unresolved capture financially complete. Measure local durable acceptance and authoritative reads against 300 ms p95, and verified-evidence receipt to local application against the five-second objective. Track external pending age separately: these targets do not bound processor completion.
Use separate, audited permissions for refunds and operational corrections. Enforce merchant isolation in queries, export paths, cache keys and reconciliation tools. Restrict provider credentials to adapters, verify callbacks and avoid logging card tokens, secrets or sensitive payloads. Immutable journals still need access controls and a deliberate retention policy.
Test crashes after processor success, after local posting and before queue acknowledgment. Race two refunds, submit two capture keys, replay different callbacks for the same capture, and restore a backup before applying old events. These tests cover separate guarantees; a balanced-total assertion alone does not prove that every external operation was recorded once.
Roll out posting changes with compatible readers and writers, a small merchant cohort and comparison against independent reconciliation. Stop unsafe posting when an invariant fails. The design deliberately accepts visible pending states and occasional unavailable decisions to preserve financial integrity and recover uncertain effects.
13Check the design against its requirements
Use the numbered requirements to check the final design. FR refers to the functional list; NFR refers to the non-functional list. Performance rows specify tests still required, not achieved benchmark results.
| Requirement | Design mechanism | Verification and remaining limit |
|---|---|---|
| FR1–2; NFR2,5: accept commands and recover status | Transactional intent/operation records, scoped keys and authoritative status; processor calls occur after commit. | Lose the local reply and recover P81. Load-test 300 ms p95 local responses; delay the processor and verify pending is returned without asserting completion. |
| FR2–3; NFR3: prevent excess financial actions | Unique full-capture claim and serialized refund reservations. | Send two capture keys and race two refunds of 2000 against a capture of 2500. The second refund cannot reserve already committed capacity. |
| FR4; NFR2–3: post and reconcile evidence | One validated posting transaction commits journal, balances, status and outgoing work; webhooks and reconciliation reuse it. | Replay distinct events for C81, compare independent processor facts and measure durable-evidence-to-local-status delay against the five-second objective. |
| NFR1,4–5: sustain and recover the chosen service | Merchant-local transaction domains, quota-bounded workers, recovery reserve and synchronous regional database failover. | Load-test the full transaction mix, fail one zone and restore operation identities. A balanced ledger alone does not prove every external charge was recovered. |
14Rapid revision
Remember: A timeout does not prove that no money moved. Keep the original operation and unresolved refund reservation until evidence resolves them.
| Concern | Complete interview mechanism |
|---|---|
| Purchase identity | Validate merchant, order amount and currency; save one matching request per merchant, operation type and request key. |
| Capture uniqueness | Allow only one capture operation per payment, even when callers send different request keys. |
| External call | Save a fixed ID for the provider operation before calling the processor. |
| Unknown result | Query or safely retry the original operation; a timeout is not proof no money moved. |
| Ledger posting | Validate evidence and atomically commit one balanced journal, status and notification work. |
| Refund limit | Lock the capture and reserve the refund amount before calling; count that amount while the result is unknown. |
| Reporting | Reporting history may lag; refund decisions check current balances in the primary transaction. |
| Reconciliation | Compare processor and settlement records with journals; apply missing results through the normal posting transaction. |
| Scale | Partition by merchant and limit provider calls; preserve operation IDs during recovery. |
Close with C81's lost response and the two competing 2000 refunds. They show why idempotency, balanced accounting and capacity reservation are separate protections rather than one generic “exactly once” feature.
15Optional prompt variant: transfer an internal wallet balance
This is an alternative prompt: Alice transfers an existing wallet balance to Bob inside the same service. It is not a merchant card capture. Choose same-currency transfers with no fees or overdraft, and keep both accounts and their journal in one database transaction domain. External deposits, withdrawals and currency exchange introduce different settlement boundaries.
A posted balance is the wallet’s committed net balance. An outgoing hold reserves part of that balance for an unfinished operation. Therefore spendable = posted balance − pending outgoing holds; expected incoming money is not spendable until posted. Store amounts as integer minor units with an explicit currency, and make every operation that changes posted or held funds follow the same account-locking rule.
Alice has $100 posted and $20 reserved for a different payment. She asks to transfer $60 to Bob, who has $15. The table displays dollars for readability; storage uses cents.
| Wallet | Posted before | Outgoing holds | Spendable before | Posted after T81 |
|---|---|---|---|---|
| Alice | $100 | $20 | $80 | $40 |
| Bob | $15 | $0 | $15 | $75 |
The $20 hold remains, so Alice has $20 spendable afterward. If another concurrent request also tries to transfer $60, it must observe the first committed transfer and reject for insufficient funds. Reading $80 before locking the accounts and spending it twice would overdraw the wallet, even if each transfer’s journal balanced.
Authenticate the caller and authorize access to the source wallet and any saved transfer result. Use a stable transfer ID scoped to the sender plus a fingerprint of source, destination, amount and currency. In the transaction, claim or read that unique identity first. Return an authorized matching completed result even if Bob’s wallet was later frozen; changed parameters conflict.
For a new transfer, use the same transaction:
Require a positive amount. Lock both accounts in stable account-ID order, then validate current debit authority, destination eligibility and matching currency.
Read posted and held balances under those locks and require Alice’s spendable funds to cover $60.
Commit a balanced journal, both balance updates, the successful transfer result and notification outbox together.
For customer wallet liabilities, debit Alice’s account by $60 and credit Bob’s by $60: the service’s total liability is unchanged. The journal and materialized balances must not be separate asynchronous commits. Never hold these locks while sending notifications.
After a timeout, retry or look up T81; do not create a new transfer identity merely because the reply was lost. TigerBeetle’s submission guidance explains stable transfer identity across retries, while PostgreSQL’s transaction example illustrates atomic account updates. If the accounts move to independent shards, an ordinary local transaction no longer covers both. Use a supported distributed transaction, or explicitly design reserved funds and a durable pending-transfer workflow with recovery. The shared distributed-workflow lesson explains those coordination choices; this variant stops at the single-database transfer.
Practise the interview questions
Say your answer aloud before opening the model answer. Then answer the follow-up and compare the reasoning.
How does the ledger differ from payment workflow state?
Reveal a model answer
Workflow records which actions are pending or known complete. Immutable journals record financial movements with balanced debit and credit totals per currency.
Interviewer follow-up
Can a posted journal be edited to correct it?
Reveal the follow-up answer
Use an auditable reversing or correcting journal rather than erasing history.
What the answer must demonstrate: Distinguish mutable workflow state from immutable balanced accounting journals.
Can two different idempotency keys capture one payment twice?
Reveal a model answer
Not if the command also claims a unique per-payment capture slot. Request-key deduplication alone only protects repeats of the same command.
Interviewer follow-up
What separately prevents duplicate local posting?
Reveal the follow-up answer
A unique financial operation/posting-kind journal applied in one transaction.
What the answer must demonstrate: Protect both one allowed capture command and one journal per financial operation.
The processor captured but the worker lost its reply. What happens?
Reveal a model answer
Keep the stored operation unknown, recover through its provider identity or lookup, and apply verified evidence locally. Do not issue a new charge to repair missing local status.
Interviewer follow-up
What if provider key retention expired?
Reveal the follow-up answer
Reconcile or review the old operation; an old local key cannot extend the processor’s deduplication window.
What the answer must demonstrate: Recover the original processor effect without issuing another charge to repair local state.
Does a balanced journal prove the payment is correct?
Reveal a model answer
No. A wrong amount or merchant can still balance. Validate the operation evidence and compare independent processor facts through reconciliation.
Interviewer follow-up
What does balance itself protect?
Reveal the follow-up answer
Internal equality of debits and credits within the journal’s currency, committed as a whole.
What the answer must demonstrate: Validate merchant, amount, currency and processor evidence in addition to balanced totals.
Two agents each refund 2000 from a 2500 capture. How is excess prevented?
Reveal a model answer
The first atomically reserves 2000 on the capture balance; the second sees only 500 available and cannot call the provider.
Interviewer follow-up
What if the first result is unknown?
Reveal the follow-up answer
Keep its reservation until definitive success or failure resolves the capacity.
What the answer must demonstrate: Reserve shared refund capacity before calls and retain it through unknown outcomes.
Why is provider-event deduplication insufficient for ledger uniqueness?
Reveal a model answer
Different events may describe the same capture. The posting identity must represent the financial operation, not only each callback.
Interviewer follow-up
How are callbacks acknowledged?
Reveal the follow-up answer
After verified durable inbox acceptance, with outcome application retried independently.
What the answer must demonstrate: Deduplicate financial operations separately from repeated provider event IDs.
Can a reporting replica approve a refund during primary failure?
Reveal a model answer
No. Its balance may omit another refund reservation. The capacity decision needs the current transactional authority.
Interviewer follow-up
What can remain available?
Reveal the follow-up answer
Known pending status or explicitly stale reporting, without pretending a new financial decision succeeded.
What the answer must demonstrate: Keep capacity decisions on current authority rather than stale reporting balances.
Why preserve operation IDs in a regional restore?
Reveal a model answer
Restored balances alone do not reveal which external effects already happened. Lost identities can cause replayed commands to charge again.
Interviewer follow-up
What establishes safe resumption?
Reveal the follow-up answer
Validated restored history, processor reconciliation and the selected durability/failover contract.
What the answer must demonstrate: Restore operation identities and reconcile external evidence before resuming commands.
What must change when the prompt asks Alice to transfer existing wallet funds to Bob?
Reveal a model answer
Recover an authorized matching retry by its stable transfer ID before applying new-transfer eligibility checks. For a new positive same-currency transfer, authenticate Alice’s debit authority, lock both accounts in stable order, recheck spendable funds after outgoing holds, and commit the balanced journal, both balances and saved result together. This internal transfer does not need a card-processor call.
Interviewer follow-up
Why do balanced journal entries alone not prevent overdraft?
Reveal the follow-up answer
Two transfers can each balance their debit and credit while both spending the same stale source balance. Serialize the current spendable-funds check with each debit. If the accounts are on different shards, use an actual distributed transaction or an explicitly pending, reserved-funds workflow.
What the answer must demonstrate: Separate internal transfers from merchant capture; protect spendable balance and journal atomicity under concurrent spends and retries.
Blank-page exercise · 45 minutes
Build the answer yourself
Design a USD 25 merchant payment, lose the capture response, and then race two USD 20 refund requests.
- Agree numbered functional and non-functional requirements, including local acceptance versus processor completion, monetary invariants and zone durability. Then trace durable payment intent, provider execution and one local ledger posting.
- Explain balanced journals, operation uniqueness and concurrent refund reservations.
- Recover unknown processor outcomes without issuing a new charge or trusting stale balances.
- Trace a timeout and a concurrent request using the actual durable records.
- Review the final architecture against every numbered requirement, including measured bottlenecks, targets still needing validation and remaining failure limits.
Check that each component and design decision follows from your requirements and workload.
Recall the key ideas
Answer from memory before opening each card. Explain why the choice works and what it costs. Revisit missed cards tomorrow.
Design a payment system and ledgerC81 may have captured USD 25.00, but its reply is lost. What should the worker recover?Recall first, then reveal
Recover C81 through supported provider lookup or idempotency, then post its verified result. A new charge could bill the customer twice.
A lost reply may hide a successful charge: recover the same operation.
Return to lessonDesign a payment system and ledgerHow much more can a captured payment refund?Recall first, then reveal
Captured amount minus completed refunds minus amounts reserved for unresolved refunds.
Unknown still reserves.
Return to lessonDesign a payment system and ledgerWhy can a balanced journal still be wrong?Recall first, then reveal
Its merchant, amount or processor evidence may be wrong. Validate those facts and reconcile them independently as well as balancing the journal.
Balance and evidence.
Return to lessonFinal revision
Summary and interview notes
Charge customers through a processor and record verified results in balanced, unchangeable journals. Reserve refund amounts before calling the processor and keep uncertain refunds reserved until their outcomes are known.
Remember these points
- Keep workflow, external operations and journals distinct.
- Prevent both duplicate capture commands and duplicate postings.
- Reserve refund capacity through uncertain outcomes.
- Apply callbacks and reconciliation evidence through the same validated posting transaction.
Interview tips
- Separate a valid balanced journal from a verified processor outcome.
- Lose C81’s reply, then race two refunds of 2000 against the 2500 capture.
Important qualifications
- Traffic and latency figures are interview assumptions, not claims about a named company's deployment.
Continue after the core interview
Explore the advanced version
The advanced lesson keeps the full detailed design. Use these sections when you want to examine the stronger requirements and failure cases.
- Full posting and refund proof
Inspect uniqueness, atomic journal lines and every reservation transition.
- Authorization, capture and void races
Provider-specific deadlines and customer-action flows require adapter detail.
- Regional disaster recovery
Zero regional loss changes acknowledgment latency and partition availability.
- Merchant partitioning and reporting
Large merchants and cross-merchant movements need a deliberate accounting boundary.
Technical references
- Stripe idempotent requestsDocuments a concrete processor retry contract, including key reuse and retention considerations.
- Stripe webhook documentationDocuments signature verification, retries, duplicate events, and ordering limitations.
- PostgreSQL constraintsDatabase mechanisms supporting unique operations and valid posting records.
- Stripe Payment Intents lifecyclePrimary provider example for a payment state machine and asynchronous outcomes; the chapter API is an abstraction, not a literal Stripe endpoint.
- PostgreSQL transaction isolationPrimary reference for locking/transaction reasoning and whole-transaction recovery.
- Stripe separate authorization and capturePrimary example of authorization eligibility, provider-specific capture deadlines, expiration and cancellation; the interview service API remains provider independent.
Practice marks stay in this browser.