Learnastra AI SYSTEM DESIGNAnup Rai

Complete design interview

Design a Conversational Customer-Support Agent

By Anup Rai11 min readReviewed September 2026

Interview problem: design a support assistant for a multi-tenant SaaS product. It must resolve a bounded set of issues, maintain useful conversation state, read authorized account information, and transfer responsibility to a person when needed.

A conversational agent uses earlier turns and current observations to decide how to respond or what permitted action to propose. Conversation text is not authoritative business state: “I have cancelled it” does not mean a subscription was cancelled.

This Learnastra case is hypothetical. Volumes, costs and targets below are interview assumptions, not measured operating results. Focus on the complete customer outcome rather than a sequence of polished chat messages.

1. Scope the product and define success

Clarify channels, languages, customer roles, supported products, business hours, existing help-desk integrations and permitted account actions. Start with authenticated web chat; anonymous users can receive public help but cannot read private account records.

Functional requirements

  1. Identify the request's intent and extract necessary details without inventing account identifiers.
  2. Answer product/policy questions using the correct version of the knowledge base.
  3. Read current account facts through scoped application APIs.
  4. Maintain the current goal, unresolved questions and user corrections across turns.
  5. Create a support ticket and execute explicitly allowed actions after server-side authorization and required confirmation.
  6. Transfer to a human on request, policy requirement, unresolved uncertainty or failed/unknown action outcome.
  7. Let the customer reconnect and see confirmed action status and case ownership.

Initially excluded: arbitrary shell/SQL access, autonomous high-value refunds, promises outside published policy, and treating past support tickets as automatically public knowledge.

Nonfunctional requirements

  1. Plan for 500,000 conversations/month across 50,000 customer organizations; verify this unusually high issue volume with the interviewer.
  2. Target p95 first useful response within one second for simple cached/read-only cases and p95 completed ordinary answers within five seconds. Long actions show truthful progress and a separate completion objective.
  3. Target 99.9% monthly request availability with a usable ticket/handoff path during model failure.
  4. Aim for at least 95% correct supported answers on a reviewed representative set; report uncertainty and risk slices, not only an average.
  5. Prevent cross-tenant and cross-user access; enforce role-specific action permissions.
  6. Preserve action identities and audit state across crashes, retries and human takeover.
  7. Bound per-conversation time, calls, tokens, queue depth and spending; keep unnecessary personal data out of diagnostic logs.

Reducing escalation toward 40% and increasing CSAT toward 85% are business hypotheses, not permission to retain cases the bot cannot solve. Define resolution as a verified task outcome, with recontact about the same issue measured over an agreed window. A silent or abandoned chat is not automatically resolved.

2. Estimate turns, concurrency and the human queue

Assume a 30-day month, six customer turns per conversation and three seconds of average active processing per turn.

Quantity Calculation Design implication
Conversations/day 500,000 / 30 ≈ 16,667 Conversations can span multiple days
Turns/month 500,000 × 6 = 3M Model calls should be budgeted per turn and whole conversation
Average turn arrival 3M / 2,592,000 ≈ 1.16/s Peaks depend on customer time zones and incidents
Illustrative 10× peak About 11.6 turns/s Measure bursts caused by outages
Mean in-flight turn work at that peak 11.6 × 3 ≈ 35 Separate from open browser sessions
Escalations at 38% 500,000 × 0.38 = 190,000/month Human capacity may dominate feasibility

A staffed fallback is a capacity requirement

Suppose a human case takes ten active minutes, each employee works 160 hours/month, and 75% of that time is available for case work. Capacity is 160 × 60 × 0.75 / 10 = 720 cases/person/month.

One hundred people can handle about 72,000 cases/month under these assumptions. Handling 190,000 requires 190,000 / 720 ≈ 264 full-time-equivalent staff, before additional coverage constraints. A claim that “100 agents will handle the remainder” is inconsistent with this workload unless handling time, staffing, deflection or scope changes.

The remedy is not forcing the bot to close unsafe cases. Validate the assumptions, reduce repeat demand, improve handoff context, adjust scope, and fund the queue. Publish realistic waiting expectations and route urgent cases by policy.

3. Build a baseline, then expose its limits

The first release retrieves public product help, drafts an answer and opens a human ticket when it cannot help. It does not change accounts. Measure whether customers complete the intended task compared with ordinary search and existing support.

Architecture / visual model
flowchart LR U[Customer message] --> A[Auth and conversation ownership] A --> K[Retrieve applicable help] K --> D[Draft evidence-supported reply] D --> C{Required checks pass?} C -->|Yes| R[Reply and track outcome] C -->|No or human requested| H[Create staffed handoff]
Read diagram source
flowchart LR
    U[Customer message] --> A[Auth and conversation ownership]
    A --> K[Retrieve applicable help]
    K --> D[Draft evidence-supported reply]
    D --> C{Required checks pass?}
    C -->|Yes| R[Reply and track outcome]
    C -->|No or human requested| H[Create staffed handoff]
Failure of the baseline Improvement Benefit Added obligation
Customer repeats account details Scoped account API reads More useful current context Identity and object-level access checks
Later turn contradicts an earlier one Ordered events and explicit current goal Correct handling of changes of mind Concurrency and summary maintenance
Customer needs a permitted account change Typed proposal plus action executor Complete useful workflows Approval, idempotency and reconciliation
Handoff loses history Durable ownership and context bundle Less repetition and faster recovery Human queue integration and clear responsibility
Fluent answer masks unresolved issue Outcome/recontact tracking Measures actual benefit Issue matching, sampling and delayed labels

4. Detailed architecture and authority boundaries

Architecture / visual model
flowchart TD C[Web chat or approved channel] --> G[Gateway: identity, tenant and admission] G --> E[Append deduplicated message event] E --> Q[Per-conversation ordered work queue] Q --> O[Orchestrator with owner epoch and deadline] O --> S[(Conversation events and task state)] O --> I[Intent and entity proposal] I --> K[Authorized policy retrieval] I --> A[Scoped account read API] K --> P[Build answer or action proposal] A --> P P --> T{Consequential action?} T -->|No| D[Draft answer from current evidence] T -->|Yes| V[Schema, ownership, policy and exact approval] V --> X[Durable action ledger and outbox] X --> W[Executor: recheck owner epoch and business state] W --> B[Authoritative business API] B --> REC[Confirmed result or unknown outcome reconciliation] REC --> D D --> CHECK[Evidence, action-result and privacy checks] CHECK --> R[Reply with confirmed status] O -->|Human requested or required| H[Atomically transfer ownership] CHECK -->|Cannot safely answer| H H --> HD[Help desk with evidence and action IDs] HD --> S
Read diagram source
flowchart TD
    C[Web chat or approved channel] --> G[Gateway: identity, tenant and admission]
    G --> E[Append deduplicated message event]
    E --> Q[Per-conversation ordered work queue]
    Q --> O[Orchestrator with owner epoch and deadline]
    O --> S[(Conversation events and task state)]
    O --> I[Intent and entity proposal]
    I --> K[Authorized policy retrieval]
    I --> A[Scoped account read API]
    K --> P[Build answer or action proposal]
    A --> P
    P --> T{Consequential action?}
    T -->|No| D[Draft answer from current evidence]
    T -->|Yes| V[Schema, ownership, policy and exact approval]
    V --> X[Durable action ledger and outbox]
    X --> W[Executor: recheck owner epoch and business state]
    W --> B[Authoritative business API]
    B --> REC[Confirmed result or unknown outcome reconciliation]
    REC --> D
    D --> CHECK[Evidence, action-result and privacy checks]
    CHECK --> R[Reply with confirmed status]
    O -->|Human requested or required| H[Atomically transfer ownership]
    CHECK -->|Cannot safely answer| H
    H --> HD[Help desk with evidence and action IDs]
    HD --> S

A fixed workflow is sufficient for many support intents. Introduce a bounded tool loop only where the next diagnostic step genuinely depends on observations. Model selection is evaluated configuration: compare current supported models on intent errors, complete-task quality, language coverage, latency and cost. There is no universal “stable support model” or guaranteed sub-100ms classifier.

The model never receives authority from text in a support article. Tools use credentials scoped to the authenticated customer and allowed operation. A ticket's tenant ID or order number in the prompt does not prove ownership.

5. Define the contracts and durable state

POST /conversations/{id}/messages
Body: {client_message_id, text, expected_conversation_version?}
Server derives principal and tenant from the authenticated session.

POST /conversations/{id}/handoff
Body: {request_id, reason}
Returns current owner, queue status and existing handoff ID on retry.

GET /actions/{operation_id}
Returns proposed | awaiting_approval | submitted | confirmed | rejected | unknown.
Every read enforces object-level authorization.
Record Important fields Why it is separate
Conversation tenant, permitted participants, owner, owner epoch, version Coordinates bot/human responsibility
Message event conversation, sequence, client message ID, actor, text reference Deduplicated ordered history
Task state intent, goal, constraints, unresolved fields Useful working memory
Action operation ID, exact payload hash, policy version, approval, status, external reference Durable side-effect tracking
Handoff reason, assigned queue/person, SLA, linked action IDs Transfers responsibility explicitly
Evidence source ID/version, retrieved time, permitted scope Reproducible explanation

Use a unique constraint on (conversation_id, client_message_id) so reconnect retries do not create duplicate messages. Append an event and enqueue its processing through a transactional outbox or an equivalent recoverable mechanism. Maintain a monotonically increasing version/sequence; two tabs must not silently overwrite each other's state.

Redis can cache recent history, but expiring chat state must not be the only record of a pending refund. Redis transactions serialize grouped commands; they do not provide SQL-style rollback of runtime command errors or solve duplicate message delivery by themselves.

6. Reason over evidence; execute through a separate gate

Intent identifies the customer's goal, such as investigating a duplicate charge. Entities identify details such as an invoice ID. A schema can validate these fields, but classification can still be wrong. Distinguish “the user requested a person” from “classification failed” in operational records even when both route to assistance.

Retrieve product-version and policy-date information, then read current account facts. Historical tickets may contain another customer's private information; use reviewed, permission-safe knowledge articles or properly scoped cases, not an unrestricted ticket dump.

For a proposed change, the executor verifies:

  1. The signed-in principal can act on this account and resource.
  2. The operation is allowed by current policy and business state.
  3. Required approval covers the exact action, amount/resource and relevant version.
  4. The conversation is still bot-owned at the expected owner epoch.
  5. The operation ID is stable across retries and has not already completed.
  6. The action budget/deadline permits starting the operation.

When the payload changes, old approval does not automatically carry over. A policy-allowed read can usually run without a new confirmation; a consequential change follows the product's defined approval policy.

Unknown is a real action state

Architecture / visual model
sequenceDiagram participant U as Customer participant O as Orchestrator participant L as Action ledger participant B as Business API participant H as Human support U->>O: Confirm the exact permitted change O->>L: Reserve operation ID and approved payload O->>B: Execute with stable operation identity B--xO: Response lost after possible commit O->>L: Record unknown outcome O-->>U: Outcome is being checked, no completion claim O->>B: Query status or retry under documented idempotency contract alt Confirmed completed B-->>O: External result ID O->>L: Mark confirmed O-->>U: Report confirmed result else Still unresolved O->>H: Transfer action ID, evidence and unknown status end
Read diagram source
sequenceDiagram
    participant U as Customer
    participant O as Orchestrator
    participant L as Action ledger
    participant B as Business API
    participant H as Human support
    U->>O: Confirm the exact permitted change
    O->>L: Reserve operation ID and approved payload
    O->>B: Execute with stable operation identity
    B--xO: Response lost after possible commit
    O->>L: Record unknown outcome
    O-->>U: Outcome is being checked, no completion claim
    O->>B: Query status or retry under documented idempotency contract
    alt Confirmed completed
        B-->>O: External result ID
        O->>L: Mark confirmed
        O-->>U: Report confirmed result
    else Still unresolved
        O->>H: Transfer action ID, evidence and unknown status
    end

A provider's idempotency window is not necessarily permanent. For example, Stripe documents key retention and reuse behavior; the application must retain its own business identity and reconcile retries outside the provider's retained window. A new key is a new request, not a way to safely “try harder.”

7. Conversation memory, corrections and human takeover

Keep recent turns plus a structured summary of current intent, constraints and unresolved issues. Preserve references to the original events. Exact approvals, amounts and action status remain in durable records outside a lossy summary. Memory/state design explains the distinction.

A correction such as “use the other subscription” invalidates an unexecuted proposal referring to the first one. If execution already started, establish its outcome before promising cancellation or attempting a replacement. Process per-conversation events in a defined order; merely appending Redis history atomically does not serialize the business workflow.

Human takeover atomically changes owner and increments owner_epoch. Serialize takeover with an executor's dispatch reservation in authoritative state: only a reservation that validates the current owner epoch may proceed. A reservation made before takeover is treated as potentially in flight; a later stale reservation is rejected. A separate read of the epoch followed by an uncoordinated dispatch leaves a race. An already dispatched remote operation may still complete: retain it as pending/unknown until reconciled. A fencing token only blocks stale workers where the receiving executor or resource actually enforces it.

The handoff bundle contains:

  1. Customer goal and unresolved questions.
  2. Applicable policy and account evidence with timestamps.
  3. Confirmed actions and their external result IDs.
  4. Pending/unknown actions and stable operation IDs.
  5. Reason for escalation, current owner and next expected step.
  6. Transcript access subject to the same tenant/user rules.

Do not require the customer to repeat everything, and do not silently leave both the bot and a person in charge.

8. Escalation and response checks

A model's self-reported “0.9 confidence” is not a calibrated success probability. Start with explicit handoff rules, then evaluate any learned acceptance score against human-reviewed outcomes by intent and language.

The following executable decision example consumes trusted application checks. It does not implement those checks or make an LLM judge infallible.

def reply_decision(*, human_requested, owner, draft, checks):
    if owner != "bot":
        return "not_bot_owned"
    if human_requested:
        return "handoff_requested"
    if not draft or not draft.strip():
        return "handoff_generation_failed"
    required = ("supported", "authorized", "action_status_truthful", "privacy")
    if any(checks.get(name) is not True for name in required):
        return "handoff_checks_failed"
    return "reply"

ok = dict(supported=True, authorized=True,
          action_status_truthful=True, privacy=True)
assert reply_decision(human_requested=False, owner="bot", draft="Done", checks=ok) == "reply"
assert reply_decision(human_requested=True, owner="bot", draft="Done", checks=ok) == "handoff_requested"
assert reply_decision(human_requested=False, owner="human", draft="Done", checks=ok) == "not_bot_owned"
assert reply_decision(human_requested=False, owner="bot", draft="Done", checks={}) == "handoff_checks_failed"

A multilingual policy cannot be replaced by a short English keyword list. A direct handoff request should not depend solely on intent classification; provide a visible “talk to a person” control. If no person is immediately available, show the real queue state and alternative, not an invented live transfer.

Streaming trades responsiveness against checks. A neutral acknowledgement can be immediate, but it is not proof of fast resolution. Buffer action-result claims and sensitive content until required validation passes. Measure time to useful content, complete response and action completion separately.

9. Reliability, quality and rollout

Failure Response Evidence to retain
Model timeout/refusal/malformed output Bounded retry where appropriate, then a defined fallback Failure category and request/configuration ID
Account API unavailable Explain inability to verify; avoid guessing Tool error and unresolved account facts
Missing policy Clarify or hand off Search scope and missing source
Unknown write outcome Reconcile stable operation ID Approved payload and external reference
Bot/human race Stop stale work; reconcile in-flight actions Owner epochs and dispatch timeline
Queue overload Admission controls and realistic waiting status Arrival/service rates and oldest case age

Test full conversations: corrections after summarization, two concurrent tabs, impersonation attempts, revoked roles, stale policy, tool failures, repeated approval and human takeover during a write. Calibrate model-assisted review against human labels. Random samples estimate common behavior; separately inspect severe events and rare slices without presenting oversampled results as population rates.

Metric What it means Common misleading interpretation
Verified resolution Issue outcome meets the task's definition Conversation closed means solved
Recontact rate Same issue returns within a defined window Every later question is a failed resolution
Escalation rate Fraction needing a person Lower is always better
CSAT Respondents' reported satisfaction Represents all customers or proves factual correctness
Supported-answer rate Reviewed answers meet evidence rubric Politeness or citation presence alone establishes support
Handoff latency Time until a capable human owns/responds Ticket creation equals human response

A hypothetical 94.3% correctness result fails a 95% point-estimate gate even if CSAT is 87%; also report case count and uncertainty. Predefine the actual statistical release rule. Launch read-only intents first, then bounded actions with rollback and business-state repair procedures. A scoped SOC 2 examination requires organizational controls and evidence; a diagram or model guardrail is not certification.

10. Cost and benefit: include the people

Use aggregate whole-conversation usage, not the price of one short completion. Illustrative assumptions:

Item Monthly calculation Amount
Automation calls/tools/review allowance 500,000 × $0.026 $13,000
Application infrastructure allowance Assumed $2,000
Human escalations 190,000 × $5 $950,000
Combined modeled cost $965,000
All-human comparison 500,000 × $5 $2,500,000
Arithmetic difference $2.5M − $965K $1,535,000/month

These are assumed unit costs, not current provider quotes or realized savings. Five dollars corresponds to ten active minutes at $30/hour before additional overhead. The staffing-capacity calculation must still hold. Include help-desk licenses, failed calls, incident recovery, staffing coverage, labeling and displaced work in the actual budget.

Reducing human handling time from ten to eight minutes for the same 190,000 cases saves 190,000 × 2 / 60 ≈ 6,333 active hours/month under these assumptions. A better handoff may therefore matter more than a tiny model-token discount. It is only a benefit if quality and actual operating capacity improve.

Interview follow-ups

1. How is this different from an FAQ chatbot? It maintains task state, reads authorized live facts, executes bounded actions and transfers responsibility. Each addition needs a separate contract and evaluation; multi-turn text alone does not provide these properties.

2. How would you choose models? Compare current supported configurations on representative complete conversations and risk slices. Use a smaller classifier only if its routing mistakes and extra call cost are acceptable. Do not infer latency or quality from a model's name or parameter count.

3. Escalation falls but recontact rises. Is that progress? Possibly false resolution or abandonment. Review same-issue outcomes, acceptance thresholds and missed handoff requests. Restore safe scope while investigating.

4. Why keep action state outside the conversation summary? Summaries can omit or misstate an exact amount, approval or result. Durable typed records support reconciliation, authorization and recovery even if chat history expires.

5. Can takeover cancel an in-flight refund? Not necessarily. Changing ownership prevents new dispatches through the enforced gate. Reconcile already dispatched work before the human repeats or compensates it.

6. What is the first launch milestone? One valuable issue class with authoritative evidence, verified outcomes and a staffed fallback. Expand only after its action/permission failures and queue capacity are understood.

60-second interview answer

I would separate conversation memory, authoritative account data and action execution. Each turn uses current identity, policy and account facts; proposed changes pass a server-side gate with exact approval and stable operation identity. Unknown outcomes are reconciled before retrying. Human takeover transfers ownership and preserves pending actions. I would launch a narrow set of intents, staff the fallback queue, and measure verified resolution, recontact and total cost rather than ticket deflection alone.

Remember: Understand → Verify → Propose → Authorize → Confirm → Hand off when needed.

For a deeper action-recovery exercise, continue to customer-support automation.

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← Design an Enterprise Knowledge Assistant
NEXT LESSONDesign a Source-Verified Financial Research Assistant →

Explore the diagram