Learnastra AI SYSTEM DESIGNAnup Rai

Concept · Understand the mechanism

Real-time voice agents: conversation, timing, and trustworthy actions

By Anup Rai19 min readReviewed September 2026

A real-time voice agent is a system that receives spoken input, interprets it, performs permitted work, and produces spoken responses during an ongoing conversation. It combines a media pipeline, conversation control, model inference, and application state. Good speech quality alone does not establish task correctness.

Real time means timing affects usefulness. In a conversational system, late audio usually degrades the interaction rather than violating a hard physical deadline, so the media path is commonly treated as soft real time. A booking or payment operation still needs ordinary correctness, authorization, and recovery guarantees.

This chapter builds the concepts first, then designs a voice service for appointment scheduling. Workloads, latency targets, and cost rates are explicit interview assumptions. Provider references were checked on September 24, 2026.

Three useful architecture patterns

1. Cascaded speech pipeline

Automatic speech recognition (ASR), also called speech-to-text (STT), converts speech into text. A language model or application workflow processes that text. Text-to-speech (TTS) synthesizes the spoken response.

Architecture / visual model
flowchart LR I[Microphone or phone audio] --> A[Streaming ASR] I --> V[Speech activity and turn detection] A --> C[Conversation controller and stable transcript] V --> C C --> L[Language model and tool workflow] L --> T[Streaming TTS] T --> P[Playback buffer and speaker] P --> C
Read diagram source
flowchart LR
    I[Microphone or phone audio] --> A[Streaming ASR]
    I --> V[Speech activity and turn detection]
    A --> C[Conversation controller and stable transcript]
    V --> C
    C --> L[Language model and tool workflow]
    L --> T[Streaming TTS]
    T --> P[Playback buffer and speaker]
    P --> C

Only the ASR/LLM/TTS text interfaces are textual: the microphone and speaker paths remain audio. VAD and transcription can operate concurrently; turn detection may combine transcript and acoustic evidence.

2. Native speech-to-speech

A model consumes audio and generates audio without requiring a separate application-level ASR → text-model → TTS chain. It can still expose transcripts, accept text, and call tools. Native speech does not remove the need for a controller, protected tool adapters, or playback tracking.

3. Voice frontend with a delegated backend

A conversational model manages listening and speaking while a separate backend performs reasoning and tool work. The frontend can remain responsive during a long lookup. Task authority, durable state, and cancellation still belong to the application.

Architecture / visual model
flowchart LR U[User audio and interruptions] <--> F[Voice frontend] F <--> C[Application controller] C <--> B[Reasoning or workflow backend] B <--> T[Scoped tools and authoritative services] C <--> S[Durable task and delivery state]
Read diagram source
flowchart LR
    U[User audio and interruptions] <--> F[Voice frontend]
    F <--> C[Application controller]
    C <--> B[Reasoning or workflow backend]
    B <--> T[Scoped tools and authoritative services]
    C <--> S[Durable task and delivery state]

Hybrid designs may also add independent transcription to a native-audio system for a particular evaluation or evidence requirement. This is optional; many providers already expose transcripts, which have their own accuracy and synchronization limits.

Decision Cascade Native speech-to-speech Delegated voice frontend
Component choice Swap ASR, model, and voice separately More coupled provider/model behavior Choose voice and backend independently within the integration contract
Debugging Inspect text at intermediate boundaries Inspect audio, events, transcripts, and tool results Correlate frontend turns with backend tasks
Timing Stream and overlap stages while respecting dependencies Integrated audio path; still measure actual latency Responsive conversation can overlap slow backend work
Speech expression Depends on retained context and TTS controls Can use acoustic cues directly Depends on frontend and information passed from backend
Tool correctness Enforced by application Enforced by application Enforced by application
Cost Mix of minutes, characters, tokens, and hosting Provider-specific audio/text/session accounting Voice session plus backend charges
Main integration risk More interfaces and queue boundaries Provider-specific state/turn semantics Stale results and frontend/backend coordination

No pattern is automatically compliant, cheapest, fastest, or best for every language. Select it using representative calls, control requirements, available integrations, and complete operating costs.

The media and conversation pipeline

Audio representation and transport

Term Definition Design implication
Sample rate Samples per second per channel Match the receiver's required rate or resample correctly
Bit depth Bits used for each uncompressed sample Determines PCM precision and byte rate
Channel Separate audio stream, such as caller and agent Preserve speaker direction when useful
Codec Encoding/decoding method Opus, PCM, and G.711 payloads are not interchangeable
Packetization Grouping audio into transport units Large chunks add buffering delay; tiny chunks add overhead
Jitter Variation in packet arrival timing A buffer trades playout stability against delay
Echo cancellation Suppression of speaker output captured by the microphone Helps prevent the assistant from reacting to itself

WebRTC provides interactive media transport, negotiation, and related mechanisms. It commonly carries audio over UDP, with alternatives such as relayed/TCP paths when required. It does not guarantee a particular latency. WebSocket provides ordered reliable messages over TCP; loss can delay later bytes, but it remains a practical server-to-provider or telephony integration. SIP is signaling for establishing/managing sessions, not an audio codec; media is carried through the negotiated path.

For example, Twilio Media Streams uses mono 8 kHz μ-law payloads. That is this interface's contract, not a rule that every SIP call has the same format. Decode and resample for a provider that requires PCM. Raising the sample rate does not restore information lost in narrowband capture. Twilio media format.

Worked size: mono 16 kHz, 16-bit PCM is 16,000 × 2 = 32,000 bytes/second, or 1.92 MB/minute before transport encoding and metadata. A 20 ms chunk contains 640 bytes. A continuous second channel doubles those raw sizes; compression and silence handling change actual storage/network usage.

Speech activity is not the end of a thought

Voice activity detection (VAD) estimates whether audio contains speech. Endpointing decides when a speech segment or user turn has ended, according to the application's contract. Turn-taking also decides when to yield, continue, backchannel, or respond.

A caller says, “Move the appointment to Tuesday … actually, Wednesday afternoon.” A short silence threshold can close the turn after Tuesday. A longer threshold increases delay on short answers. Learned turn detection can use acoustic and semantic signals, but it can still misclassify the pause.

Strategy Useful when Tradeoff
Fixed silence threshold Simple tasks with predictable pauses Delay versus premature turn endings
Learned semantic/acoustic detector Varied conversational speech Model/language sensitivity and tuning effort
ASR-integrated endpoint event Provider exposes useful turn state Vendor-specific events and semantics
Explicit push-to-talk/button Accessibility or controlled workflows Less natural hands-free interaction
Hybrid Different call states need different behavior More policy and testing complexity

LiveKit currently supports acoustic/semantic turn detection, provider-side turns, STT endpointing, and explicit control, along with adaptive interruption handling. Silero VAD is an available component, not a universal prerequisite. LiveKit turn handling.

Deepgram Flux distinguishes confirmed and eager turn events, including a resumed-turn signal. Eager generation can reduce delay but produces discarded work when the caller continues. AssemblyAI also documents end-of-turn semantics and timing controls; an endpoint event is not simply “the transcript is formatted.” Flux documentation, AssemblyAI turn detection.

Streaming recognition and synthesis

A partial transcript is provisional and may change. A final transcript is stable under the recognizer's current result contract; it is not proof that the words were heard correctly. Word timing, speaker labels, and confidence availability vary by provider.

For TTS, distinguish request-to-first-byte from request-to-playable-audio and from the user's end-of-turn to audible response. “TTFA” is used for different intervals, so always name the start and end timestamps. A small model-inference figure excludes other delays. ElevenLabs latency definitions.

Do not necessarily synthesize the first isolated token. Buffer a short, meaningful phrase when needed for pronunciation, prosody, or safety. Streaming later phrases can overlap earlier playback. Never speak “Your booking is confirmed” before the authoritative booking result exists.

Interruptions require two state machines

Barge-in is a user interrupting while the assistant is speaking. A backchannel such as “mm-hmm” may acknowledge rather than interrupt. Noise or echo may produce a false speech trigger.

Handle a real interruption by:

  1. Stopping or attenuating playback promptly according to the interaction policy.
  2. Canceling obsolete generation and clearing queued audio where supported.
  3. Recording the best available played-audio boundary, including uncertainty, and updating conversation history accordingly.
  4. Preserving the new user input, including a short prefix so initial sounds are not clipped.
  5. Deciding separately whether any backend task should continue, cancel, or reconcile.
Architecture / visual model
sequenceDiagram participant U as Caller participant P as Playback and turn controller participant M as Voice generation participant B as Booking workflow M->>P: Audio chunks for response 7 P-->>U: Play first part U->>P: Interrupt with a correction P->>P: Clear pending audio and record played boundary P->>M: Cancel obsolete speech and synchronize history P->>B: Evaluate correction against current operation state alt Booking not submitted B-->>P: Replace stale proposal after validation else Booking already submitted B-->>P: Reconcile outcome before any replacement end
Read diagram source
sequenceDiagram
    participant U as Caller
    participant P as Playback and turn controller
    participant M as Voice generation
    participant B as Booking workflow
    M->>P: Audio chunks for response 7
    P-->>U: Play first part
    U->>P: Interrupt with a correction
    P->>P: Clear pending audio and record played boundary
    P->>M: Cancel obsolete speech and synchronize history
    P->>B: Evaluate correction against current operation state
    alt Booking not submitted
        B-->>P: Replace stale proposal after validation
    else Booking already submitted
        B-->>P: Reconcile outcome before any replacement
    end

Canceling speech is not rolling back a booking. Keep separate response_id, playback_generation, and operation_id values. Drop late audio chunks from an obsolete playback generation; keep the durable operation result even if its original spoken response was interrupted.

For OpenAI Realtime, the documented WebRTC/SIP path manages output buffering and interruption truncation on the server. A WebSocket client manages playback and sends conversation.item.truncate at the played boundary. Do not delete the whole conversation or assume generated audio was heard. Realtime interruption contract.

In Twilio bidirectional streams, clear flushes buffered audio and also causes pending mark events to return. A returned mark after a clear is therefore not, by itself, evidence that the associated audio played. Track cleared versus completed chunks. Twilio buffer events.

Latency follows the critical path

Streaming overlaps work on different chunks. It does not turn the latency of a causally dependent ASR → model → TTS response into max(all stages). That intuition can describe ideal steady-state throughput in some pipelines; first-response latency follows the actual dependencies.

Worked trace for one straightforward turn

Assume the caller finishes speaking at time zero. ASR has already processed earlier audio. Endpointing and final recognition run concurrently after the final inbound audio arrives.

Event Increment or dependency Time from caller's final speech
Final inbound audio available 50 ms transport 50 ms
Turn and needed transcript ready max(220 ms endpointing, 100 ms ASR tail) 270 ms
First useful model token 260 ms 530 ms
Speakable phrase assembled 90 ms 620 ms
First playable synthesized chunk available 130 ms 750 ms
Chunk reaches client 60 ms 810 ms
Playout buffer/device begins audio 40 ms 850 ms

These are illustrative durations for one trace, not vendor guarantees or percentile measurements. 50 + max(220, 100) + 260 + 90 + 130 + 60 + 40 = 850 ms. Do not add component p95 values and call the result the system p95; measure end-to-end percentiles under representative load.

If a required tool adds 600 ms before the substantive answer can be generated, the useful answer takes longer. An honest acknowledgment can improve the experience but must be measured separately; filler audio does not make the result arrive sooner.

Optimization order

  1. Instrument timestamps and identify the measured critical path.
  2. Remove unnecessary buffering and sequential network hops.
  3. Tune turn decisions against premature endings and long pauses.
  4. Stream model output and synthesize useful phrases without waiting for the entire answer.
  5. Consider speculative read/draft work on stable partial input; count discarded work and cancel obsolete results.
  6. Reduce model/tool latency while preserving required quality and authorization.
  7. Load-test tail latency, packet loss, jitter, and provider quotas.

Speculation must not commit a consequential action from an unfinished utterance. A small average gap also does not establish natural conversation for people who pause, speak slowly, use assistive technology, or switch languages.

Current implementation options

These are examples to evaluate, not a ranking or a claim of identical support.

Layer Current examples What to verify
Pipeline orchestration LiveKit Agents; Pipecat Versioned turn/interrupt APIs, supported adapters, lifecycle and backpressure
Streaming ASR Deepgram Flux; AssemblyAI streaming; ElevenLabs Scribe v2 Realtime Exact language, endpoint, channel, confidence, and turn-event support
Streaming TTS Cartesia Sonic 3.6; ElevenLabs Flash or conversational v3 Model-specific latency, pronunciation, voice rights, and audio format
Integrated native audio OpenAI gpt-realtime-2.1; Google gemini-3.8-live; Amazon Nova 2 Sonic Current model lifecycle, regions, tools, interruption and context behavior
Delegated conversation OpenAI gpt-live-1 with a separate backend Session duration billing, delegation contract, task/result synchronization
Managed service Hosted voice platforms or cloud contact-center integrations Full charges, data controls, provider access, export and handoff support

Pipecat's current documentation describes frame processors, pipelines, and workers; pin the framework release rather than mixing older examples with new lifecycle APIs. Pipecat architecture. Verify model options at Cartesia and ElevenLabs.

OpenAI: the current Realtime quickstart uses gpt-realtime-2.1; its configurable reasoning can affect latency. Browser access uses short-lived session credentials from a trusted server, with permanent API keys kept server-side. GPT-Live separates speech interaction from backend work; application controls still govern functions and task progress. Realtime guide, model details, GPT-Live guide.

Google: Gemini 3.8 Live became the stable default in September 2026. Migration requires more than changing the name: the standard model omits thinking_level, defaults to nonblocking function behavior, and uses audio responses with transcription when text is needed. Its Extended Thinking variant has a different configuration contract. Gemini 3.8 Live model and migration notes.

AWS: use Nova 2 Sonic documentation for a current design. The original Nova Sonic model's listed end-of-life date was September 14, 2026. Nova 2's May refresh was an in-place deployment, illustrating why evaluations must also watch provider updates behind an unchanged model ID. Original lifecycle, Nova 2 release notes.

Production concerns that change the design

Identity and critical values

Hearing “yes” is not identity verification. Caller ID, a familiar voice, and a high ASR confidence score do not establish account ownership. Use the approved account-verification flow, with alternative input such as a keypad or secure link when appropriate.

Confirm critical slots—date, timezone, amount, address, customer identity—using structured values and current source data. Voice repetition can repeat the same error, so offer spelling, a screen, or another appropriate channel when ambiguity persists. ASR errors are important; they are not universally the dominant failure for every application.

Tool calls and background results

Give a short truthful acknowledgment when work takes time. Avoid repeated filler, false completion, or claiming a tool is running when it never started. Set a deadline and offer a handoff or supported follow-up when the wait exceeds it.

Bind a tool result to the task and proposal version. If the caller changes the date during a lookup, the old result may be irrelevant. Read-only work can often be canceled/discarded; a submitted mutation requires outcome reconciliation. Native audio and cascaded systems both need these controls.

State, context, and reconnects

State Owner Recovery rule
Media packets/playout buffer Transport/runtime Bound the buffer; do not replay old speech blindly
Turn and transcript revisions Conversation controller Preserve final versus provisional state and speaker identity
Spoken delivery Playback controller Track generated, queued, played, and cleared portions
Task and proposal Application backend Persist exact scope, versions, and required confirmations
External operation Business adapter/ledger Retain stable identity and known/unknown outcome
Durable user preferences Authorized memory store Scope and correct separately from transient conversation

A disconnected call is not necessarily a failed booking. On reconnect, verify the session/user association and load the current operation state before proposing a retry. Keep call identity distinct from task identity; one task may span more than one call.

Long conversation history consumes context. The transport may stream new audio incrementally while the provider reuses retained history for inference. That is different from the client uploading the entire recording every turn. Summarize where appropriate, but preserve structured critical values and authoritative task state outside the summary. See agent memory and state.

Privacy, accessibility, and operations

  • Apply the required disclosure, recording/processing permissions, retention, and data access policy for the actual users and jurisdictions.
  • Keep permanent provider credentials off clients; restrict session credentials and privileged tools server-side.
  • Treat transcripts, screenshots, and spoken instructions as potentially untrusted inputs.
  • Support text/keypad alternatives, repetition, language fallback, and a reachable human path.
  • Bound output queues and stop obsolete audio. Under sustained overload, admit fewer sessions or transfer rather than accumulating seconds of unusable audio.
  • Drain long-lived calls during deployment where possible; record recovery state before forced termination.

Evaluate conversation and task success separately

Metric Definition or measurement Common misleading shortcut
Task completion Verified required outcome under domain policy Fluent final response
Word error rate (substitutions + deletions + insertions) / reference words Treating low overall WER as correct names/codes
Critical-slot accuracy Correct important structured values Averaging them away among easy words
Premature turn endings Turn closures that cut off intended speech Only measuring average response gap
Interruption response User onset to effective playback stop/yield Time to cancel model generation only
Useful response latency End of relevant user turn to substantive audible response Time to a generic filler phrase
False interruption rate Unnecessary yields under an explicit annotation rule Assuming every VAD trigger means a real interruption
Recovery quality Correct state after dropout, timeout, correction, or reconnect Counting reconnected sessions alone
Full cost All attempts, call time, tools, review and operations TTS price alone

For WER, insertions can make the value exceed 100%; specify normalization and evaluation language. Include names, digit strings, accents, noise, overlapping speech, long pauses, and code-switching. Measure the same business task in text and voice to isolate where the voice path adds errors.

The March 2026 τ-Voice paper evaluates 278 grounded tasks with full-duplex interaction. It reported substantially lower task completion for the tested voice systems than its text baseline, especially with noise and varied accents. Those results belong to its model versions and simulator protocol; they are not the capability of every September 2026 voice model. Primary benchmark paper, agent evaluation.

Billing and complete economics

Realtime voice products do not all bill the same way. Some charge duration, some characters or generated speech, and some audio/text tokens. Retained history, caching, optional transcription, and tool calls can change the total.

OpenAI Realtime conversational responses bill modality-specific tokens with retained history and eligible caching; input transcription has separate accounting. GPT-Live bills active session duration, including silence/backend waits, plus backend usage. Its reported duration updates are cumulative snapshots, not increments to sum. Close completed sessions and record final usage. Voice accounting.

Compare managed and self-operated options using the same task quality and support scope. If an operated stack adds $3,000/month in fixed work and saves $0.03/minute, the simple crossover is 100,000 minutes/month, not a universal 10,000-minute rule. Real contracts, staffing, usage shape, and failures change it.

Interview design: appointment scheduling by phone

Assumptions: a home-services company receives 3,000 calls/day over an eight-hour service window, across 22 days/month. The agent handles four minutes/call on average. Callers can ask about availability and book or reschedule an appointment. Human staff handle unsupported or ambiguous cases.

Functional requirements

  1. Accept inbound calls and establish the appropriate user/account context.
  2. Explain capabilities and provide required notices/options.
  3. Understand the requested service, location, time window, and timezone.
  4. Retrieve current eligible availability and quote applicable conditions.
  5. Prepare and confirm the exact booking or reschedule proposal.
  6. Commit through the scheduling service and report its authoritative result.
  7. Handle interruptions, corrections, retries, disconnects, and human transfers.
  8. Preserve task/effect state and send a confirmation only through an authorized channel.

Non-functional requirements

  1. Correctness: no double booking, unauthorized account change, or false confirmation.
  2. Responsiveness: initial target of p95 useful responses within 1.2 seconds for turns without required external work; measure tool-dependent turns separately.
  3. Interruption: initial target of p95 effective yield within 250 ms of real user interruption, tested on supported call paths.
  4. Capacity: support a sustained 3× arrival peak and provider/session quotas.
  5. Durability: recover operation state across worker and connection failures.
  6. Privacy/accessibility: protect recordings and support alternate input/human assistance.
  7. Economics: measure complete cost per attempted and verified completed task.

Basic design and first flaw

Phone stream → ASR → model → booking API → TTS.

Flaw: “Tuesday … actually Wednesday” creates a booking from a provisional transcript. The caller then hears “confirmed” before the API finishes.

Repair: a conversation controller owns transcript/turn versions; a backend prepares a structured proposal and obtains the required confirmation before commit. A brief acknowledgment may stream early, but confirmation speech is released only after the authoritative result.

Detailed architecture

Architecture / visual model
flowchart TD P[Phone or browser caller] <--> G[Media gateway and codec conversion] G --> A[Streaming recognition and turn events] A --> C[Conversation controller and revision state] C <--> L[Language model and bounded response generation] C <--> I[Identity and account verification] C --> D[Domain workflow and current availability] D --> Q[Exact booking proposal and confirmation state] Q --> K[Authorized commit with version and operation key] K <--> S[Authoritative scheduling service] K --> O[Known result or reconciliation queue] O --> C C --> T[Streaming speech and output-generation IDs] T --> B[Bounded playback buffer and delivery tracking] B --> G B --> C C <--> H[Human handoff with verified context] C <--> R[Durable task, proposal and operation records]
Read diagram source
flowchart TD
    P[Phone or browser caller] <--> G[Media gateway and codec conversion]
    G --> A[Streaming recognition and turn events]
    A --> C[Conversation controller and revision state]
    C <--> L[Language model and bounded response generation]
    C <--> I[Identity and account verification]
    C --> D[Domain workflow and current availability]
    D --> Q[Exact booking proposal and confirmation state]
    Q --> K[Authorized commit with version and operation key]
    K <--> S[Authoritative scheduling service]
    K --> O[Known result or reconciliation queue]
    O --> C
    C --> T[Streaming speech and output-generation IDs]
    T --> B[Bounded playback buffer and delivery tracking]
    B --> G
    B --> C
    C <--> H[Human handoff with verified context]
    C <--> R[Durable task, proposal and operation records]

Choose a cascaded implementation for this worked design because separately testing critical slot recognition and response wording is useful. Native audio or a delegated frontend can use the same domain boundary if evaluation shows a better result. Compliance and tool correctness do not follow from this choice alone.

Failure review and decision costs

Failure Repair Cost/benefit
Old availability arrives after a correction Bind results to proposal/turn revision Bookkeeping; prevents stale choices
Another caller takes the slot Conditional/transactional booking in the scheduling service Retry/alternative selection; prevents double allocation
Timeout after booking succeeds Query/reconcile using the logical operation identity Extra durable state; avoids duplicate retry
Caller interrupts after submission Stop speech, resolve submitted effect, then handle requested change More workflow states; accurate recovery
Buffered old audio arrives late Playback-generation guard and buffer clear Runtime complexity; avoids contradictory speech
Name/code repeatedly misheard Alternate input and verified account flow Longer interaction; less consequential ambiguity
Human transfer fails Retain queue ownership and offer supported fallback Staffing/queue cost; avoids silent abandonment

For rescheduling, determine whether the service supports an atomic move or reserve-then-release workflow. Canceling the old appointment before securing the new one can leave the caller with neither. Define compensation and explain any unresolved state rather than assuming two API calls form one transaction.

Capacity and storage

Arrival rate is 3,000 / 28,800 = 0.104 calls/second. With a four-minute mean agent duration, average active calls are 25. A sustained 3× arrival peak gives 75 active calls on the same duration assumption. At 70% planned occupancy, begin with capacity for 108 concurrent agent sessions, then test tail call lengths, transfers, connection setup, and provider quotas.

Do not divide every model's token throughput by this session count and call it sufficient capacity. ASR streams, TTS bursts, model requests, media bandwidth, and tool quotas each have different demand. Long human waits should release unnecessary model resources while preserving the call/handoff state.

There are 66,000 calls and 264,000 agent minutes/month. Retaining one continuous 16 kHz, 16-bit mono PCM track would use about 506.9 GB/month, before metadata and backups. A second track doubles it; compressed recordings differ. Retention is a policy decision, not an automatic requirement to save every call forever.

Cost assumptions and sensitivity

Assume 10% of calls transfer to a human for six additional minutes. That is 39,600 human minutes, or 660 hours. At 120 productive hours/person/month, this requires 5.5 people of capacity before absence and peak coverage.

The rates below are illustrative unit costs, not current quotes from any named vendor. They make the accounting explicit.

Component Calculation Monthly cost
Telephony, including extra transfer time 303,600 minutes × $0.015 $4,554.00
ASR on the agent portion 264,000 minutes × $0.005 $1,320.00
Model/tool usage 66,000 calls × $0.012 $792.00
Synthesized speech, assumed 60% of agent duration 158,400 minutes × $0.018 $2,851.20
Runtime and monitoring Assumed fixed cost $800.00
Maintenance 24 hours × $100 $2,400.00
Human handling 660 hours × $50 $33,000.00
Total $45,717.20

This is about $0.69 per attempted call. At 90% verified completion across automatic and human-assisted outcomes, it is about $0.77 per completion. Add initial development, applicable taxes/fees, storage beyond the assumed infrastructure allowance, and exceptional incident work where relevant. If a vendor bundles ASR/TTS/model time, use its combined rate rather than double-counting the components.

Reducing human transfer can materially change cost, but a false booking is not a successful cost saving. Optimize confirmed task quality, transfer quality, and full cost together.

Closing remarks

“I would separate conversational responsiveness from booking authority. The media path handles turn detection, audio queues, and interruptions; the backend validates and commits exact proposals, then reconciles uncertain outcomes. I would launch with a narrow service scope and expand after measuring slot accuracy, task completion, latency tails, interruption recovery, and human workload on representative phone calls.”

Interview questions and answer notes

  1. What distinguishes VAD from endpointing? VAD detects speech activity; endpointing decides a segment/turn boundary using the chosen policy.
  2. Why does streaming not make total latency the slowest stage? The first useful response still depends on earlier information becoming available.
  3. Is first audio byte a sufficient user-experience metric? No. Track audible and substantive response timing at the client.
  4. Does native speech-to-speech remove tool authorization work? No. Application boundaries still control effects.
  5. What does a delegated voice frontend add? It can converse while backend work proceeds, at the cost of explicit task/result coordination.
  6. Does canceling TTS undo a tool call? No. Speech and business-operation state must be handled separately.
  7. Why track played audio? Generated or queued words may never have reached the caller.
  8. Does a Twilio mark always mean speech played? No. Clearing the buffer also returns pending marks.
  9. Why is a final transcript not ground truth? Recognition errors remain even after revision stops.
  10. Can caller ID or voice familiarity authorize an account change? No. Use the approved verification process.
  11. Does resampling 8 kHz audio recover missing high frequencies? No. It changes representation, not the original captured information.
  12. Can WER exceed 100%? Yes, because insertions contribute to the numerator.
  13. How should reconnect affect booking retries? Restore authoritative operation state before deciding whether a retry is safe.
  14. What determines managed-versus-operated crossover? Comparable quality, all fixed/variable costs, workload, and support obligations.
  15. What should the interview close emphasize? Correct task outcomes, measured timing, interruption/reconnect recovery, and complete economics.

Final notes

Recall hear → decide the turn → interpret → act correctly → speak → track what was heard. Tune each boundary using real evidence. Natural timing improves the conversation; durable, authorized, verified operations make it a dependable service.

Next: Multimodal generation.

Your notes

Write the decision you would make and the uncertainty you would investigate next. Saved only in this browser.

PREVIOUS LESSON← Safety and governance for tool-using agents
NEXT LESSONMultimodal Generation: From Prompt to Publishable Media →

Explore the diagram