Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q34IntermediateSystem design

Design a real-time voice AI agent (e.g. a phone-based customer support bot).

30-second answerSay your answer out loud first, then reveal.
Voice agent pipeline: caller audio flows through VAD and endpointing, streaming STT, a streaming LLM with CRM and order tools, phrase chunking and streaming TTS back to the caller; the LLM logs transcripts and traces, and caller speech during bot audio triggers barge-in.

Latency budget (illustrative)

StageTarget
Endpointing (deciding the user finished)200–500ms
Final STT~100–300ms
LLM time to first token200–500ms
TTS time to first audio100–300ms
Network50–150ms

Techniques: stream everything; start TTS on the first sentence; use fast models for routine turns; filler phrases while tools run ("Let me check your order..."); pre-fetch customer data at call start (caller ID lookup).

Voice-specific design

  • Turn-taking: tune endpointing (too aggressive cuts users off, too slow feels laggy); semantic end-of-turn detection.
  • Barge-in: stop playback immediately and discard unspoken text from context.
  • Speech-friendly output: no markdown or lists; spell out numbers; confirm critical details ("That's order 4-5-7-2, correct?").
  • Accents, noise, code-switching (Hindi/English): evaluate STT word error rate per segment.
  • Escalation to a human with a transcript summary.
  • Compliance: call recording consent, PII handling, payment card data must not pass through the LLM.

Metrics: containment rate, average handle time, latency percentiles, interruption rate, CSAT, escalation correctness.

This is what real progress feels like.