Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q36HardSystem design

Design a customer-support agent for an e-commerce company with 1M monthly users.

30-second answerSay your answer out loud first, then reveal.

1. Clarify requirements

Always start here in a design round:

  • Channels (web chat, WhatsApp, email)? Languages (English, Hindi, Hinglish)?
  • Actions allowed: track order, cancel, return, refund, address change?
  • SLAs: first response < 3s; resolution rate target, e.g. 60–70% without a human.
  • Scale: 1M users/month ≈ maybe 300K conversations/month, with peaks during sales.

2. Architecture

Customer-support architecture: user traffic passes the API gateway and input guardrails to an intent router that sends FAQs to RAG, order actions to a tool-calling support agent with a policy engine and approval queue, and angry, legal or low-confidence cases to a human queue, with output guardrails before the reply.

3. Key design decisions

  • Router first: most traffic is simple ("where is my order?"). A cheap classifier keeps cost low and latency short.
  • Tools scoped to the authenticated user: the tool layer injects user_id from the session, never from model arguments. The agent cannot look up other users' orders even if tricked.
  • Policy in code: refund eligibility and limits live in the policy engine (Q26).
  • Handoff to humans with a summary of the conversation, steps tried and customer sentiment, so the human doesn't restart.
  • State: conversation thread with checkpoints; resumes across channels.
  • Knowledge: RAG over help articles with freshness. Re-index on change, and cite articles.

4. Reliability and safety: timeouts and fallbacks for order APIs, circuit breakers, idempotent return/refund calls (Q32), injection defences for content pasted by users.

5. Evaluation and monitoring:

  • Offline: 300+ scenario conversations (with a simulated user) covering policies and edge cases; pass^k.
  • Online: resolution rate, escalation rate, CSAT, refund anomalies, cost per resolved conversation, p95 latency.

6. Cost: small model for routing and FAQs, larger model for complex agent cases; prompt caching of system prompt and tools; response cache for top FAQs.

7. Rollout: shadow mode (agent drafts, humans send) → 5% canary on low-risk intents → expand by intent.

Little by little, you're building something great.