Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q38HardSystem design

Design the guardrails architecture for a bank's customer chatbot with under 300ms of added latency.

30-second answerSay your answer out loud first, then reveal.
A bank chatbot message passes inline deterministic checks, then injection and topic classifiers run in parallel with retrieval into a decision that returns a template response or streams LLM generation through chunk-level output checks, with tool requests going to action guardrails in code and async groundedness judging after the user sees the answer.

Latency budget (illustrative)

CheckPlacementAdded latency
Deterministic input checksInline~5ms
Injection + topic classifiersParallel with retrieval~0 extra (hidden behind retrieval)
Output chunk checksStreaming~10–30ms per chunk
Action guardrailsCode at tool boundaryNegligible
Groundedness judgeHigh-risk answers only (rates/fees) or asyncSelective

Design principles

  • Fail closed for actions (transfers, card changes); fail open with logging for low-risk informational guardrail outages (documented decision).
  • Self-hosted small classifiers near the app for low latency and data residency.
  • Template answers for sensitive topics (fraud reporting steps, complaints) instead of free generation.
  • Measure the false positive and false negative rates of each guardrail, plus p95 added latency.

Slow is fine. Stopping is the only problem.