Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q28IntermediateScenario

Your guardrails are blocking many legitimate user requests (over-refusal). How do you fix this without weakening safety?

30-second answerSay your answer out loud first, then reveal.

Steps

  1. Quantify: false refusal rate on the benign set; block rate in production; user complaints and rephrase rates after blocks.
  2. Attribute: log which guardrail fired (input classifier, output filter, model refusal) and why.
  3. Common causes and fixes:
CauseFix
Generic toxicity model flags medical termsDomain-tuned classifier or allow-listed contexts
Keyword filters (e.g. "kill process" flagged as violence)Replace with context-aware classifiers
Thresholds set for consumer chat applied to a B2B toolPer-product thresholds
System prompt over-cautious ("never discuss health")Precise policies: what is allowed and how
Model's own refusalsPrompting with explicit allowed scope; a different model

4. Graduated responses: answer with safety framing, ask a clarifying question, or provide high-level info without risky specifics. That's better than a blunt refusal.

5. Monitor both directions: harmful-pass rate on red-team sets must not increase.

Business framing: over-refusal drives users away and erodes trust just as unsafe outputs do. Both are product failures.

Slow is fine. Stopping is the only problem.