1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
Your guardrails are blocking many legitimate user requests (over-refusal). How do you fix this without weakening safety?
30-second answerSay your answer out loud first, then reveal.
Steps
- Quantify: false refusal rate on the benign set; block rate in production; user complaints and rephrase rates after blocks.
- Attribute: log which guardrail fired (input classifier, output filter, model refusal) and why.
- Common causes and fixes:
| Cause | Fix |
|---|---|
| Generic toxicity model flags medical terms | Domain-tuned classifier or allow-listed contexts |
| Keyword filters (e.g. "kill process" flagged as violence) | Replace with context-aware classifiers |
| Thresholds set for consumer chat applied to a B2B tool | Per-product thresholds |
| System prompt over-cautious ("never discuss health") | Precise policies: what is allowed and how |
| Model's own refusals | Prompting with explicit allowed scope; a different model |
4. Graduated responses: answer with safety framing, ask a clarifying question, or provide high-level info without risky specifics. That's better than a blunt refusal.
5. Monitor both directions: harmful-pass rate on red-team sets must not increase.
Business framing: over-refusal drives users away and erodes trust just as unsafe outputs do. Both are product failures.
Related
Slow is fine. Stopping is the only problem.