1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
How do you evaluate the guardrails themselves? How do you measure bypass resistance?
30-second answerSay your answer out loud first, then reveal.
Evaluation dimensions
| Dimension | Metric |
|---|---|
| Detection quality | Recall (harmful caught), precision (benign allowed), per category |
| Over-refusal | False positive rate on benign look-alikes |
| Robustness | Bypass rate under adversarial transformations |
| Coverage | Languages, modalities (text in images), channels (documents, tool outputs) |
| Performance | Added latency p95, cost per request, availability |
| Drift | Performance over time as user behaviour and attacks change |
Adversarial transformation suite
- Paraphrase / synonym substitution.
- Encoding (base64, ROT13), character tricks (zero-width spaces, homoglyphs, spacing).
- Translation into other languages or code-mixed text.
- Payload splitting across turns or documents.
- Role-play / hypothetical framing.
- Injection hidden in markup (HTML comments, white text, alt text).
Interpretation: a guardrail with 98% recall on the clean dataset but a 40% bypass rate under encoding needs fixes (normalise or decode inputs before classification, add more layers). Report both numbers.
Related
- Previous: Q48. Design a human review and annotation operation that supports evals at scale (thousands of items per week).
- Next: Q50. You're running the launch-readiness review for a new customer-facing AI feature. What evaluation and guardrail evidence do you require before go-live?
- Guardrails AI
- NeMo Guardrails
This is what real progress feels like.