Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q49HardConcept

How do you evaluate the guardrails themselves? How do you measure bypass resistance?

30-second answerSay your answer out loud first, then reveal.

Evaluation dimensions

DimensionMetric
Detection qualityRecall (harmful caught), precision (benign allowed), per category
Over-refusalFalse positive rate on benign look-alikes
RobustnessBypass rate under adversarial transformations
CoverageLanguages, modalities (text in images), channels (documents, tool outputs)
PerformanceAdded latency p95, cost per request, availability
DriftPerformance over time as user behaviour and attacks change

Adversarial transformation suite

  • Paraphrase / synonym substitution.
  • Encoding (base64, ROT13), character tricks (zero-width spaces, homoglyphs, spacing).
  • Translation into other languages or code-mixed text.
  • Payload splitting across turns or documents.
  • Role-play / hypothetical framing.
  • Injection hidden in markup (HTML comments, white text, alt text).

Interpretation: a guardrail with 98% recall on the clean dataset but a 40% bypass rate under encoding needs fixes (normalise or decode inputs before classification, add more layers). Report both numbers.

This is what real progress feels like.