Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q34IntermediateScenario

A jailbreak of your public chatbot goes viral on social media, with screenshots of it saying offensive things. What do you do?

30-second answerSay your answer out loud first, then reveal.

Timeline

  1. First hours: contain.
    • Reproduce the jailbreak; find the vulnerable path (system prompt, model, missing output filter).
    • Hotfix: input pattern detection for the attack family, a stricter output classifier threshold for that category, or a temporary block of the feature.
    • Check the logs for how many users tried it and whether anything worse happened (PII, tool misuse).
  2. Communication: coordinate with comms/PR and legal; a factual, brief public response if appropriate; internal update.
  3. Days: robust fix.
    • Generalise beyond the exact prompt (attackers will create variants). Use automated generation of variants to test the fix.
    • Consider model-level changes (a different model, a safety-tuned system prompt) plus output classifiers as a backstop.
  4. Post-mortem: why didn't red-teaming catch it? Add the category to the red-team plan; add monitoring (spikes in flagged outputs, social listening).

Key principle: output classifiers are a crucial backstop for public bots, because model alignment alone will eventually be bypassed.

This is what real progress feels like.