1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
A jailbreak of your public chatbot goes viral on social media, with screenshots of it saying offensive things. What do you do?
30-second answerSay your answer out loud first, then reveal.
Timeline
- First hours: contain.
• Reproduce the jailbreak; find the vulnerable path (system prompt, model, missing output filter).
• Hotfix: input pattern detection for the attack family, a stricter output classifier threshold for that category, or a temporary block of the feature.
• Check the logs for how many users tried it and whether anything worse happened (PII, tool misuse). - Communication: coordinate with comms/PR and legal; a factual, brief public response if appropriate; internal update.
- Days: robust fix.
• Generalise beyond the exact prompt (attackers will create variants). Use automated generation of variants to test the fix.
• Consider model-level changes (a different model, a safety-tuned system prompt) plus output classifiers as a backstop. - Post-mortem: why didn't red-teaming catch it? Add the category to the red-team plan; add monitoring (spikes in flagged outputs, social listening).
Key principle: output classifiers are a crucial backstop for public bots, because model alignment alone will eventually be bypassed.
Related
This is what real progress feels like.