1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
How are LLMs made safe? Discuss alignment techniques, jailbreaks, and red-teaming.
30-second answerSay your answer out loud first, then reveal.
Training-level methods
- SFT + RLHF on helpful/harmless data, including refusal examples (and examples of not over-refusing).
- Constitutional AI / RLAIF (Anthropic): the model critiques and revises outputs against written principles, and AI feedback scales preference labelling.
- Specification-guided training: models trained to reason about written safety policies.
- Adversarial training: include jailbreak attempts and the correct responses.
Common jailbreak types
| Type | Example |
|---|---|
| Role-play / persona | "Pretend you're an AI without rules..." |
| Obfuscation | Encoding requests in base64, other languages, leetspeak |
| Many-shot | Long context filled with fake dialogue of compliance |
| Multi-turn escalation | Gradually steering across turns |
| Prompt injection (apps) | Instructions hidden in documents or web pages |
| Optimised adversarial suffixes | Automatically found token strings |
System-level defences
- Input/output classifiers for harmful categories.
- Separation of trusted instructions from untrusted content (for apps and agents).
- Monitoring and abuse detection across sessions; rate limiting.
- Human review for high-risk domains.
Balancing helpfulness: over-refusal is a failure too. Measure both harmful-compliance rate and false-refusal rate.
Red-teaming: domain experts plus automated attack generation, before and after release, with results feeding back into training and filters.
Related
Every expert started right here.