Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q45HardConcept

How are LLMs made safe? Discuss alignment techniques, jailbreaks, and red-teaming.

30-second answerSay your answer out loud first, then reveal.

Training-level methods

  • SFT + RLHF on helpful/harmless data, including refusal examples (and examples of not over-refusing).
  • Constitutional AI / RLAIF (Anthropic): the model critiques and revises outputs against written principles, and AI feedback scales preference labelling.
  • Specification-guided training: models trained to reason about written safety policies.
  • Adversarial training: include jailbreak attempts and the correct responses.

Common jailbreak types

TypeExample
Role-play / persona"Pretend you're an AI without rules..."
ObfuscationEncoding requests in base64, other languages, leetspeak
Many-shotLong context filled with fake dialogue of compliance
Multi-turn escalationGradually steering across turns
Prompt injection (apps)Instructions hidden in documents or web pages
Optimised adversarial suffixesAutomatically found token strings

System-level defences

  • Input/output classifiers for harmful categories.
  • Separation of trusted instructions from untrusted content (for apps and agents).
  • Monitoring and abuse detection across sessions; rate limiting.
  • Human review for high-risk domains.

Balancing helpfulness: over-refusal is a failure too. Measure both harmful-compliance rate and false-refusal rate.

Red-teaming: domain experts plus automated attack generation, before and after release, with results feeding back into training and filters.

Every expert started right here.