Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q48HardSystem design

Design the safety and abuse-prevention architecture for a public, free-to-use generative AI app.

30-second answerSay your answer out loud first, then reveal.
A user request passes an account layer, rate limits and input classifiers (block to refusal), then an aligned model and output classifiers (block to refusal) before the response; refusals and responses feed behavioural monitoring and user reports, which lead to enforcement and classifier retraining.

Key design considerations

  1. Severity tiers: some categories are zero tolerance with legal reporting obligations (e.g. child sexual abuse material), others are context-dependent (medical or security education). Policies define each tier.
  2. Account-level signals catch what single requests can't: many borderline requests in sequence, automated scraping, coordinated misuse.
  3. Latency vs safety: fast classifiers inline; heavier analysis asynchronous (can lead to later account action).
  4. Image/media generation: prompt filtering + output image classifiers + provenance (watermarking, content credentials metadata).
  5. Over-refusal: track false refusal rate on benign test sets; overly strict filters push users away.
  6. Red-teaming: internal and external, before launches and continuously; adversarial test suites in CI.
  7. Privacy: safety logging balanced with data minimisation; access controls for reviewers.
  8. Incident response: playbooks for viral jailbreaks, quick classifier or prompt hotfixes, communication plans.

Slow is fine. Stopping is the only problem.