Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q33IntermediateConcept

How do you run a red-teaming program for an LLM application?

30-second answerSay your answer out loud first, then reveal.
A red-teaming program flows from a threat model to a per-category test plan, through manual and automated attacks to findings, mitigations and a CI regression suite, with production monitoring looping back into the test plan.

Attack categories to cover: direct jailbreaks, indirect injection via documents, emails and web pages, system prompt extraction, data exfiltration (other users' data, secrets), unsafe tool actions, harmful content in your domain, bias and discrimination, resource abuse (very long inputs, loops), multilingual and encoded attacks (base64, leetspeak, Hinglish).

Good practice

  • Severity rubric (e.g. critical: data leakage or unauthorised action; low: mildly off-brand output).
  • Diverse testers: languages, cultures, domain experts.
  • Measure the attack success rate over time; fixes should reduce it without causing over-refusal.
  • Tools: Promptfoo red-teaming, PyRIT, garak, custom generators.

Slow is fine. Stopping is the only problem.