Red-team report: attacking the support assistant before a customer does
A red-team run that probes your own assistant, scores what breaks it, and writes it up as a report a team can act on.
The problem
Before an assistant ships, someone will ask whether it is safe, and "it seemed fine when I tried it" is not an answer. You need evidence: what makes it leak, refuse wrongly, or follow an injected instruction, and how often.
You want a red-team run that probes the assistant with known attacks, presses it with adaptive multi-turn ones, scores what got through, and writes it up as a report a team can act on. Only test a system you are allowed to test.
Architecture
An attack pipeline that ends in a document. The target is scanned with a library of known probes, pressed with scripted multi-turn attacks, the responses are scored for whether the attack succeeded, and the findings become a report. The scoring is what turns a pile of transcripts into something a team can prioritise.
A finding is only useful if it is reproducible. Record the exact prompt, the response, and the score for each one, so a fix can be tested against the same case rather than a vibe of what went wrong.
What it draws on
Everything here comes from this stage of the roadmap; the project is where those courses meet.
- garak: scanning for known attacks and reading attack success rate
- PyRIT: scripted attacks, scorers and converters
- Guardrails AI or NeMo Guardrails: the fix on inputs and outputs
- Agent Governance Toolkit: the fix when the risk is a tool call
What done looks like
| Requirement | Done when |
|---|---|
| Scan | garak runs against the assistant itself, not a bare model, and its report is kept |
| Targeted attacks | At least two PyRIT attacks specific to this assistant's tools and data |
| Scoring | Each attack's success is judged by a scorer, and the scorer was checked on a few examples by hand |
| Fix | One finding is fixed and the same attacks are run again |
| Report | States attack success rates before and after, what is still open, and the risk accepted |
Where to start
Point a scanner at your own assistant first and read the attack success rate; that is your baseline. Then script the two or three attacks that matter most for your system, score them, and write the report around what got through. Keep every successful attack as a case, so the fix has something to prove itself against.
Every expert started right here.