Dashboard

Red-team report: attacking the support assistant before a customer does

A red-team run that probes your own assistant, scores what breaks it, and writes it up as a report a team can act on.

The problem

Before an assistant ships, someone will ask whether it is safe, and "it seemed fine when I tried it" is not an answer. You need evidence: what makes it leak, refuse wrongly, or follow an injected instruction, and how often.

You want a red-team run that probes the assistant with known attacks, presses it with adaptive multi-turn ones, scores what got through, and writes it up as a report a team can act on. Only test a system you are allowed to test.

Architecture

An attack pipeline that ends in a document. The target is scanned with a library of known probes, pressed with scripted multi-turn attacks, the responses are scored for whether the attack succeeded, and the findings become a report. The scoring is what turns a pile of transcripts into something a team can prioritise.

A target in, a scored report of what breaks it out
Probesgarak scansAttacksPyRIT, multi-turnThe targetyour assistantFindingseach attempt scoredThe reportbefore and after a fix
Hover or tap a piece to see what it does.

A finding is only useful if it is reproducible. Record the exact prompt, the response, and the score for each one, so a fix can be tested against the same case rather than a vibe of what went wrong.

What it draws on

Everything here comes from this stage of the roadmap; the project is where those courses meet.

  • garak: scanning for known attacks and reading attack success rate
  • PyRIT: scripted attacks, scorers and converters
  • Guardrails AI or NeMo Guardrails: the fix on inputs and outputs
  • Agent Governance Toolkit: the fix when the risk is a tool call

What done looks like

RequirementDone when
Scangarak runs against the assistant itself, not a bare model, and its report is kept
Targeted attacksAt least two PyRIT attacks specific to this assistant's tools and data
ScoringEach attack's success is judged by a scorer, and the scorer was checked on a few examples by hand
FixOne finding is fixed and the same attacks are run again
ReportStates attack success rates before and after, what is still open, and the risk accepted

Where to start

Point a scanner at your own assistant first and read the attack success rate; that is your baseline. Then script the two or three attacks that matter most for your system, score them, and write the report around what got through. Keep every successful attack as a case, so the fix has something to prove itself against.

Only run this against a system you own or are authorised to test. A free Groq key covers scanning a hosted model.
Back toAll frameworks

Every expert started right here.