AI SecurityNeMo Guardrails 0.24 · RAGAS 0.4 · OpenAI SDK 3.3 · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your path

AI security checklist for production

An AI security checklist is a list of checks a team runs before an LLM application goes to production, covering how it can be attacked, how it can be wrong, what it remembers and how it behaves under load.

Last updated: 09 Oct, 2026 · Python 3.12

Horizontal pod autoscaling (HPA) ended with a deployment that looked scaled and was not. Most production problems of LLM applications have that shape: a control exists on paper and nobody checked that it works. A checklist turns each control into one question with a yes or no answer.

Four groups of checks

An agent in production has four kinds of risk: it can be attacked, it can be wrong, it forgets or remembers the wrong thing, and it falls over under load. The checks below are grouped the same way.

An LLM application in the centre with four groups of checks pointing at it: guardrails because it can be attacked, evaluation because it can be wrong, memory because it forgets or remembers the wrong thing, and operations because it falls over under load.

Guardrails: it can be attacked

Check before launchWhyTaught in
The risks of the application are mapped to the OWASP Top 10 for LLM ApplicationsA list nobody wrote down is a list nobody testsLLM security risks (OWASP Top 10)
An input check runs before the model and an output check after itThe model's own refusals are not a controlAI guardrails, Input and output rails
Direct and indirect prompt injection are both tested, including text hidden in retrieved documentsRAG and tools bring in text the user never typedPrompt injection and jailbreaks
Off-topic, jailbreak and sensitive-topic requests each have a rail, and the refusals seen are the scripted onesA model that refuses by itself today may not tomorrowTopic, jailbreak and sensitive-topic rails
Personal data is checked in both directions, with the misses of the patterns knownA regex has false positives and gapsInput and output rails
Block rate and false-positive rate are measured on a labelled setAn unmeasured guardrail is a guessMeasuring a guardrail
Timeouts, retries and a fallback model are in placeA provider outage or a retired model id should not be an outage of yoursLLM gateways
A managed guardrail sits on the model path where the platform offers oneA second layer with different blind spotsAmazon Bedrock Guardrails

Evaluation: it can be wrong

Check before launchWhyTaught in
A golden set built from real questions exists and is versioned"It looks right" is not a testGoldens
The judge model's scores were checked for repeatabilityA judge is a model tooLLM as a judge
Answers are scored for faithfulness to the retrieved contextIt catches invented claimsFaithfulness
Retrieval is scored by itself: context precision and context recallA good generator cannot fix a bad retrieverContext precision, Context recall
Answers are compared with reference answers where references existFaithful is not the same as correctAnswer correctness
Thresholds come from a baseline run and known-bad samplesA threshold picked from the air passes everything or nothingReading evaluation results
A deterministic eval gate runs on every commit, and the judged metrics on a scheduleRegressions arrive with ordinary code changesEvals in CI

Memory: it forgets, or remembers the wrong thing

Check before launchWhyTaught in
The conversation history sent to the model is bounded by tokensCost and latency grow with every turn otherwiseToken buffer memory
Summaries are checked for dropped or changed factsA summary of a summary driftsSummary memory
Long-term memories are stored and read per userOne user's facts must never answer another user's questionSecuring agent memory
A check runs before anything is written to long-term memoryAn injected instruction saved as a fact comes back in later sessionsSecuring agent memory
Stored memories hold no personal data that was not meant to be keptMemory is a database, with a database's dutiesSecuring agent memory
Old memories decay and are prunedA store that only grows returns stale factsForgetting and decay in agent memory

Operations: it falls over under load

Check before launchWhyTaught in
Every public endpoint asks for credentials, the MCP endpoint includedAn open endpoint spends your model budget for anyoneMCP server for an agentic RAG API
Traffic is encrypted, and only what must be public has a public addressDashboards and schedulers are admin toolsDeploying on Amazon EKS
No long-lived cloud keys in manifests or CIA leaked static key works until someone noticesDeploying on Amazon EKS
Every request is traced, with tokens and cost, and sensitive values scrubbedYou cannot investigate what you did not recordTracing agents with Langfuse, LLM observability with Pydantic Logfire
Cache keys include everything that changes the answerA cache hit must never serve one user's answer to anotherRedis caching for RAG
A load test was run with a target for p95 and for the failure shareThe first real traffic should not be the testLoad testing with Locust
Autoscaling was watched end to end: new pods reach RunningA wanted pod count is not capacityHorizontal pod autoscaling (HPA)
Memory and CPU requests are set from measured usageAn oversized request blocks the schedulerHorizontal pod autoscaling (HPA)

Scoring a configuration against the checklist

A checklist that lives in a document is read once. One that lives in code runs on every release. The example keeps a few of the checks above as small functions over a configuration dictionary and prints what fails.

A check is a name and a rule

Each check is a group, a sentence, and a function that takes the configuration and returns true or false. Checks with a number in them carry their own threshold.

python
CHECKS = [
    ("Operations", "traffic is encrypted (TLS)", lambda c: c["tls"]),
    ("Operations", "failure share at peak load at most 1%",
     lambda c: c["failure_share_at_peak"] <= 0.01),
]

In the sample configuration the six operations values describe the deployment in the video: no credentials on the public API, no TLS, a 6.0 GiB memory request against 3.12 GiB in use, no node autoscaler, and 186 failed requests of 629 at 50 users. The other values are an invented starting point. The thresholds (20 goldens, 5% false positives, 1.5 times the measured memory, 1% failures) are this example's; set your own.

ExampleA release check: fifteen checks over one configuration
CHECKS = [
    ("Guardrails", "input check before the model", lambda c: c["input_check"]),
    ("Guardrails", "output check after the model", lambda c: c["output_check"]),
    ("Guardrails", "indirect injection tested on retrieved text", lambda c: c["indirect_injection_tested"]),
    ("Guardrails", "false-positive rate measured, at most 5%",
     lambda c: c["false_positive_rate"] is not None and c["false_positive_rate"] <= 0.05),
    ("Evaluation", "at least 20 goldens", lambda c: c["goldens"] >= 20),
    ("Evaluation", "faithfulness at or above its threshold", lambda c: c["faithfulness"] >= c["faithfulness_threshold"]),
    ("Evaluation", "an eval gate runs in CI", lambda c: c["eval_gate_in_ci"]),
    ("Memory", "history is bounded", lambda c: c["history_limit_tokens"] is not None),
    ("Memory", "long-term memory is separated per user", lambda c: c["memory_per_user"]),
    ("Memory", "a check runs before a memory is saved", lambda c: c["memory_write_check"]),
    ("Operations", "the public API asks for credentials", lambda c: c["public_api_requires_auth"]),
    ("Operations", "traffic is encrypted (TLS)", lambda c: c["tls"]),
    ("Operations", "memory request within 1.5x of measured use",
     lambda c: c["memory_request_gib"] <= 1.5 * c["memory_used_gib"]),
    ("Operations", "something adds nodes when pods do not fit", lambda c: c["node_autoscaler"]),
    ("Operations", "failure share at peak load at most 1%", lambda c: c["failure_share_at_peak"] <= 0.01),
]

config = {
    "input_check": True, "output_check": True, "indirect_injection_tested": False, "false_positive_rate": None,
    "goldens": 5, "faithfulness": 0.75, "faithfulness_threshold": 0.70, "eval_gate_in_ci": True,
    "history_limit_tokens": 2000, "memory_per_user": True, "memory_write_check": False,
    # the last six values describe the deployment in the video
    "public_api_requires_auth": False, "tls": False, "memory_request_gib": 6.0, "memory_used_gib": 3.12,
    "node_autoscaler": False, "failure_share_at_peak": 186 / 629,
}

failed, groups = [], {}
for group, text, rule in CHECKS:
    passed = bool(rule(config))
    done, total = groups.get(group, (0, 0))
    groups[group] = (done + passed, total + 1)
    if not passed:
        failed.append(f"{group}: {text}")

for group, (done, total) in groups.items():
    print(f"{group:11} {done} of {total}")
print(f"score: {sum(d for d, _ in groups.values())} of {len(CHECKS)}")
print("ship:", "yes" if not failed else "no, fix these first")
for line in failed:
    print("  -", line)

What the score says

  • 6 of 15 checks pass, and the answer to "ship?" is no.
  • Operations scores 0 of 5. Each of the five comes from a reading of the video's screen or files: an open API, plain HTTP, a request almost twice the measured use, no way to add a node, and a failure share of about 30% at the peak.
  • The other groups each miss one or two. Indirect injection was never tested, the false-positive rate was never measured (None fails the check, it does not pass it), there are 5 goldens where the rule asks for 20, and nothing checks a memory before it is saved.
  • A missing measurement is a failure. The rule for the false-positive rate tests for None first. A check that passes when the number is absent rewards not measuring.

What this checklist leaves out

TopicWhat it isWhere to start
Knowledge-graph memoryMemories stored as entities and relations instead of text or vectorsGraph stores such as Neo4j, and graph-based memory libraries
Red-teaming toolsAutomated attack suites that probe a model or an applicationgarak, PyRIT, promptfoo
Fine-tuned safety classifiersSmall models trained to label prompts and replies, run as a guardrailGuardrail frameworks names Llama Guard and Prompt Guard
Azure and Google Cloud equivalentsThe managed guardrail and Kubernetes services of the other cloudsAzure AI Content Safety and AKS; Model Armor and GKE on Google Cloud
MCP authorizationThe OAuth 2.1 flow the MCP specification defines for remote serversThe authorization pages of the MCP specification
Node autoscaling and custom metricsKarpenter or Cluster Autoscaler, and scaling on queue length with KEDAThe autoscaling section of the Kubernetes docs

Before launch vs after launch

Before launchAfter launch
GuardrailsRun the labelled attack setWatch the block rate and review a sample of blocked and allowed messages
EvaluationPass the golden setRe-run it on a schedule and add real failures to it
MemoryTest isolation between two usersAudit what was written, prune and decay
OperationsLoad test to the expected peakAlert on p95, failure share, Pending pods and cost

Where you use an AI security checklist

  • As a release gate, in CI, next to the deterministic evaluation tests.
  • In a design review, before the first line of code, to decide which checks apply.
  • After an incident, to add the check that would have caught it.
Watch out. A checklist records that a control exists, not that it works. "Autoscaler: yes" was true of the cluster in the video. Wherever a check can be a measurement (a block rate, a faithfulness score, a failure share at peak load), store the number and its threshold instead of a yes.

Start again from the overview: AI security

Try it yourself
  • Set "public_api_requires_auth" and "tls" to True in the configuration: Operations rises to 2 of 5 and the score to 8 of 15.
  • Set "false_positive_rate" to 0.08: the check still fails, now because 8% is above the 5% threshold, and the score does not move.
  • Add your own check to CHECKS, for example ("Operations", "p95 under 30 s", lambda c: c["p95_ms"] <= 30000), and add "p95_ms": 30000 to the configuration: the total becomes 16 checks.

Little by little, you're building something great.