AI security checklist for production
An AI security checklist is a list of checks a team runs before an LLM application goes to production, covering how it can be attacked, how it can be wrong, what it remembers and how it behaves under load.
Last updated: 09 Oct, 2026 · Python 3.12
Horizontal pod autoscaling (HPA) ended with a deployment that looked scaled and was not. Most production problems of LLM applications have that shape: a control exists on paper and nobody checked that it works. A checklist turns each control into one question with a yes or no answer.
Four groups of checks
An agent in production has four kinds of risk: it can be attacked, it can be wrong, it forgets or remembers the wrong thing, and it falls over under load. The checks below are grouped the same way.
Guardrails: it can be attacked
| Check before launch | Why | Taught in |
|---|---|---|
| The risks of the application are mapped to the OWASP Top 10 for LLM Applications | A list nobody wrote down is a list nobody tests | LLM security risks (OWASP Top 10) |
| An input check runs before the model and an output check after it | The model's own refusals are not a control | AI guardrails, Input and output rails |
| Direct and indirect prompt injection are both tested, including text hidden in retrieved documents | RAG and tools bring in text the user never typed | Prompt injection and jailbreaks |
| Off-topic, jailbreak and sensitive-topic requests each have a rail, and the refusals seen are the scripted ones | A model that refuses by itself today may not tomorrow | Topic, jailbreak and sensitive-topic rails |
| Personal data is checked in both directions, with the misses of the patterns known | A regex has false positives and gaps | Input and output rails |
| Block rate and false-positive rate are measured on a labelled set | An unmeasured guardrail is a guess | Measuring a guardrail |
| Timeouts, retries and a fallback model are in place | A provider outage or a retired model id should not be an outage of yours | LLM gateways |
| A managed guardrail sits on the model path where the platform offers one | A second layer with different blind spots | Amazon Bedrock Guardrails |
Evaluation: it can be wrong
| Check before launch | Why | Taught in |
|---|---|---|
| A golden set built from real questions exists and is versioned | "It looks right" is not a test | Goldens |
| The judge model's scores were checked for repeatability | A judge is a model too | LLM as a judge |
| Answers are scored for faithfulness to the retrieved context | It catches invented claims | Faithfulness |
| Retrieval is scored by itself: context precision and context recall | A good generator cannot fix a bad retriever | Context precision, Context recall |
| Answers are compared with reference answers where references exist | Faithful is not the same as correct | Answer correctness |
| Thresholds come from a baseline run and known-bad samples | A threshold picked from the air passes everything or nothing | Reading evaluation results |
| A deterministic eval gate runs on every commit, and the judged metrics on a schedule | Regressions arrive with ordinary code changes | Evals in CI |
Memory: it forgets, or remembers the wrong thing
| Check before launch | Why | Taught in |
|---|---|---|
| The conversation history sent to the model is bounded by tokens | Cost and latency grow with every turn otherwise | Token buffer memory |
| Summaries are checked for dropped or changed facts | A summary of a summary drifts | Summary memory |
| Long-term memories are stored and read per user | One user's facts must never answer another user's question | Securing agent memory |
| A check runs before anything is written to long-term memory | An injected instruction saved as a fact comes back in later sessions | Securing agent memory |
| Stored memories hold no personal data that was not meant to be kept | Memory is a database, with a database's duties | Securing agent memory |
| Old memories decay and are pruned | A store that only grows returns stale facts | Forgetting and decay in agent memory |
Operations: it falls over under load
| Check before launch | Why | Taught in |
|---|---|---|
| Every public endpoint asks for credentials, the MCP endpoint included | An open endpoint spends your model budget for anyone | MCP server for an agentic RAG API |
| Traffic is encrypted, and only what must be public has a public address | Dashboards and schedulers are admin tools | Deploying on Amazon EKS |
| No long-lived cloud keys in manifests or CI | A leaked static key works until someone notices | Deploying on Amazon EKS |
| Every request is traced, with tokens and cost, and sensitive values scrubbed | You cannot investigate what you did not record | Tracing agents with Langfuse, LLM observability with Pydantic Logfire |
| Cache keys include everything that changes the answer | A cache hit must never serve one user's answer to another | Redis caching for RAG |
| A load test was run with a target for p95 and for the failure share | The first real traffic should not be the test | Load testing with Locust |
| Autoscaling was watched end to end: new pods reach Running | A wanted pod count is not capacity | Horizontal pod autoscaling (HPA) |
| Memory and CPU requests are set from measured usage | An oversized request blocks the scheduler | Horizontal pod autoscaling (HPA) |
Scoring a configuration against the checklist
A checklist that lives in a document is read once. One that lives in code runs on every release. The example keeps a few of the checks above as small functions over a configuration dictionary and prints what fails.
A check is a name and a rule
Each check is a group, a sentence, and a function that takes the configuration and returns true or false. Checks with a number in them carry their own threshold.
CHECKS = [
("Operations", "traffic is encrypted (TLS)", lambda c: c["tls"]),
("Operations", "failure share at peak load at most 1%",
lambda c: c["failure_share_at_peak"] <= 0.01),
]In the sample configuration the six operations values describe the deployment in the video: no credentials on the public API, no TLS, a 6.0 GiB memory request against 3.12 GiB in use, no node autoscaler, and 186 failed requests of 629 at 50 users. The other values are an invented starting point. The thresholds (20 goldens, 5% false positives, 1.5 times the measured memory, 1% failures) are this example's; set your own.
CHECKS = [
("Guardrails", "input check before the model", lambda c: c["input_check"]),
("Guardrails", "output check after the model", lambda c: c["output_check"]),
("Guardrails", "indirect injection tested on retrieved text", lambda c: c["indirect_injection_tested"]),
("Guardrails", "false-positive rate measured, at most 5%",
lambda c: c["false_positive_rate"] is not None and c["false_positive_rate"] <= 0.05),
("Evaluation", "at least 20 goldens", lambda c: c["goldens"] >= 20),
("Evaluation", "faithfulness at or above its threshold", lambda c: c["faithfulness"] >= c["faithfulness_threshold"]),
("Evaluation", "an eval gate runs in CI", lambda c: c["eval_gate_in_ci"]),
("Memory", "history is bounded", lambda c: c["history_limit_tokens"] is not None),
("Memory", "long-term memory is separated per user", lambda c: c["memory_per_user"]),
("Memory", "a check runs before a memory is saved", lambda c: c["memory_write_check"]),
("Operations", "the public API asks for credentials", lambda c: c["public_api_requires_auth"]),
("Operations", "traffic is encrypted (TLS)", lambda c: c["tls"]),
("Operations", "memory request within 1.5x of measured use",
lambda c: c["memory_request_gib"] <= 1.5 * c["memory_used_gib"]),
("Operations", "something adds nodes when pods do not fit", lambda c: c["node_autoscaler"]),
("Operations", "failure share at peak load at most 1%", lambda c: c["failure_share_at_peak"] <= 0.01),
]
config = {
"input_check": True, "output_check": True, "indirect_injection_tested": False, "false_positive_rate": None,
"goldens": 5, "faithfulness": 0.75, "faithfulness_threshold": 0.70, "eval_gate_in_ci": True,
"history_limit_tokens": 2000, "memory_per_user": True, "memory_write_check": False,
# the last six values describe the deployment in the video
"public_api_requires_auth": False, "tls": False, "memory_request_gib": 6.0, "memory_used_gib": 3.12,
"node_autoscaler": False, "failure_share_at_peak": 186 / 629,
}
failed, groups = [], {}
for group, text, rule in CHECKS:
passed = bool(rule(config))
done, total = groups.get(group, (0, 0))
groups[group] = (done + passed, total + 1)
if not passed:
failed.append(f"{group}: {text}")
for group, (done, total) in groups.items():
print(f"{group:11} {done} of {total}")
print(f"score: {sum(d for d, _ in groups.values())} of {len(CHECKS)}")
print("ship:", "yes" if not failed else "no, fix these first")
for line in failed:
print(" -", line)Guardrails 2 of 4 Evaluation 2 of 3 Memory 2 of 3 Operations 0 of 5 score: 6 of 15 ship: no, fix these first - Guardrails: indirect injection tested on retrieved text - Guardrails: false-positive rate measured, at most 5% - Evaluation: at least 20 goldens - Memory: a check runs before a memory is saved - Operations: the public API asks for credentials - Operations: traffic is encrypted (TLS) - Operations: memory request within 1.5x of measured use - Operations: something adds nodes when pods do not fit - Operations: failure share at peak load at most 1%
What the score says
- 6 of 15 checks pass, and the answer to "ship?" is no.
- Operations scores 0 of 5. Each of the five comes from a reading of the video's screen or files: an open API, plain HTTP, a request almost twice the measured use, no way to add a node, and a failure share of about 30% at the peak.
- The other groups each miss one or two. Indirect injection was never tested, the false-positive rate was never measured (
Nonefails the check, it does not pass it), there are 5 goldens where the rule asks for 20, and nothing checks a memory before it is saved. - A missing measurement is a failure. The rule for the false-positive rate tests for
Nonefirst. A check that passes when the number is absent rewards not measuring.
What this checklist leaves out
| Topic | What it is | Where to start |
|---|---|---|
| Knowledge-graph memory | Memories stored as entities and relations instead of text or vectors | Graph stores such as Neo4j, and graph-based memory libraries |
| Red-teaming tools | Automated attack suites that probe a model or an application | garak, PyRIT, promptfoo |
| Fine-tuned safety classifiers | Small models trained to label prompts and replies, run as a guardrail | Guardrail frameworks names Llama Guard and Prompt Guard |
| Azure and Google Cloud equivalents | The managed guardrail and Kubernetes services of the other clouds | Azure AI Content Safety and AKS; Model Armor and GKE on Google Cloud |
| MCP authorization | The OAuth 2.1 flow the MCP specification defines for remote servers | The authorization pages of the MCP specification |
| Node autoscaling and custom metrics | Karpenter or Cluster Autoscaler, and scaling on queue length with KEDA | The autoscaling section of the Kubernetes docs |
Before launch vs after launch
| Before launch | After launch | |
|---|---|---|
| Guardrails | Run the labelled attack set | Watch the block rate and review a sample of blocked and allowed messages |
| Evaluation | Pass the golden set | Re-run it on a schedule and add real failures to it |
| Memory | Test isolation between two users | Audit what was written, prune and decay |
| Operations | Load test to the expected peak | Alert on p95, failure share, Pending pods and cost |
Where you use an AI security checklist
- As a release gate, in CI, next to the deterministic evaluation tests.
- In a design review, before the first line of code, to decide which checks apply.
- After an incident, to add the check that would have caught it.
Related
- Previous: Horizontal pod autoscaling (HPA)
- Reference: OWASP Top 10 for LLM Applications
Start again from the overview: AI security
- Set
"public_api_requires_auth"and"tls"toTruein the configuration: Operations rises to 2 of 5 and the score to 8 of 15. - Set
"false_positive_rate"to0.08: the check still fails, now because 8% is above the 5% threshold, and the score does not move. - Add your own check to
CHECKS, for example("Operations", "p95 under 30 s", lambda c: c["p95_ms"] <= 30000), and add"p95_ms": 30000to the configuration: the total becomes 16 checks.
Little by little, you're building something great.