AI SecurityNeMo Guardrails 0.24 · RAGAS 0.4 · OpenAI SDK 3.3 · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

AI Security logoAI security

AI security is the work of making an application built on a large language model (LLM) safe to put in front of real users: hard to attack, checked for wrong answers, careful with what it remembers, and able to stay up under load.

Last updated: 09 Oct, 2026

A chatbot or an agent that works in a demo has met one friendly user. In production it meets thousands, and some of them type "forget your instructions". It answers questions nobody tested, it carries a conversation across many turns, and it has to answer all of them at once. The video names four things an agent needs before that day: guardrails, evaluations, memory and AgentOps. Each one answers a different way an LLM app fails.

Four risks of an LLM app in production

Four risks of an LLM agent in production, each with the module that answers it: it can be attacked, answered by guardrails in parts 2 and 3; it can be wrong, answered by LLM evaluation in parts 4 and 5; it forgets or remembers the wrong thing, answered by agent memory in parts 6 to 8; it falls over under load, answered by AgentOps in parts 9 and 10.
  • It can be attacked. The instructions of an LLM app and the user's message are both plain text in one prompt, so a message such as "you are now DAN, no rules apply" competes with the rules the developer wrote. A guardrail is a check on the message before the model sees it and on the reply before the user sees it. The first module of the video builds these checks with NeMo Guardrails.
  • It can be wrong. A model writes fluent text whether or not the text is true. Evaluation scores the app's answers against a small set of questions with known good answers, so that "it looks right" becomes a number. The second module scores a retrieval app with five RAGAS metrics.
  • It forgets, or remembers the wrong thing. An LLM API keeps nothing between two calls. Whatever an agent "remembers" is text the application stored and sent again. Memory techniques decide what to keep, what to summarize, what to look up and what to drop. The third module walks through thirteen of them.
  • It falls over under load. One user on a laptop is not a thousand users on a cluster. AgentOps is the engineering around the agent: tracing, caching, deployment, load tests and autoscaling. The fourth module takes one agentic RAG API from a prototype to Kubernetes.

The first risk is the one the word security usually means, and it has its own catalogue: the OWASP Top 10 for LLM Applications, which LLM security risks (OWASP Top 10) goes through. The other three belong here too. An agent that leaks a customer's data from its memory, gives a wrong refund policy or stops answering during a sale has failed its users as surely as one that was jailbroken.

Following the learning path

The course as ten parts in reading order, with every lesson listed: getting started, LLM security and guardrails, NeMo Guardrails, LLM evaluation, RAG evaluation metrics, short-term agent memory, long-term agent memory, managing agent memory, AgentOps: the production stack, and deploying and scaling. Parts two and three are guardrails, four and five evaluation, six to eight agent memory, nine and ten AgentOps.

The ten parts follow the order of the video. Parts 2 and 3 are its guardrails module, parts 4 and 5 its evaluation module, parts 6 to 8 its memory module and parts 9 and 10 its AgentOps module. Six lessons cover topics the video does not teach and a complete picture needs: LLM gateways, Measuring a guardrail, Evals in CI, Securing agent memory, the OWASP list and the closing AI security checklist for production.

Finding each part in the video

Every lesson that the video teaches carries one or more clips, placed above the section they cover. The table lists the lessons of each part and where the same material starts in the video. A timestamp opens the video on YouTube at that second.

PartLessonsWhere it is in the video
1. Getting startedAI security · Installing Python for AI security0:00:00 the four modules in three minutes
2. LLM security and guardrailsLLM security risks (OWASP Top 10) · AI guardrails · Prompt injection and jailbreaks · Guardrail frameworks · LLM gateways0:03:08 LLM security and guardrails
0:16:38 guardrail frameworks
0:20:50 the demo: prompt injection, off-topic questions and jailbreaks
3. NeMo GuardrailsNeMo Guardrails · Colang · Intent detection in NeMo Guardrails · Topic, jailbreak and sensitive-topic rails · Input and output rails · Measuring a guardrail · LLM observability with Pydantic Logfire0:36:20 NeMo Guardrails and Colang
0:51:04 observability with Pydantic Logfire
4. LLM evaluationLLM evaluation · Benchmarks vs custom evaluation · Goldens · LLM as a judge · RAGAS metrics1:13:30 evaluating a production RAG app
1:23:18 custom evaluations and benchmarks
1:30:44 goldens
1:49:54 an LLM as a judge
1:52:46 the RAGAS metrics
5. RAG evaluation metricsFaithfulness · Answer relevancy · Context precision · Context recall · Answer correctness · Reading evaluation results · Evals in CI2:04:58 faithfulness
2:12:35 answer relevancy
2:18:02 context precision
2:24:48 context recall
2:30:46 answer correctness
2:40:31 the test results
6. Short-term agent memoryAgent memory · Conversation buffer memory · Sliding window memory · Summary memory · Summary buffer memory · Token buffer memory2:47:50 agent memory
3:01:00 conversation buffer memory
3:12:43 sliding window memory
3:37:24 summary memory
3:56:30 summary buffer memory
4:20:05 token buffer memory
7. Long-term agent memoryVector store memory · Entity memory · Episodic memory · Semantic memory · Procedural memory4:24:41 vector store memory
4:41:29 entity memory
4:56:29 episodic memory
5:15:54 semantic memory
5:20:14 procedural memory
8. Managing agent memorySelf-reflection memory · Memory routing · Forgetting and decay in agent memory · Securing agent memory5:25:56 self-reflection memory
5:33:13 memory routing
5:40:23 forgetting and decay
9. AgentOps: the production stackAgentOps · Agentic RAG API with FastAPI and LangGraph · Tracing agents with Langfuse · Amazon Bedrock Guardrails · Hybrid search with BM25 and vector search · Redis caching for RAG5:50:47 AgentOps
5:55:27 Airflow, Neon and OpenSearch
6:08:10 the FastAPI endpoints
6:10:58 Langfuse tracing
6:19:11 Amazon Bedrock Guardrails
6:40:41 hybrid search
6:45:01 Redis caching
10. Deploying and scalingMCP server for an agentic RAG API · Deploying on Amazon EKS · Load testing with Locust · Horizontal pod autoscaling (HPA) · AI security checklist for production6:53:12 the MCP server
7:08:50 Amazon EKS
7:22:25 load testing with Locust
7:31:42 autoscaling

Who this course is for

  • Developers who have built a chatbot, a RAG app or an agent and now have to put it in front of users. Every topic here is something a demo skips and a launch review asks about.
  • People preparing for AI engineering interviews. Prompt injection, guardrails, RAG evaluation metrics, memory techniques and scaling are the questions the video says interviewers ask about agents.
  • Backend and platform engineers who are handed an LLM service to run and want to know what is different about it: non-deterministic output, token cost, and attacks written in plain English.

Preparing what you need

  • Python: functions, dictionaries, classes, importing a library and reading a traceback.
  • One call to an LLM API. If you have sent a list of messages to a chat model and printed the reply, you have enough. Installing Python for AI security makes that call with you.
  • Two free API keys: a Groq key for the chat models and a Google Gemini key for embeddings. The same lesson sets both up and shows how to keep them out of your code.
  • Python 3.12 or 3.13 on Windows, macOS or Linux, or Google Colab in a browser. No GPU and no model downloads.
  • Optional company: RAG and agents are explained again where they are used. For more depth, the LangChain tutorial, the LangGraph tutorial, the NeMo Guardrails tutorial and the RAGAS tutorial each go further into one tool.

Reading a lesson

A lesson opens with a one-sentence definition, then a clip where the video teaches the idea. The text under the clip follows it, with the video's own example. Then the example is run again for real: the code calls a hosted model, and the reply printed under it is the reply that came back, including the times a model ignored a rule or a guardrail let an attack through. Those runs are the most useful ones to read. Numbers such as a score, a token count or a block rate come from plain Python that gives the same result every time. Each lesson closes with a few changes to try.

Following the video and its repositories

The lessons follow The Complete AI Security Course In 8 Hours-AI Guardrails, LLM Evals & Memory And AgentOps (June 2026, 7 hours 48 minutes), four modules recorded as live classes. Theory is drawn on a digital whiteboard and the practical parts run in demo apps, notebooks and a terminal.

The video's description links the code behind the guardrails, evaluation and AgentOps modules, and these repositories serve as the notes:

  • guardrails-webinar: the NeMo Guardrails demo app of the first module, with its Colang rules, custom actions and a notebook of saved runs.
  • LIVE-WEBINAR-25-MAY-GATEWAYS: the LLM gateway demos from the class before the guardrails one, used in LLM gateways.
  • ragas: the retrieval app, the goldens and the metric code of the evaluation module.
  • Agentic-RAG-project: the agentic RAG API of the AgentOps module, with its Kubernetes manifests and load tests.

The memory module has no linked repository. It works through thirteen notebooks on screen, one per technique, and the memory lessons show the code of each.

Hosted models and libraries have moved since June 2026. These are the differences the lessons point out where they matter:

The video (June 2026)Today
Groq models llama-3.3-70b-versatile and llama-3.1-8b-instantBoth are retired on Groq and return a 404. The lessons run openai/gpt-oss-120b and openai/gpt-oss-20b
OpenAI models and an OpenAI key in the memory notebooksThe same OpenAI SDK pointed at Groq with base_url, so only two lines change
Local embedding models: FastEmbed in NeMo Guardrails, HuggingFace models in the evaluation app, ChromaDB's defaultHosted Gemini embeddings, gemini-embedding-2, one text per call. Nothing is downloaded
NeMo Guardrails with the model passed in from LangChainNeMo Guardrails 0.24.1 with the model named in config.yml
RAGAS installed without pinsragas==0.4.3 with "langchain-community<0.4"; the metrics come from ragas.metrics.collections
Logfire 4.35 behind the demo app's tracesLogfire 5.1, where a project has its own write tokens
Back toAll courses

Every expert started right here.