AI SecurityNeMo Guardrails 0.24 · RAGAS 0.4 · OpenAI SDK 3.3 · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

AgentOps

AgentOps is the operational discipline of deploying, scaling, observing and governing autonomous AI agents in production.

Last updated: 09 Oct, 2026

Securing agent memory closed the memory module. The last module of the video takes one agent out of the notebook and runs it as a service that people call: it has to start on a server, stay up, show what it did on every request, refuse what it should refuse, and keep answering when many people ask at once.

AgentOps: from prototype to production · from the Complete AI Security Course in 8 Hours video · 5:51:04 to 5:54:26

This part of the video starts at 5:51:04. It opens with the application the whole module uses, a question-answering agent over research papers that is already deployed on Amazon EKS, and then names what an agent needs once it leaves the laptop: tracing, evaluation, security and deployment.

Comparing a prototype with production

The slide in the video is titled "Your Agent Works on Localhost. Now What?" and puts two lists side by side. The left list is how most agents start. The right list is what the same agent meets when real users arrive.

Two cards side by side compare a prototype (a notebook with keys typed into the code, one test query, it works on my machine, no concurrency and no failure modes) with production (many users at once, non-deterministic outputs at scale, cascading failures and the need for a rollback plan), above a note that MLOps was designed for models that predict, while agents act.
  • A notebook with hardcoded API keys becomes a service with secrets kept outside the code. A key typed into a cell ends up in a screen share or a repository.
  • One test query becomes many users at the same moment. The slide's example figure is 10,000 concurrent users; the load test later in the video uses 10, 20 and 50 users and already finds the limit of the cluster.
  • "It works on my machine" becomes non-deterministic outputs at scale. The same question can produce a different answer, a different tool call or a different number of steps on the next run.
  • No failure modes becomes cascading failures. An agent calls a guardrail, a search engine and a model, one after the other. When one of them slows down, every request behind it waits.
  • No rollback plan becomes a rollback you have rehearsed. On Kubernetes, kubectl rollout undo returns a Deployment to its previous image.

The yellow note on the slide gives the reason the old playbook is not enough: "Traditional MLOps was designed for models that predict. Agents act. That changes everything." A model that predicts takes an input and returns a number or a label. An agent decides which tool to call, loops, rewrites its own question and writes free text, so there is more to watch and more that can go wrong on each request.

The six pillars of AgentOps

The next slide defines the term and splits the work into six pillars. Two of them are taught in the lessons that follow this one. Three belong to deployment and scaling. One is only named.

A two by three grid of the six pillars of AgentOps (deployment and orchestration, scaling and reliability, agentic CI/CD, A2A multi-agent coordination, observability and tracing, governance and guardrails), with observability and guardrails marked as taught in this part, three pillars as taught in the deployment part, and A2A coordination as named only.
PillarThe question it answersWhere it is taught
Observability and tracingWhat did the agent do on this request, and how long did each step take?Tracing agents with Langfuse
Governance and guardrailsWhat must the agent refuse, and what must it never send back?Amazon Bedrock Guardrails
Deployment and orchestrationWhere does the agent run, and what restarts it when it dies?Deploying on Amazon EKS
Scaling and reliabilityWhat happens at 50 users instead of one?Load testing with Locust, Horizontal pod autoscaling (HPA)
Agentic CI/CDHow does a code change reach the cluster without a person copying files?Deploying on Amazon EKS
A2A multi-agent coordinationHow do several agents call each other?The project has an agent-to-agent service; the video does not walk through it. MCP server for an agentic RAG API shows how the same API is offered to other agents as tools.

The project used in this module

Every lesson of this part reads the same project: an agentic RAG service over arXiv papers in computer science, AI and machine learning. RAG, retrieval-augmented generation, means the service first searches its own documents and then asks a model to answer from what it found. The repository builds the service in seven phases, and each phase has a notebook and a workflow document.

Four containers on the Docker network rag-network, rag-api (FastAPI, port 8000), rag-opensearch (OpenSearch 2.19.5, ports 9200 and 9600), rag-dashboards (port 5601) and rag-airflow (Airflow 2.10.3, port 8080), with the API searching OpenSearch and Airflow indexing chunks into it, and four cloud services outside the network reached over HTTPS: Neon Postgres, Upstash Redis, Langfuse Cloud, and the LLM and embedding providers.

On a laptop the project is four containers started by Docker Compose. Everything that holds state for longer than a container lives in a managed cloud service: the paper table in Neon (a serverless Postgres), cached answers in Upstash (a serverless Redis), traces in Langfuse Cloud. The model is either Amazon Bedrock or OpenAI, chosen by one setting in the .env file.

PhaseWhat it addsLesson
1The containers, the health checks and the cloud accountsAgentic RAG API with FastAPI and LangGraph
2The Airflow DAG that fetches arXiv papers, parses the PDFs and stores themAgentic RAG API with FastAPI and LangGraph
3Keyword search with BM25 in OpenSearchHybrid search with BM25 and vector search
4Chunking, embeddings and hybrid searchHybrid search with BM25 and vector search
5The RAG routes /ask and /streamRedis caching for RAG
6Tracing with Langfuse and the Redis answer cacheTracing agents with Langfuse, Redis caching for RAG
7The LangGraph agent with its guardrail nodesAgentic RAG API with FastAPI and LangGraph, Amazon Bedrock Guardrails

MLOps vs AgentOps

MLOpsAgentOps
What runsA trained model that predictsAn agent that calls tools, loops and writes text
One requestOne model callSeveral calls: guardrail, search, grading, generation
What you traceLatency and the predictionEvery step, its input, its output and its tokens
What you testAccuracy on a held-out setAnswers against goldens, and refusals against attacks
What can go wrongA wrong predictionA wrong answer, a harmful answer, a leaked secret, an endless loop
What a rollback restoresThe previous model fileThe previous image; the index, the cache and the guardrail settings stay as they are

Where you use AgentOps

  • Before the first external user. Tracing and guardrails go in while the agent is still small, because adding them after an incident means you have no record of the incident.
  • When latency becomes a complaint. A trace shows which step holds the time, so the fix goes to the right place.
  • When cost becomes a line in a budget. Token counts per step and a cache for repeated questions are the two levers this part teaches.
Watch out. A Kubernetes rollback restores the container image and stops there. The search index, the cached answers, the prompts stored outside the image and the guardrail configuration keep their current state, so a bad change to any of them needs its own way back.
Try it yourself
  • Take an agent you have built and write one line per pillar: where it runs, what restarts it, what records each step, what it refuses. Every empty line is a job still to do.
  • List what a rollback of your agent's code would leave unchanged: the index, the cache, the prompts. Write next to each one how you would restore it.
  • Count the calls your agent makes for one question (guardrail, search, model). That count is the number of places a slow service can hold a request.

Little by little, you're building something great.