AgentOps
AgentOps is the operational discipline of deploying, scaling, observing and governing autonomous AI agents in production.
Last updated: 09 Oct, 2026
Securing agent memory closed the memory module. The last module of the video takes one agent out of the notebook and runs it as a service that people call: it has to start on a server, stay up, show what it did on every request, refuse what it should refuse, and keep answering when many people ask at once.
This part of the video starts at 5:51:04. It opens with the application the whole module uses, a question-answering agent over research papers that is already deployed on Amazon EKS, and then names what an agent needs once it leaves the laptop: tracing, evaluation, security and deployment.
Comparing a prototype with production
The slide in the video is titled "Your Agent Works on Localhost. Now What?" and puts two lists side by side. The left list is how most agents start. The right list is what the same agent meets when real users arrive.
- A notebook with hardcoded API keys becomes a service with secrets kept outside the code. A key typed into a cell ends up in a screen share or a repository.
- One test query becomes many users at the same moment. The slide's example figure is 10,000 concurrent users; the load test later in the video uses 10, 20 and 50 users and already finds the limit of the cluster.
- "It works on my machine" becomes non-deterministic outputs at scale. The same question can produce a different answer, a different tool call or a different number of steps on the next run.
- No failure modes becomes cascading failures. An agent calls a guardrail, a search engine and a model, one after the other. When one of them slows down, every request behind it waits.
- No rollback plan becomes a rollback you have rehearsed. On Kubernetes,
kubectl rollout undoreturns a Deployment to its previous image.
The yellow note on the slide gives the reason the old playbook is not enough: "Traditional MLOps was designed for models that predict. Agents act. That changes everything." A model that predicts takes an input and returns a number or a label. An agent decides which tool to call, loops, rewrites its own question and writes free text, so there is more to watch and more that can go wrong on each request.
The six pillars of AgentOps
The next slide defines the term and splits the work into six pillars. Two of them are taught in the lessons that follow this one. Three belong to deployment and scaling. One is only named.
| Pillar | The question it answers | Where it is taught |
|---|---|---|
| Observability and tracing | What did the agent do on this request, and how long did each step take? | Tracing agents with Langfuse |
| Governance and guardrails | What must the agent refuse, and what must it never send back? | Amazon Bedrock Guardrails |
| Deployment and orchestration | Where does the agent run, and what restarts it when it dies? | Deploying on Amazon EKS |
| Scaling and reliability | What happens at 50 users instead of one? | Load testing with Locust, Horizontal pod autoscaling (HPA) |
| Agentic CI/CD | How does a code change reach the cluster without a person copying files? | Deploying on Amazon EKS |
| A2A multi-agent coordination | How do several agents call each other? | The project has an agent-to-agent service; the video does not walk through it. MCP server for an agentic RAG API shows how the same API is offered to other agents as tools. |
The project used in this module
Every lesson of this part reads the same project: an agentic RAG service over arXiv papers in computer science, AI and machine learning. RAG, retrieval-augmented generation, means the service first searches its own documents and then asks a model to answer from what it found. The repository builds the service in seven phases, and each phase has a notebook and a workflow document.
On a laptop the project is four containers started by Docker Compose. Everything that holds state for longer than a container lives in a managed cloud service: the paper table in Neon (a serverless Postgres), cached answers in Upstash (a serverless Redis), traces in Langfuse Cloud. The model is either Amazon Bedrock or OpenAI, chosen by one setting in the .env file.
| Phase | What it adds | Lesson |
|---|---|---|
| 1 | The containers, the health checks and the cloud accounts | Agentic RAG API with FastAPI and LangGraph |
| 2 | The Airflow DAG that fetches arXiv papers, parses the PDFs and stores them | Agentic RAG API with FastAPI and LangGraph |
| 3 | Keyword search with BM25 in OpenSearch | Hybrid search with BM25 and vector search |
| 4 | Chunking, embeddings and hybrid search | Hybrid search with BM25 and vector search |
| 5 | The RAG routes /ask and /stream | Redis caching for RAG |
| 6 | Tracing with Langfuse and the Redis answer cache | Tracing agents with Langfuse, Redis caching for RAG |
| 7 | The LangGraph agent with its guardrail nodes | Agentic RAG API with FastAPI and LangGraph, Amazon Bedrock Guardrails |
MLOps vs AgentOps
| MLOps | AgentOps | |
|---|---|---|
| What runs | A trained model that predicts | An agent that calls tools, loops and writes text |
| One request | One model call | Several calls: guardrail, search, grading, generation |
| What you trace | Latency and the prediction | Every step, its input, its output and its tokens |
| What you test | Accuracy on a held-out set | Answers against goldens, and refusals against attacks |
| What can go wrong | A wrong prediction | A wrong answer, a harmful answer, a leaked secret, an endless loop |
| What a rollback restores | The previous model file | The previous image; the index, the cache and the guardrail settings stay as they are |
Where you use AgentOps
- Before the first external user. Tracing and guardrails go in while the agent is still small, because adding them after an incident means you have no record of the incident.
- When latency becomes a complaint. A trace shows which step holds the time, so the fix goes to the right place.
- When cost becomes a line in a budget. Token counts per step and a cache for repeated questions are the two levers this part teaches.
Related
- Previous: Securing agent memory
- Next: Agentic RAG API with FastAPI and LangGraph
- See also: LLM observability with Pydantic Logfire
- Reference: the project repository, branch agentops
- Take an agent you have built and write one line per pillar: where it runs, what restarts it, what records each step, what it refuses. Every empty line is a job still to do.
- List what a rollback of your agent's code would leave unchanged: the index, the cache, the prompts. Write next to each one how you would restore it.
- Count the calls your agent makes for one question (guardrail, search, model). That count is the number of places a slow service can hold a request.
Little by little, you're building something great.