AI SecurityNeMo Guardrails 0.24 · RAGAS 0.4 · OpenAI SDK 3.3 · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

Deploying on Amazon EKS

Amazon EKS (Elastic Kubernetes Service) is a managed Kubernetes service: AWS runs the control plane, and Kubernetes keeps the containers you describe in manifest files running on a group of EC2 machines called nodes.

Last updated: 09 Oct, 2026 · Kubernetes Python client 37.0

On a laptop the whole project of AgentOps starts with one Docker Compose command. In production it has to survive a dead machine, take a new version without downtime and be reachable from outside. The video deploys the API from MCP server for an agentic RAG API on EKS, and the files it applies are short enough to read line by line.

The EKS cluster file · from the Complete AI Security Course in 8 Hours video · 7:11:26 to 7:12:38

This part of the video starts at 7:11:26. The minSize: 2 and maxSize: 4 lines on screen set the number of nodes; the 2 to 6 pods of the autoscaler are set in another file, hpa.yaml, and that horizontal autoscaler is the only scaling the project configures.

The EKS cluster: control plane, nodes and pods

A pod is one running copy of a container. A node is a machine that runs pods. The control plane is the part of Kubernetes that decides which pod goes on which node and notices when one dies. On EKS the control plane is AWS's job and the nodes are EC2 instances in your account.

An EKS cluster in us-east-1 with a control plane run by AWS and two m5.xlarge nodes holding the rag-api pods, opensearch-0, airflow and opensearch-dashboards (one possible placement); the internet reaches rag-api, airflow and the dashboards through three LoadBalancer Services, OpenSearch is reachable only inside the cluster, and the pods call services outside it such as ECR, Neon Postgres, Upstash Redis, Amazon Bedrock, Jina, Langfuse, Logfire and Grafana Cloud.

The video creates the cluster with eksctl from one file. The video names Terraform, the AWS CDK and Pulumi as other ways to do the same.

Shown as it ran in the video, not run here: it needs an AWS account and the eksctl tool.

yaml
apiVersion: eksctl.io/v1alpha5
kind: ClusterConfig

metadata:
  name: agentic-rag-cluster
  region: us-east-1
  version: "1.31"

iam:
  withOIDC: true

managedNodeGroups:
  - name: rag-workers
    instanceType: m5.xlarge   # 4 vCPU, 16 GB RAM
    minSize: 2                # Keep 2 nodes for pod anti-affinity spread
    maxSize: 4
    desiredCapacity: 2
    volumeSize: 50
    privateNetworking: true
  • managedNodeGroups asks for two m5.xlarge machines, each with 4 vCPUs and 16 GiB of memory. maxSize: 4 is only a ceiling: nothing in the project adds a third node by itself.
  • withOIDC: true lets a pod take an AWS role through its service account (IRSA) instead of carrying access keys.
  • privateNetworking: true gives the nodes no public address. Traffic comes in through a load balancer.
  • version: "1.31" is the Kubernetes version. The EKS console in the video shows a banner that extended support for 1.31 ends on November 26, 2026, so choose a version in standard support for a new cluster.

Deployment: the API pods

A Deployment says "keep N copies of this container running" and replaces them one by one when the image changes. This is the Deployment of the API, with the labels and the health checks trimmed to the lines discussed here.

yaml
apiVersion: apps/v1
kind: Deployment
metadata:
  name: rag-api
  namespace: production
spec:
  replicas: 2   # HPA (hpa.yaml) scales this between 2 and 6
  strategy:
    type: RollingUpdate
    rollingUpdate:
      maxUnavailable: 0
      maxSurge: 1
  template:
    spec:
      serviceAccountName: rag-api-sa
      containers:
        - name: rag-api
          image: <ACCOUNT_ID>.dkr.ecr.us-east-1.amazonaws.com/agentic-rag/api:latest
          envFrom:
            - secretRef:
                name: rag-app-secrets
          resources:
            requests:
              memory: "6Gi"
              cpu: "500m"
            limits:
              memory: "8Gi"
              cpu: "2000m"
          readinessProbe:
            httpGet:
              path: /api/v1/health
              port: 8000
            initialDelaySeconds: 40
            periodSeconds: 15
  • replicas: 2 keeps two API pods, so one node can fail and the API still answers.
  • maxUnavailable: 0, maxSurge: 1 is a rolling update that never drops below two pods: a third pod starts, passes its check, and only then an old one stops.
  • readinessProbe calls /api/v1/health. A pod gets traffic only after this check passes, first tried 40 seconds after start.
  • envFrom.secretRef loads every key of the Secret rag-app-secrets as an environment variable, the cluster's version of an .env file.
  • image points at a private registry (ECR). The account id in the address is replaced by <ACCOUNT_ID> here.

Requests and limits

The resources block holds the two numbers that decide how the cluster behaves under load.

A bar from 0 to 8Gi for one rag-api pod: 3.12Gi in use (52% of the request), a request of 6Gi that the scheduler reserves on a node, a limit of 8Gi above which the container is OOM-killed, and the autoscaler's memory target at 80% of the request, 4.8Gi.
  • A request is a reservation. The scheduler places a pod only on a node that still has the requested amount unreserved, whether or not the pod ever uses it. Autoscaler percentages are a share of the request.
  • A limit is a cap. A container that goes over its memory limit is killed (OOM kill). One that goes over its CPU limit is slowed down.
  • Between the two is burst room. A pod with a request below its limit may use more than it reserved when the node has memory to spare.
  • Neither number resizes anything. A request of 6Gi with a limit of 8Gi describes one pod of a fixed size. Changing a pod's size is a different operation, covered in Horizontal pod autoscaling (HPA).

Service types: how traffic reaches a pod

Pods come and go, and each new one has a new address. A Service is a stable name and port in front of all pods that carry a label.

yaml
apiVersion: v1
kind: Service
metadata:
  name: rag-api
  namespace: production
spec:
  type: LoadBalancer
  selector:
    app: rag-api         # Routes to pods with this label
  ports:
    - name: http
      port: 80           # clients connect here
      targetPort: 8000   # uvicorn listens here
Service typeReachable fromIn the project
ClusterIPInside the cluster onlyopensearch on port 9200, used by the API and Airflow
ClusterIP with clusterIP: None (headless)Inside the cluster, one DNS name per podopensearch-headless, which the StatefulSet needs
NodePortA fixed port on every nodeNot used
LoadBalancerThe internet, through a cloud load balancerThree: rag-api (80 to 8000), airflow (8080), opensearch-dashboards (5601)

StatefulSet: OpenSearch and its disk

An API pod holds no data, so any copy can replace any other. A search index is different: its pod needs the same disk and the same name after every restart. A StatefulSet gives each pod a fixed name (opensearch-0) and, through volumeClaimTemplates, its own disk that outlives the pod.

yaml
kind: StatefulSet
metadata:
  name: opensearch
spec:
  serviceName: opensearch-headless
  replicas: 1
  volumeClaimTemplates:
    - metadata:
        name: opensearch-data
      spec:
        accessModes: [ReadWriteOnce]
        storageClassName: gp3
        resources:
          requests:
            storage: 20Gi

The claim asks for a 20Gi gp3 volume. On EKS that needs the EBS CSI driver add-on and a gp3 storage class, and the project's setup script has a step for both.

The idle cluster: nodes and the autoscaler row · from the Complete AI Security Course in 8 Hours video · 7:14:30 to 7:15:13

This part of the video starts at 7:14:30. In the TARGETS column the 70% is the CPU target and the 80% is the memory target.

Reading the running cluster with kubectl

The outputs of two commands are on screen in the clip, both taken before any load. The first lists what each node is using.

bash
kubectl top nodes
Captured from a real run
NAME     CPU(cores)   CPU(%)   MEMORY(bytes)   MEMORY(%)
node-a   263m         6%       6184Mi          42%
node-b   52m          1%       6578Mi          44%

The second watches the autoscaler of the API. Both pods are idle: CPU is at 2% to 3% of its request and memory at 52%.

bash
kubectl get hpa -n production -w
Captured from a real run
NAME          REFERENCE            TARGETS                        MINPODS   MAXPODS   REPLICAS   AGE
rag-api-hpa   Deployment/rag-api   cpu: 3%/70%, memory: 52%/80%   2         6         2          3h46m
rag-api-hpa   Deployment/rag-api   cpu: 2%/70%, memory: 52%/80%   2         6         2          3h46m

Checking the manifests without a cluster

None of the files above can be applied without AWS, but they can be checked on any machine. The video applies its files with kubectl apply; the code below loads the same three objects into the typed classes of the official Kubernetes Python client, with no cluster and no network, to read the values back and to see which typing mistakes it catches. It needs the client and a YAML parser.

pip install kubernetes==37.0.0 pyyaml==6.0.3

One line per object

Each class has a from_dict method that takes the parsed YAML and returns a typed object, or raises an error.

python
import yaml
from kubernetes import client

doc = yaml.safe_load(open("deployment.yaml"))
dep = client.V1Deployment.from_dict(doc)
dep.spec.replicas          # a typed attribute, not a dict lookup
ExampleThe project's manifests, shortened, parsed with the Kubernetes Python client
import yaml
from kubernetes import client

MANIFESTS = """
apiVersion: apps/v1
kind: Deployment
metadata: {name: rag-api, namespace: production}
spec:
  replicas: 2
  selector: {matchLabels: {app: rag-api}}
  strategy:
    type: RollingUpdate
    rollingUpdate: {maxUnavailable: 0, maxSurge: 1}
  template:
    metadata: {labels: {app: rag-api}}
    spec:
      containers:
        - name: rag-api
          image: <ACCOUNT_ID>.dkr.ecr.us-east-1.amazonaws.com/agentic-rag/api:latest
          resources:
            requests: {memory: 6Gi, cpu: 500m}
            limits: {memory: 8Gi, cpu: 2000m}
          readinessProbe:
            httpGet: {path: /api/v1/health, port: 8000}
            initialDelaySeconds: 40
---
apiVersion: v1
kind: Service
metadata: {name: rag-api, namespace: production}
spec:
  type: LoadBalancer
  selector: {app: rag-api}
  ports: [{port: 80, targetPort: 8000}]
---
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata: {name: rag-api-hpa, namespace: production}
spec:
  scaleTargetRef: {apiVersion: apps/v1, kind: Deployment, name: rag-api}
  minReplicas: 2
  maxReplicas: 6
  metrics:
    - type: Resource
      resource: {name: cpu, target: {type: Utilization, averageUtilization: 70}}
    - type: Resource
      resource: {name: memory, target: {type: Utilization, averageUtilization: 80}}
"""
MODELS = {"Deployment": client.V1Deployment, "Service": client.V1Service,
          "HorizontalPodAutoscaler": client.V2HorizontalPodAutoscaler}

objs = {}
for doc in yaml.safe_load_all(MANIFESTS):
    objs[doc["kind"]] = MODELS[doc["kind"]].from_dict(doc)      # typed object, or an error
    print(f"{doc['apiVersion']:15} {doc['kind']:24} -> {type(objs[doc['kind']]).__name__}")

dep, svc, hpa = objs["Deployment"], objs["Service"], objs["HorizontalPodAutoscaler"]
box = dep.spec.template.spec.containers[0]
roll = dep.spec.strategy.rolling_update
print("replicas  :", dep.spec.replicas, "| surge", roll.max_surge, "| unavailable", roll.max_unavailable)
print("requests  :", box.resources.requests)
print("limits    :", box.resources.limits)
print("readiness :", box.readiness_probe.http_get.path, "after", box.readiness_probe.initial_delay_seconds, "s")
print("service   :", svc.spec.type, [(p.port, p.target_port) for p in svc.spec.ports])
print("autoscaler:", hpa.spec.min_replicas, "to", hpa.spec.max_replicas, "pods,",
      [(m.resource.name, m.resource.target.average_utilization) for m in hpa.spec.metrics])

first = MANIFESTS.split("---")[0]
try:
    client.V1Deployment.from_dict(yaml.safe_load(first.replace("replicas: 2", "replicas: two")))
except ValueError as err:
    lines = str(err).splitlines()
    print("wrong type:", lines[0], "|", lines[2].strip())
misspelt = client.V1Deployment.from_dict(yaml.safe_load(first.replace("replicas: 2", "replcas: 2")))
print("wrong key : replicas is", misspelt.spec.replicas)

What the typed objects show

  • All three kinds load: apps/v1 Deployment, v1 Service and autoscaling/v2 HorizontalPodAutoscaler are current API versions in this client.
  • The numbers read back as written: 2 replicas, a surge of 1, requests of 6Gi and 500m, limits of 8Gi and 2000m, a Service from port 80 to 8000, and an autoscaler from 2 to 6 pods on CPU 70 and memory 80.
  • A wrong type is caught: replicas: two raises a validation error that names the field's class and the bad value.
  • A misspelt key is not: replcas: 2 loads without complaint and replicas reads None. The client ignores keys it does not know, so a typed load is a first check, not a full one.

What two pods reserve on two nodes

With the requests known, a few lines of arithmetic show how full the cluster already is before any test. The requests and limits are those of the four manifests of the namespace.

ExampleRequests, limits and node capacity from the manifests and the kubectl top reading
# requests and limits per pod, as the four manifests set them (cpu cores, memory GiB)
PODS = {
    "rag-api":               {"count": 2, "req": (0.5, 6.0), "lim": (2.0, 8.0)},
    "airflow":               {"count": 1, "req": (0.5, 2.0), "lim": (2.0, 5.0)},
    "opensearch":            {"count": 1, "req": (0.5, 2.0), "lim": (2.0, 3.0)},
    "opensearch-dashboards": {"count": 1, "req": (0.2, 0.5), "lim": (0.5, 1.0)},
}


def total(kind, i):
    return sum(p["count"] * p[kind][i] for p in PODS.values())


print("namespace totals")
print(f"  cpu    requests {total('req', 0):.2f} cores   limits {total('lim', 0):.1f} cores")
print(f"  memory requests {total('req', 1):.1f} GiB    limits {total('lim', 1):.0f} GiB")

# kubectl top nodes in the video: memory in use and its share of what pods may use
nodes = {"node-a": (6184, 42), "node-b": (6578, 44)}
low = max(used / ((pct + 0.5) / 100) for used, pct in nodes.values())
high = min(used / ((pct - 0.5) / 100) for used, pct in nodes.values())
print(f"allocatable memory per node: {low / 1024:.2f} to {high / 1024:.2f} GiB of the 16 GiB instance")

node = (low + high) / 2 / 1024
print(f"both nodes together: {2 * node:.1f} GiB, of which the requests reserve {total('req', 1):.1f} GiB")
used = sum(u for u, _ in nodes.values()) / 1024
print(f"memory in use on both nodes: {used:.1f} GiB")
print(f"one API pod: request 6 GiB, in use at 52% of the request: {0.52 * 6:.2f} GiB")
print(f"reserved and unused by the two API pods: {2 * (6 - 0.52 * 6):.2f} GiB")
  • The totals match the dashboard. The Grafana panel in the video shows 16.5 GiB of memory requests, 25 GiB of memory limits and 2.20 cores of CPU requests for the namespace, the same sums.
  • A node offers about 14.5 GiB, not 16: the system keeps a share. The two nodes together offer about 29.0 GiB.
  • More than half is already reserved. 16.5 of 29.0 GiB is spoken for by four workloads, while the nodes report 12.5 GiB in use. One API pod uses about 3.12 GiB of the 6 GiB it reserves, so the two API pods hold 5.76 GiB that is reserved and unused.

The CD pipeline: from a push to a rollout

The video does not run kubectl apply by hand for each release. A GitHub Actions workflow in the repository builds the images and applies the manifests on every push to a deploy branch.

The CD workflow as five steps: a git push to a deploy branch, a first job that builds the API and Airflow images and pushes them to ECR tagged with the commit SHA, and a second job that runs kubectl apply for the Secret, OpenSearch, the API with its HPA and Airflow, then waits on kubectl rollout status for up to 300 seconds.

Shown as it ran in the video, not run here: it needs GitHub Actions and AWS credentials stored as repository secrets.

yaml
jobs:
  build-and-push:
    name: Build & Push Docker Images to ECR
    # builds ./Dockerfile and ./airflow/Dockerfile, tags each with the commit SHA
  deploy:
    name: Deploy to EKS Production
    needs: build-and-push
    steps:
      - run: kubectl apply -f deployment/eks/namespace.yaml
      - run: kubectl apply -f deployment/k8s/api/service.yaml
      - run: kubectl apply -f deployment/k8s/api/deployment.yaml
      - run: kubectl apply -f deployment/k8s/api/hpa.yaml
      - run: kubectl rollout status deployment/rag-api -n production --timeout=300s

The run shown in the video lists the deploy job as succeeded in 34 seconds. kubectl rollout status is what makes a failed release visible: it waits until the new pods are ready and fails the job after 300 seconds if they are not. To go back, kubectl rollout undo deployment/rag-api returns to the previous version.

What this deployment leaves open

Reading the same files as an attacker would gives a short list. Each line is in the repository's manifests.

  • Three public doors over plain HTTP. The API, the Airflow UI and OpenSearch Dashboards each get their own internet-facing load balancer with no TLS. One Ingress with a certificate in front of ClusterIP Services is the usual design, and the repository's own comments recommend it for production.
  • No authentication on the API, including the /mcp endpoint from MCP server for an agentic RAG API.
  • OpenSearch runs with its security plugin disabled (DISABLE_SECURITY_PLUGIN=true), which is acceptable only while it stays unreachable from outside the cluster.
  • Long-lived AWS keys. The Secret still carries a Bedrock access key although the service account can take a role, and the workflow signs in with static keys instead of a short-lived OIDC role.
  • A Kubernetes Secret is only base64-encoded unless encryption at rest is turned on, so who can read Secrets in the namespace matters.

Deployment vs StatefulSet

DeploymentStatefulSet
Pod namesRandom suffix, such as rag-api-7948f7fb7-hq8lqFixed and numbered, such as opensearch-0
StorageNone of its ownOne disk per pod, kept across restarts
Replacing a podAny new copy will doThe same name and the same disk come back
In the projectrag-api, airflow, opensearch-dashboardsopensearch

Where you use Amazon EKS

  • Several services that scale differently, as here: an API with two or more copies, a scheduler with one, a search index with a disk.
  • Teams that already run Kubernetes and want the same manifests on AWS.
  • Not for one small container. The video itself says a simpler service such as ECS with Fargate is enough for many applications; a cluster adds a control plane to pay for and to keep patched.
Watch out. A request reserves memory whether the pod uses it or not. This API reserves 6Gi and uses about 3.12Gi, so the two pods hold back 5.76 GiB that no other pod can be scheduled into. Set requests from measured usage, then leave headroom in the limit.
Try it yourself
  • In the first example, change memory: 6Gi under requests to memory: 3Gi and run it again: the typed object now prints 'memory': '3Gi'.
  • In the second example, change the count of rag-api from 2 to 3: memory requests rise from 16.5 to 22.5 GiB and limits from 25 to 33 GiB.
  • Replace type: LoadBalancer with type: ClusterIP in the Service and run the first example: the service line changes, and on a real cluster the API would no longer be reachable from outside.

Slow is fine. Stopping is the only problem.