Cloud-deployed agent platform: live, traced and affordable
The month 4 project of the AI Forward Deployed Engineer roadmap: take the ops workflow agent from month 3 and run it on a cloud the way a client would, deployed by a pipeline, traced on every run, with a cost you can show and a fallback when a model is down.
The problem
An agent that runs on your laptop is a demo. To be used it has to run behind an endpoint, reach models through keys it does not hold in code, survive a provider outage, and deploy from CI rather than from your terminal. And when the client asks what it cost last week, or why it answered something on Tuesday, you need the answer to hand.
Architecture
The agent gets an operational shell. The repository's agent reaches models through a gateway that holds the keys and falls back when one is down, runs behind a server, ships through CI that tests, evaluates and then deploys, and is served as an endpoint someone can call. Tracing sits beside every run and feeds the cost dashboard.
Route every model call through the gateway, never directly. That one indirection lets you swap a model, add a fallback and track cost without touching the agent, and it keeps the keys in one place.
What it draws on
- LangGraph: streaming and serving the agent
- Claude Code: agents in GitHub Actions
- RAGAS: evals in CI
- LangGraph: tracing with LangSmith
- LangChain: retries and fallback models
- FastAPI in production, Docker, Kubernetes, a cloud, GitHub Actions pipelines, Terraform, Langfuse and the LiteLLM gateway: courses upcoming. Use their official docs; until the LiteLLM course opens, LangChain's fallback middleware gives you the fallback model
What done looks like
| Requirement | Done when |
|---|---|
| Serving | The agent runs on AWS, Azure or GCP behind an authenticated endpoint |
| CI/CD | Every pull request runs tests and evals; a merge to main deploys to staging, and a tagged release to prod |
| Secrets | No key in the repository or the image; keys come from the cloud's secret store |
| Tracing | Every run is traced, and one user report can be traced to its run in under a minute |
| Cost dashboard | Tokens and cost per day and per ticket type are visible on one page |
| Fallback | Turning off the main model switches traffic to the backup with no code change |
Where to start
Put the fallback in first, even locally: every model call goes through one place that reads its keys from the environment. Then containerise the service, get one test running in CI, and deploy the smallest version to staging. Add tracing before you add features, so the cost dashboard has data from the first day.
Every expert started right here.