Ops workflow agent: from ticket to approved action
The month 3 project of the AI Forward Deployed Engineer roadmap: an agent that reads an operations ticket, looks up what it needs in a database, drafts the action to take, and waits for a person to approve it. Every step of every run is logged, so any decision can be explained afterwards.
The problem
An operations team works a queue of tickets: a failed nightly job, a stuck order, an account locked by mistake. Each one means reading the ticket, checking a few tables, and deciding whether to retry, fix or escalate. A script that tries to do this grows into a tangle of flags nobody wants to touch at 3am, and a chatbot cannot be trusted to change anything.
You want an agent that does the reading and the lookups, proposes the action with its reasons, and changes nothing until a person says yes. When the team lead asks why it proposed something, the trace answers.
Architecture
The workflow is a LangGraph graph. A ticket comes in, the graph carries state through named steps, a branch decides between retry, fix and escalate, and the proposed action stops at an interrupt for approval. Only an approved action runs. Because the state is checkpointed, a run that dies halfway resumes from its last step.
Put everything a later step needs into the graph's state, not into local variables, so a resumed run picks up exactly where it stopped. Give the database tool read access only; the action that changes data runs after approval, as its own step.
What it draws on
- LangChain: the agent loop and tool design
- LangGraph: state, nodes, edges and checkpoints
- LangGraph: human-in-the-loop approvals
- MCP: put the database lookups on a server, if you want the tools outside the agent
- LangChain: guardrails and call limits
- LangGraph: the support agent project, a worked example of the same shape
- Tracing: LangSmith tracing is in month 4 (LangGraph: LangSmith). A JSON log of every node's input and output is enough for this project
What done looks like
| Requirement | Done when |
|---|---|
| Routing | Each of at least three ticket types takes its own path through the graph |
| Data | Lookups come from a small SQL database through a read-only tool, not from the prompt |
| Drafts | Each ticket ends with a drafted action and the reasons for it |
| Approval gate | Nothing that changes data runs until a person approves, edits or rejects it |
| Trace logging | Every run records each step's input, output and decision, and one ticket can be replayed from the log |
| Resume | Stopping and restarting mid-ticket continues where it left off |
| Limits | A run that loops is stopped by a recursion or call limit and reported, not left running |
Where to start
Draw the graph before you write a node. Name every step and branch, and decide what piece of state each one reads and writes. Build the smallest version first: one ticket type, one lookup, one drafted action, one interrupt. Add the checkpointer early so you can kill a run and resume it, then add the other ticket types.
Every expert started right here.