Agentic support capstone: one agent that uses everything
A support agent for an online shop that answers from policy documents, acts through MCP tools, remembers each customer, refuses what it should, asks a person before any refund, and comes with a score that proves it works.
The problem
A shop's support inbox gets the same few kinds of message all day: where is my order, what is the return policy, I want my money back. A plain chatbot gets these wrong in expensive ways. It invents a policy, forgets the customer told it their order number yesterday, answers a question about a competitor, or promises a refund nobody approved.
You want one agent that does the job properly: answers policy questions from the real documents, looks orders up and issues refunds through tools it does not own, remembers each customer between conversations, stays on topic, and never moves money without a person saying yes. Then you want a number that says how good its answers are, and a trace of every run so you can explain any one of them.
Architecture
A message passes the input rails, then reaches the LangGraph agent. The graph can search the policy documents, call the order and refund tools on an MCP server, and read and write the customer's memories. A refund goes to a person first. The answer passes the output rails before the customer sees it. Off to the side, RAGAS scores the agent on goldens and LangSmith records every run.
Keep the rails outside the graph. The input rail runs before the agent spends a model call, and the output rail checks the final answer whatever path the graph took. A rail inside the graph only guards the path it sits on.
It must
- Answer policy questions only from the policy documents, and say so when they do not cover the question
- Look up orders and issue refunds through tools on an MCP server, never through functions inside the agent
- Remember each customer across conversations with LangMem, keyed so one customer never sees another's memories
- Block off-topic and unsafe messages with a NeMo input rail, and check every answer with an output rail
- Pause before any refund and resume only on a person's approval, surviving a restart in between
- Report a RAGAS score on at least ten goldens, and trace every run in LangSmith
- Route model calls through a LiteLLM gateway once that course opens; until then, use LangChain's retry and fallback middleware
What it draws on
Every piece comes from a checkpoint earlier on the roadmap; the capstone is where they meet.
- LangGraph: the support agent, the graph you start from
- LangGraph: agentic RAG for the policy documents
- MCP: the support desk for the order and refund tools
- LangMem: the support assistant for memory across conversations
- NeMo Guardrails: the guarded assistant for input and output rails
- RAGAS: the TechNest evaluation for the score
- LangGraph: deploy and LangSmith for serving and tracing
What done looks like
| Requirement | Done when |
|---|---|
| Grounded | A policy question gets an answer from the documents, and an uncovered one gets a clear I don't know |
| Tools | Order lookups and refunds go through the MCP server, and the agent has no other way to reach them |
| Memory | A returning customer's order number or preference is used without being asked for again |
| Rails | An off-topic or unsafe message is refused before the graph runs, and a bad answer never reaches the customer |
| Approval | No refund runs without a person's yes, even after a restart |
| Score | RAGAS faithfulness, answer relevancy and context recall are reported on ten or more goldens |
| Traces | Any run can be found and read back step by step in LangSmith |
Where to start
Start from the LangGraph support agent and get one thing working end to end before adding the next. Write the goldens first, so every change after that has a score. Then add retrieval, move the tools onto the MCP server, add memory, and put the rails around the graph last, checking the score after each step so you know which piece moved it.
Every expert started right here.