Approval-gated agent: nothing risky without a human
An agent whose safe calls run freely and whose risky ones stop for a human, all validated and logged.
The problem
An agent that can move money, delete records, or run code is one hallucinated tool call away from real damage. Letting it act freely is reckless; making a human approve everything makes it useless. The line has to be drawn by risk.
You want an agent where safe calls run on their own and risky ones stop for a human, where a policy validates every call, and where anything that does run does so in a sandbox with an audit trail. Control without turning the agent back into a form.
Architecture
Layers of control between the request and the effect. Guardrails validate and score the call, a human approves the risky ones, the approved call runs in a sandbox, and the result is logged. Each layer only stops what it must, so the cheap safe path stays fast.
Decide risk by the tool and its arguments, not by the model's confidence. A refund of forty pounds and one of four thousand are the same call to the model; the gate is what tells them apart before either runs.
What it draws on
Everything here comes from this stage of the roadmap; the project is where those courses meet.
- OpenAI Agents SDK: guardrails and approvals
- LangGraph: pause, resume and rewind
- Pydantic AI: human approval before a refund
- Guardrails AI: validating what the model says
- Agent Governance Toolkit: policy on every tool call and an audit log
What done looks like
| Requirement | Done when |
|---|---|
| Input | Requests that are abusive or off-topic are stopped early |
| Approval | No refund runs without an explicit approval |
| Edit | A reviewer can change the amount before approving |
| Audit | Every request, decision and approver is logged |
| Kill switch | One setting stops all actions immediately |
Where to start
Start with the choke point: route every tool call through one function before anything runs. Add a default-deny policy, then a risk rule that sends only the dangerous calls to a human, then the sandbox, then the log. Build it so a new tool is dangerous until a rule says otherwise, not the other way round.
Every expert started right here.