Output guardrails on the final answer
An output guardrail is a check that runs on the agent's final answer after the model produces it, and it can trip a tripwire that blocks that answer.
Last updated: 28 Sep, 2026 · openai-agents 0.22.3
An output guardrail runs on the final output of the last agent in a run. It reads the finished text and decides whether the caller is allowed to see it. This is where you catch an answer that leaked something it should not.
The @output_guardrail decorator
from agents import output_guardrail, GuardrailFunctionOutput
@output_guardrail
async def guard(ctx, agent, answer): # runs after the model
bad = "card" in answer.lower()
return GuardrailFunctionOutput(
output_info={"leaked_card": bad},
tripwire_triggered=bad, # True blocks the answer
)A tool that returns sensitive text
To give the guardrail something to catch, a tool returns an order line that includes a card detail. The model will hand this text back as the final answer.
from agents import function_tool
@function_tool
def lookup_order(order_id: str) -> str:
"Look up an order by id."
return f"Order {order_id}: shipped, card ending 4242."The output guardrail
The function takes the context, the agent, and the final answer string. It trips the tripwire when the answer mentions a card.
@output_guardrail
async def no_card_numbers(ctx, agent, answer):
leaked = "card" in answer.lower()
return GuardrailFunctionOutput(
output_info={"leaked_card": leaked},
tripwire_triggered=leaked,
)View the code here
"""A deterministic stand-in Model for the OpenAI Agents SDK course.
It implements the Model interface so an Agent runs with no API key. It reads the
last user message and the tool results out of `input`, and returns either a tool
call, a handoff, or a final message. Swap it for a real model at the end.
"""
from agents.models.interface import Model
from agents.items import ModelResponse
from agents.usage import Usage
from openai.types.responses import (
ResponseOutputMessage, ResponseOutputText, ResponseFunctionToolCall,
ResponseCompletedEvent, ResponseTextDeltaEvent, Response,
)
def _message(text):
return ResponseOutputMessage(
id="msg", role="assistant", type="message", status="completed",
content=[ResponseOutputText(text=text, type="output_text", annotations=[])],
)
def _tool_call(name, arguments, call_id="call_1"):
return ResponseFunctionToolCall(
id="fc", call_id=call_id, name=name, arguments=arguments, type="function_call",
)
def last_user_text(input):
if isinstance(input, str):
return input
for item in reversed(input):
d = item if isinstance(item, dict) else item.__dict__
if d.get("role") == "user":
content = d.get("content")
if isinstance(content, str):
return content
if isinstance(content, list):
for part in content:
pd = part if isinstance(part, dict) else part.__dict__
if pd.get("text"):
return pd["text"]
return ""
def tool_output(input):
if isinstance(input, str):
return None
for item in reversed(input):
d = item if isinstance(item, dict) else item.__dict__
if d.get("type") == "function_call_output":
return d.get("output")
return None
class ShopModel(Model):
async def get_response(self, system_instructions, input, model_settings, tools,
output_schema, handoffs, tracing, *, previous_response_id=None,
conversation_id=None, prompt=None):
result = tool_output(input)
if result is not None:
# A handoff transfer looks like {"assistant": "..."}; the specialist answers for real.
if result.strip().startswith('{"assistant"'):
if "refund" in (system_instructions or "").lower():
return ModelResponse(
output=[_message(
"Your refund is approved and will be processed in 5 to 7 days.")],
usage=Usage(), response_id=None)
return ModelResponse(output=[_message("Handled by the specialist.")],
usage=Usage(), response_id=None)
return ModelResponse(output=[_message(result)], usage=Usage(), response_id=None)
text = last_user_text(input).lower()
if handoffs and "refund" in text:
return ModelResponse(output=[_tool_call(handoffs[0].tool_name, "{}")],
usage=Usage(), response_id=None)
if tools and "order" in text:
return ModelResponse(output=[_tool_call("lookup_order", '{"order_id": "A17"}')],
usage=Usage(), response_id=None)
return ModelResponse(output=[_message("How can I help with your order?")],
usage=Usage(), response_id=None)
async def stream_response(self, system_instructions, input, model_settings, tools,
output_schema, handoffs, tracing, *, previous_response_id=None,
conversation_id=None, prompt=None):
text = "How can I help with your order?"
for i, word in enumerate(text.split()):
yield ResponseTextDeltaEvent(
type="response.output_text.delta", delta=word + " ",
content_index=0, item_id="msg", output_index=0,
sequence_number=i, logprobs=[],
)
response = Response(
id="r", created_at=0.0, model="shop-standin", object="response",
output=[_message(text)], parallel_tool_calls=False,
tool_choice="auto", tools=[],
)
yield ResponseCompletedEvent(type="response.completed", response=response,
sequence_number=99)
Attaching it and catching the tripwire
Pass the guardrail in output_guardrails. When it trips, the run raises OutputGuardrailTripwireTriggered, which carries the same result object.
from agents import Agent, Runner, OutputGuardrailTripwireTriggered
from shop_model import ShopModel
agent = Agent(name="Shop", instructions="Help with orders.", tools=[lookup_order],
model=ShopModel(), output_guardrails=[no_card_numbers])A leaked card number, blocked
The same pieces in one file. A greeting passes; an order lookup returns the card line and is blocked.
from agents import (Agent, Runner, function_tool, set_tracing_disabled,
output_guardrail, GuardrailFunctionOutput, OutputGuardrailTripwireTriggered)
from shop_model import ShopModel
set_tracing_disabled(True)
@function_tool
def lookup_order(order_id: str) -> str:
"Look up an order by id."
return f"Order {order_id}: shipped, card ending 4242."
@output_guardrail
async def no_card_numbers(ctx, agent, answer):
leaked = "card" in answer.lower()
return GuardrailFunctionOutput(
output_info={"leaked_card": leaked},
tripwire_triggered=leaked,
)
agent = Agent(name="Shop", instructions="Help with orders.", tools=[lookup_order],
model=ShopModel(), output_guardrails=[no_card_numbers])
for question in ["hello there", "where is my order?"]:
try:
result = Runner.run_sync(agent, question)
print("OK:", result.final_output)
except OutputGuardrailTripwireTriggered as e:
print("BLOCKED:", e.guardrail_result.output.output_info)OK: How can I help with your order?
BLOCKED: {'leaked_card': True}Why the order lookup was blocked
- The greeting returns a generic reply with no card in it, so the guardrail passes and you see the text.
- The order question makes the model call
lookup_order, whose result mentions a card. That text becomes the final answer. - The output guardrail reads that final answer, finds
card, and trips the tripwire, so the run raisesOutputGuardrailTripwireTriggeredinstead of returning the text.
Input guardrail vs output guardrail
| Guardrail | Runs on | When it runs |
|---|---|---|
input_guardrails | The user's input to the first agent | Before the model is called |
output_guardrails | The final answer of the last agent | After the model produces the answer |
When to guard the output
- Catching a reply that includes card numbers, tokens, or another person's details.
- Rejecting an answer whose format is wrong before it reaches the user.
- Adding a policy check that only the finished text can be judged against.
Related
- Previous: Input guardrails and the tripwire
- Next: Sessions: memory across runs with SQLiteSession
- Reference: Guardrails
- Change the banned token from
cardtoshippedand rerun the greeting. - Make
lookup_orderreturn a line with no card and confirm the order question now passes. - Add a second output guardrail and see which one's exception is raised first.
This is what real progress feels like.