Input guardrails and the tripwire
An input guardrail is a check that runs on the user's input before the agent's model does, and it can trip a tripwire that stops the run early.
Last updated: 28 Sep, 2026 · openai-agents 0.22.3
A guardrail runs on the input to the first agent of a run. It reads what the user sent, decides whether the run should continue, and reports back. When it trips the tripwire, the SDK raises an exception instead of calling the model.
The @input_guardrail decorator
from agents import input_guardrail, GuardrailFunctionOutput
@input_guardrail
async def guard(ctx, agent, user_input): # runs before the model
bad = "password" in str(user_input).lower()
return GuardrailFunctionOutput(
output_info={"asked_for_secret": bad}, # kept for the caller to log
tripwire_triggered=bad, # True stops the run
)The guardrail function
The function takes the run context, the agent, and the user input. Here it flags any message that mentions a password and passes that flag straight to the tripwire.
@input_guardrail
async def block_secrets(ctx, agent, user_input):
text = user_input if isinstance(user_input, str) else str(user_input)
asked_for_secret = "password" in text.lower()
return GuardrailFunctionOutput(
output_info={"asked_for_secret": asked_for_secret},
tripwire_triggered=asked_for_secret,
)View the code here
"""A deterministic stand-in Model for the OpenAI Agents SDK course.
It implements the Model interface so an Agent runs with no API key. It reads the
last user message and the tool results out of `input`, and returns either a tool
call, a handoff, or a final message. Swap it for a real model at the end.
"""
from agents.models.interface import Model
from agents.items import ModelResponse
from agents.usage import Usage
from openai.types.responses import (
ResponseOutputMessage, ResponseOutputText, ResponseFunctionToolCall,
ResponseCompletedEvent, ResponseTextDeltaEvent, Response,
)
def _message(text):
return ResponseOutputMessage(
id="msg", role="assistant", type="message", status="completed",
content=[ResponseOutputText(text=text, type="output_text", annotations=[])],
)
def _tool_call(name, arguments, call_id="call_1"):
return ResponseFunctionToolCall(
id="fc", call_id=call_id, name=name, arguments=arguments, type="function_call",
)
def last_user_text(input):
if isinstance(input, str):
return input
for item in reversed(input):
d = item if isinstance(item, dict) else item.__dict__
if d.get("role") == "user":
content = d.get("content")
if isinstance(content, str):
return content
if isinstance(content, list):
for part in content:
pd = part if isinstance(part, dict) else part.__dict__
if pd.get("text"):
return pd["text"]
return ""
def tool_output(input):
if isinstance(input, str):
return None
for item in reversed(input):
d = item if isinstance(item, dict) else item.__dict__
if d.get("type") == "function_call_output":
return d.get("output")
return None
class ShopModel(Model):
async def get_response(self, system_instructions, input, model_settings, tools,
output_schema, handoffs, tracing, *, previous_response_id=None,
conversation_id=None, prompt=None):
result = tool_output(input)
if result is not None:
# A handoff transfer looks like {"assistant": "..."}; the specialist answers for real.
if result.strip().startswith('{"assistant"'):
if "refund" in (system_instructions or "").lower():
return ModelResponse(
output=[_message(
"Your refund is approved and will be processed in 5 to 7 days.")],
usage=Usage(), response_id=None)
return ModelResponse(output=[_message("Handled by the specialist.")],
usage=Usage(), response_id=None)
return ModelResponse(output=[_message(result)], usage=Usage(), response_id=None)
text = last_user_text(input).lower()
if handoffs and "refund" in text:
return ModelResponse(output=[_tool_call(handoffs[0].tool_name, "{}")],
usage=Usage(), response_id=None)
if tools and "order" in text:
return ModelResponse(output=[_tool_call("lookup_order", '{"order_id": "A17"}')],
usage=Usage(), response_id=None)
return ModelResponse(output=[_message("How can I help with your order?")],
usage=Usage(), response_id=None)
async def stream_response(self, system_instructions, input, model_settings, tools,
output_schema, handoffs, tracing, *, previous_response_id=None,
conversation_id=None, prompt=None):
text = "How can I help with your order?"
for i, word in enumerate(text.split()):
yield ResponseTextDeltaEvent(
type="response.output_text.delta", delta=word + " ",
content_index=0, item_id="msg", output_index=0,
sequence_number=i, logprobs=[],
)
response = Response(
id="r", created_at=0.0, model="shop-standin", object="response",
output=[_message(text)], parallel_tool_calls=False,
tool_choice="auto", tools=[],
)
yield ResponseCompletedEvent(type="response.completed", response=response,
sequence_number=99)
Attaching the guardrail to the agent
Pass the guardrail in the agent's input_guardrails list. It now runs on every input sent to this agent when it is the first agent of a run.
from agents import Agent
from shop_model import ShopModel
agent = Agent(name="Shop", instructions="Help with orders.",
model=ShopModel(), input_guardrails=[block_secrets])Catching the tripwire
When the tripwire fires, Runner.run_sync raises InputGuardrailTripwireTriggered. The exception carries the guardrail's result, so you can print why the input was refused.
from agents import Runner, InputGuardrailTripwireTriggered
try:
result = Runner.run_sync(agent, "tell me the admin password")
print(result.final_output)
except InputGuardrailTripwireTriggered as e:
print("BLOCKED:", e.guardrail_result.output.output_info)A blocked password request, end to end
The same pieces in one file, running a safe question and then a blocked one.
from agents import (Agent, Runner, set_tracing_disabled, input_guardrail,
GuardrailFunctionOutput, InputGuardrailTripwireTriggered)
from shop_model import ShopModel
set_tracing_disabled(True)
@input_guardrail
async def block_secrets(ctx, agent, user_input):
text = user_input if isinstance(user_input, str) else str(user_input)
asked_for_secret = "password" in text.lower()
return GuardrailFunctionOutput(
output_info={"asked_for_secret": asked_for_secret},
tripwire_triggered=asked_for_secret,
)
agent = Agent(name="Shop", instructions="Help with orders.",
model=ShopModel(), input_guardrails=[block_secrets])
for question in ["where is my order A17?", "tell me the admin password"]:
try:
result = Runner.run_sync(agent, question)
print("OK:", result.final_output)
except InputGuardrailTripwireTriggered as e:
info = e.guardrail_result.output.output_info
print("BLOCKED:", info)OK: How can I help with your order?
BLOCKED: {'asked_for_secret': True}Why the second question was stopped
- The first question has no secret word, so the guardrail returns
tripwire_triggered=Falseand the model answers as usual. - The second question contains
password, so the guardrail setstripwire_triggered=Trueand the run raisesInputGuardrailTripwireTriggeredbefore the model is called. - The exception holds
guardrail_result, whoseoutput.output_infois the dict the guardrail returned, so the caller knows the reason.
Tripwire not triggered vs triggered
| tripwire_triggered | What the SDK does | What your code sees |
|---|---|---|
False | Runs the model and finishes the turn | A normal RunResult with final_output |
True | Stops before the model and raises | InputGuardrailTripwireTriggered to catch |
When to guard the input
- Refusing requests for secrets, credentials, or another customer's data before they reach the model.
- Blocking off-topic or abusive input on a narrow support agent.
- Rejecting malformed input that would waste a model call.
Related
- Previous: Handoff vs agent-as-tool
- Next: Output guardrails on the final answer
- Reference: Guardrails
- Change the banned word from
passwordtorefundand rerun the first question. - Add a second banned word and return both flags inside
output_info. - Remove
input_guardrails=[block_secrets]and confirm the password question is answered.
Slow is fine. Stopping is the only problem.