Guardrails
A guardrail is a check around an agent that catches unwanted input or output: PIIMiddleware rewrites card numbers before the model sees them, and a custom hook can end a request early.
Last updated: 27 Sep, 2026 · LangChain 1.4
Checks at the edges of the agent
A guardrail helps you build a safe, compliant AI application by validating and filtering content at key points in the agent's run: before the agent starts (an input guardrail), after it completes (an output guardrail), or around model and tool calls. In LangChain every guardrail is middleware. Common uses are preventing PII leaks, blocking prompt injection and dangerous requests, requiring approval for financial operations, and making every response meet a safety standard. There are two ways to decide whether something is allowed:
- Deterministic guardrails use rules: regular expressions, keyword lists, explicit checks. They are fast, predictable and free, but miss anything the rules did not foresee.
- Model-based guardrails ask a model, often a small cheap one, whether the content is safe. They catch subtle cases, but add time and cost to every request.
# Quick illustration of the two approaches
import re
# --- Deterministic approach ---
def deterministic_guardrail(text: str) -> bool:
"""Returns True if content is blocked."""
banned_keywords = ["hack", "exploit", "malware", "bomb"]
return any(kw in text.lower() for kw in banned_keywords)
test_inputs = [
"How do I hack into a database?",
"What is the capital of France?",
"Explain how malware spreads",
]
print("=== Deterministic Guardrail Demo ===")
for inp in test_inputs:
blocked = deterministic_guardrail(inp)
status = "🚫 BLOCKED" if blocked else "✅ ALLOWED"
print(f"{status}: {inp}")=== Deterministic Guardrail Demo === 🚫 BLOCKED: How do I hack into a database? ✅ ALLOWED: What is the capital of France? 🚫 BLOCKED: Explain how malware spreads
A keyword list is the simplest deterministic check: no model, only a fixed set of banned keywords. It blocks "hack" and "malware" and lets the geography question through. The video then runs the same three questions through a model-based check, gpt-4o-mini asked to reply safe or unsafe. That check reads the context: it flags the hacking question as unsafe but treats the question about how malware spreads as general information, which the keyword list blocked.
PII redaction with PIIMiddleware
LangChain ships PIIMiddleware for personally identifiable information. It detects emails, credit cards, IP addresses, MAC addresses and URLs, and applies one of four strategies: redact replaces the value with a label, mask puts stars over it, hash replaces it with a hash, and block raises an exception. It runs on the input before the agent does. The video's agent has a dummy customer_lookup tool and one PIIMiddleware per type: redact for emails, mask for credit cards, and block for API keys, found by a detector regular expression for 32 characters after sk-:
from langchain.agents import create_agent
from langchain.agents.middleware import PIIMiddleware
from langchain_core.tools import tool
# Define a simple dummy tool
@tool
def customer_lookup(query: str) -> str:
"""Look up customer information."""
return f"Customer record found for query: {query}"
# Create agent with PII Middleware
agent = create_agent(
model="groq:openai/gpt-oss-120b",
tools=[customer_lookup],
middleware=[
# Redact emails in user input before sending to model
PIIMiddleware(
"email",
strategy="redact",
apply_to_input=True,
),
# Mask credit cards in user input
PIIMiddleware(
"credit_card",
strategy="mask",
apply_to_input=True,
),
# Block API keys - raise error if detected
PIIMiddleware(
"api_key",
detector=r"sk-[a-zA-Z0-9]{32}",
strategy="block",
apply_to_input=True,
),
],
)
print("Agent with PII middleware created successfully!")Send the agent a message with an email address and a card number, then print the reply and the message the agent stored:
# Test PII Redaction
result = agent.invoke({
"messages": [{
"role": "user",
"content": "My email is john.doe@example.com and my card is 5105-1051-0510-5100. Can you help me?"
}]
})
print("=== Agent Response ===")
print(result["messages"][-1].content)
print("STORED:", result["messages"][0].content)=== Agent Response === I’m happy to help you with your account. To protect your security, I’ll never ask for your full card number or any other sensitive details. Could you let me know what you need assistance with (e.g., a recent transaction, a login problem, a billing question, etc.)? With a bit more information I can look up the relevant details and guide you through a solution. STORED: My email is [REDACTED_EMAIL] and my card is ****-****-****-5100. Can you help me?
The STORED line is what the agent kept: the email replaced by [REDACTED_EMAIL] and the card masked to its last four digits, both before the model was called, so the model never saw the real values. The block strategy raises an exception instead, so the video wraps the call in try and sends a made-up key that looks like an OpenAI key:
# Test API Key Blocking
try:
result = agent.invoke({
"messages": [{
"role": "user",
"content": "Here is my key: sk-abcdefghijklmnopqrstuvwxyz123456"
}]
})
except Exception as e:
print(f"🚫 Blocked as expected: {e}")🚫 Blocked as expected: Detected 1 instance(s) of api_key in text content
Now the shop. Customers paste card numbers into support chats. The number should not reach the model, the logs, or the saved conversation, and some requests should be turned away before the model runs at all.
PIIMiddleware and a before_agent check
PIIMiddleware("credit_card", strategy="redact") # redact | mask | hash | block
@before_agent(can_jump_to=["end"]) # a check that may end the run
def check(state, runtime):
... # return {"messages": [...], "jump_to": "end"} to stopAn agent that redacts cards
This lesson's agent answers order questions with lookup_order, the tool built in Tools: a function the model can call. Start the file with it.
from langchain.tools import tool
ORDERS = {"A17": "shipped on 3 March", "C40": "waiting for stock"}
@tool
def lookup_order(order_id: str) -> str:
"""Look up an order's shipping status by its id, such as A17."""
status = ORDERS.get(order_id)
return f"{order_id} {status}." if status else f"{order_id} is not an order we have."Add PIIMiddleware for credit cards to the shop agent: the Groq model and the lookup_order tool. It rewrites the message before the model call.
from langchain.agents import create_agent
from langchain.agents.middleware import PIIMiddleware
from langchain.chat_models import init_chat_model
model = init_chat_model("groq:openai/gpt-oss-120b", temperature=0) # uses your GROQ_API_KEY
# find card numbers in the input and replace them before the model sees them
agent = create_agent(model, tools=[lookup_order], middleware=[PIIMiddleware("credit_card")],
system_prompt="You are the support assistant for a small online shop. Answer in one or two short sentences, using only what the tools returned.")The card redacted before the model
Send a message with a card number, then print the customer's message as the agent's state now holds it. The model's reply is not the point here, so it is not printed.
result = agent.invoke({"messages": [{"role": "user", "content": "My card 4111 1111 1111 1111 was charged twice for A17"}]})
print(result["messages"][0].text)My card [REDACTED_CREDIT_CARD] was charged twice for A17
The card number was replaced before the model was called, in the conversation itself, so the saved state never holds it. By default PIIMiddleware checks the input only; apply_to_output and apply_to_tool_results turn on checks of the model's replies and of tool results.
Other strategies
The strategy decides what replaces the number. Run the same message through mask and hash.
for strategy in ["mask", "hash"]:
pii = PIIMiddleware("credit_card", strategy=strategy)
agent = create_agent(model, tools=[lookup_order], middleware=[pii],
system_prompt="You are the support assistant for a small online shop. Answer in one or two short sentences, using only what the tools returned.")
print(agent.invoke({"messages": [{"role": "user", "content": "My card 4111 1111 1111 1111 was charged twice for A17"}]})["messages"][0].text)My card **** **** **** 1111 was charged twice for A17 My card <credit_card_hash:6a7e0e79> was charged twice for A17
redact, the default, replaces the whole number. mask keeps the last four digits, which support staff often need. hash replaces it with a short hash, so two messages with the same card can be matched without storing it. block raises an error instead.
A custom guardrail with before_agent
A guardrail of your own is a middleware class. A before_agent hook works as an input filter: it runs as soon as the input arrives, before any model call, which suits keyword filtering, authentication checks, rate limiting and blocking whole categories of request. The video's ContentFilterMiddleware inherits AgentMiddleware and takes a list of banned keywords in __init__. Its before_agent method reads the first message, skips anything that is not from the human, lowercases the text and checks it against every keyword. @hook_config(can_jump_to=["end"]) lets it end the run early with its own reply. In a class, can_jump_to goes on @hook_config; as a decorated function, it goes straight on @before_agent, as the shop version below does.
from typing import Any
from langchain.agents.middleware import AgentMiddleware, AgentState, hook_config
from langgraph.runtime import Runtime
from langchain.agents import create_agent
from langchain_core.tools import tool
class ContentFilterMiddleware(AgentMiddleware):
"""
Deterministic guardrail: Block requests containing banned keywords.
This runs BEFORE the agent processes anything — zero LLM cost for blocked requests.
"""
def __init__(self, banned_keywords: list[str]):
super().__init__()
self.banned_keywords = [kw.lower() for kw in banned_keywords]
@hook_config(can_jump_to=["end"])
def before_agent(self, state: AgentState, runtime: Runtime) -> dict[str, Any] | None:
if not state["messages"]:
return None
first_message = state["messages"][0]
if first_message.type != "human":
return None
content = first_message.content.lower()
for keyword in self.banned_keywords:
if keyword in content:
print(f"🚫 Blocked — keyword detected: '{keyword}'")
return {
"messages": [{
"role": "assistant",
"content": (
"I cannot process requests containing inappropriate content. "
"Please rephrase your request."
)
}],
"jump_to": "end"
}
return None
@tool
def search_tool(query: str) -> str:
"""Search for information."""
return f"Results for: {query}"
# Create agent with content filter
filtered_agent = create_agent(
model="groq:openai/gpt-oss-120b",
tools=[search_tool],
middleware=[
ContentFilterMiddleware(
banned_keywords=["hack", "exploit", "malware", "jailbreak", "bypass"]
),
],
)
print("Content filter agent created!")# Test 2: Unsafe request, should be blocked
result = filtered_agent.invoke({
"messages": [{"role": "user", "content": "How do I hack into a server?"}]
})
print("🚫 Unsafe request response:")
print(result["messages"][-1].content)🚫 Blocked — keyword detected: 'hack' 🚫 Unsafe request response: I cannot process requests containing inappropriate content. Please rephrase your request.
"What is machine learning?" matches no keyword and goes on to the model as usual. "How do I hack into a server?" does: the filter printed its log line and returned a fixed reply, and the model was never called.
The video's filter reads the first message, which is fine for one-off requests. On a thread with memory, check the newest message instead, state["messages"][-1], as the shop version below does.
A check of your own
The same idea as a decorated function: return jump_to and the message the hook adds becomes the reply.
from langchain.agents.middleware import before_agent
from langchain.messages import AIMessage
@before_agent(can_jump_to=["end"])
def no_passwords(state, runtime):
if "password" in state["messages"][-1].text.lower():
answer = AIMessage("I cannot help with passwords. Please use the reset link.")
return {"messages": [answer], "jump_to": "end"}Wire the hook into the same shop agent as middleware.
from langchain.agents import create_agent
from langchain.chat_models import init_chat_model
model = init_chat_model("groq:openai/gpt-oss-120b", temperature=0) # uses your GROQ_API_KEY
agent = create_agent(model, tools=[lookup_order], middleware=[no_passwords],
system_prompt="You are the support assistant for a small online shop. Answer in one or two short sentences, using only what the tools returned.")result = agent.invoke({"messages": [{"role": "user", "content": "What is my password?"}]})
for message in result["messages"]:
print(f"{message.type:<5} {message.text}")human What is my password? ai I cannot help with passwords. Please use the reset link.
Two messages and no model call: the Groq model was never reached. A check like this costs nothing to run and cannot be talked out of its rule, which is its advantage over asking the model to refuse.
Layering guardrails
Guardrails stack in the middleware list, one by one, so each layer handles what the one before let through. The crash course's combined agent has tools for search and for sending email, and stacks a content filter as layer 1, PIIMiddleware as layer 2, human-in-the-loop approval next, and a model-based output safety check last. Laid out in full:
Read it as the order a request meets the checks. The input checks, layers 1 and 2, run in list order. The output checks are after hooks, which run in reverse, so in the middleware list they go in the opposite order to the one you want them to run in. SafetyGuardrailMiddleware is the video's own model-based check from the section before: an after_agent class that asks gpt-4o-mini whether each reply is SAFE or UNSAFE before it reaches the user.
What the guardrails did
- PIIMiddleware rewrites the message before the model call, so the card never reaches the model, the logs, or the saved conversation.
- A before_agent hook runs once, before the model, and can end the run with its own reply, so the rule costs no model call and cannot be talked around.
- can_jump_to lists the jumps the hook may make; without it the jump is ignored and the model runs anyway.
redact vs mask vs hash vs block
| strategy | What the model sees | What it keeps |
|---|---|---|
| redact | [REDACTED_CREDIT_CARD] | Nothing |
| mask | **** **** **** 1111 | The last four digits |
| hash | <credit_card_hash:...> | A repeatable token to match repeats |
| block | Nothing; an error is raised | Nothing |
Where guardrails fit
- Support chat where customers paste card numbers or emails that must not be stored.
- A hard rule the agent must never break, such as refusing password requests, enforced before any model call.
Related
- Previous: Edit and respond
- Next: Documents and splitting
- Reference: Middleware
- Use
strategy="block"and read the error. - Add
PIIMiddleware("email")to the list and include an email address in the message. - Remove
can_jump_tofromno_passwordsand count the AI messages.
This is what real progress feels like.