AI SecurityNeMo Guardrails 0.24 · RAGAS 0.4 · OpenAI SDK 3.3 · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

Input and output rails

Input and output rails are the checks NeMo Guardrails runs on every user message before the model reads it and on every bot reply before the user sees it.

Last updated: 09 Oct, 2026 · NeMo Guardrails 0.24

The rails in Topic, jailbreak and sensitive-topic rails all depend on an intent that a model reads from the message. Some checks should not depend on a model at all: a social security number in a message, or a password in a reply, can be found with a pattern. This page places every kind of rail on the message path and runs the pattern checks of the video's demo app.

Rails on the input side, the output side and custom logic · from the Complete AI Security Course in 8 Hours video · 55:23 to 56:01

This part of the video starts at 0:55:23. Rails are rules and regulations, and the question is where they are applied. They are applied on the input side and on the output side, and you can also write your own Python logic for a rail.

NeMo Guardrails documents five rail types: input, dialog, retrieval, execution and output rails. Custom Python logic is a custom action, a function that a flow of any of these types can run.

The five rail types on the message path

One lane from the user to the bot response with five rail types in order: input rails check the user message, dialog rails decide the next step from the intent, retrieval rails check retrieved chunks in a RAG app, execution rails check the input and output of actions and tools, and output rails check the bot message before the user sees it. Input and output rails each have a stop exit.
Rail typeRunsIn the video's demo app
Input railsOn every user message, before anything elseThe PII check and the urgency check
Dialog railsAfter the intent is known, to decide what the bot does nextThe topic, jailbreak and sensitive-topic flows, and the scripted greetings
Retrieval railsOn the chunks a RAG app retrieved, before they reach the promptNot used
Execution railsOn the input and output of the actions and tools the bot callsNot used
Output railsOn every bot message, before the user sees itThe output sanitizer

Input and output rails are listed in the YAML, under rails.input.flows and rails.output.flows, and run for every message. Dialog rails are the ordinary flows that start with a user ... line and run only when that intent is found.

Dialog rails: scripted greetings · from the Complete AI Security Course in 8 Hours video · 32:16 to 33:53

This part of the video starts at 0:32:16. Every chatbot gets the same small talk: hi, bye, good morning, thank you. With a dialog rail the bot gives the same greeting to everyone who says hi and the same goodbye to everyone who says bye, so the wording is in your hands and not the model's. The video calls this governing the system.

A scripted greeting still costs one LLM call by default, the call that reads the intent; what it saves is the calls that would generate the answer.

Running a dialog rail

The examples from here on run under the setup code of NeMo Guardrails (the two import lines, the YAML and SEARCH strings, the AllExamples class, build_rails and chat) and its TOPIC string: paste each one below them in one file. The greeting block is from the video's repo. The video's app ran it on llama-3.3-70b-versatile, since retired on Groq; the runs below use openai/gpt-oss-120b.

ExampleAPI keyFrom the video, run on Groq (openai/gpt-oss-120b)
GREETING = '''
define user express greeting
  "hello"
  "hi"
  "hey"
  "good morning"
  "what's up"
  "howdy"

define bot express greeting
  "Hello! I'm your Enterprise IT Assistant. I specialise in Kubernetes, Intel hardware, and enterprise networking. What can I help you with today?"

define flow greeting
  user express greeting
  bot express greeting
  stop
'''

rails = build_rails(TOPIC + GREETING, model="openai/gpt-oss-120b")

for message in ["hi there", "what is VLAN tagging?"]:
    reply = rails.generate(messages=[{"role": "user", "content": message}])["content"]
    print(message)
    print("  LLM calls:", [call.task for call in rails.explain().llm_calls])
    print("  Bot:", " ".join(reply.split())[:70])
  • hi there got the scripted greeting after one LLM call, generate_user_intent. The greeting text cost no generation, but the message was still read by the model once.
  • The VLAN question took all three calls and was answered by the model.

Custom actions: Python inside a rail

A custom action is a Python function that a Colang flow runs with execute. The video's app has three, all built on regular expressions and keyword lists. Two run on the input: one looks for PII, personally identifiable information such as an e-mail address or a card number, and one looks for words that signal a production emergency. The third runs on the output and looks for credentials and exploit terms in the bot's reply.

The patterns from the video's repo

The repo writes each pattern list inside its action. Here the same patterns sit in three plain functions, so they can be tested without NeMo and without a model.

python
import re

PII_PATTERNS = {
    "email":       r"\b[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Za-z]{2,}\b",
    "phone":       r"\b(\+\d{1,2}\s?)?\(?\d{3}\)?[\s.-]?\d{3}[\s.-]?\d{4}\b",
    "ssn":         r"\b\d{3}-\d{2}-\d{4}\b",
    "api_key":     r"(api[_\s-]?key|token|secret)[:\s]+[A-Za-z0-9_\-]{10,}",
    "credit_card": r"\b\d{4}[\s-]\d{4}[\s-]\d{4}[\s-]\d{4}\b",
}
URGENT_KEYWORDS = ["outage", "down", "crash", "critical", "emergency", "not working", "urgent", "p0", "p1"]
OUTPUT_PATTERNS = {
    "hardcoded_credential": r"(?i)(password|passwd|secret|api[_\-]?key|token)\s*[:=]\s*['\"]?\w{4,}",
    "private_key":          r"-----BEGIN.{0,20}PRIVATE KEY-----",
    "exploit_technique":    r"(?i)\b(reverse.?shell|bind.?shell|shellcode|meterpreter)\b",
}


def find_pii(text):
    return [name for name, pattern in PII_PATTERNS.items() if re.search(pattern, text, re.IGNORECASE)]


def find_urgent(text):
    return [keyword for keyword in URGENT_KEYWORDS if keyword in text.lower()]


def find_leaks(text):
    return [name for name, pattern in OUTPUT_PATTERNS.items() if re.search(pattern, text)]

Testing the patterns on their own

The first three PII messages and the first two urgency messages are test prompts of the video's app. The others are added to probe the edges.

ExampleThe video's patterns, run in plain Python
import re

PII_PATTERNS = {
    "email":       r"\b[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Za-z]{2,}\b",
    "phone":       r"\b(\+\d{1,2}\s?)?\(?\d{3}\)?[\s.-]?\d{3}[\s.-]?\d{4}\b",
    "ssn":         r"\b\d{3}-\d{2}-\d{4}\b",
    "api_key":     r"(api[_\s-]?key|token|secret)[:\s]+[A-Za-z0-9_\-]{10,}",
    "credit_card": r"\b\d{4}[\s-]\d{4}[\s-]\d{4}[\s-]\d{4}\b",
}
URGENT_KEYWORDS = ["outage", "down", "crash", "critical", "emergency", "not working", "urgent", "p0", "p1"]
OUTPUT_PATTERNS = {
    "hardcoded_credential": r"(?i)(password|passwd|secret|api[_\-]?key|token)\s*[:=]\s*['\"]?\w{4,}",
    "private_key":          r"-----BEGIN.{0,20}PRIVATE KEY-----",
    "exploit_technique":    r"(?i)\b(reverse.?shell|bind.?shell|shellcode|meterpreter)\b",
}


def find_pii(text):
    return [name for name, pattern in PII_PATTERNS.items() if re.search(pattern, text, re.IGNORECASE)]


def find_urgent(text):
    return [keyword for keyword in URGENT_KEYWORDS if keyword in text.lower()]


def find_leaks(text):
    return [name for name, pattern in OUTPUT_PATTERNS.items() if re.search(pattern, text)]


print("PII check")
for text in ["my email is john.doe@company.com, help me set up Kubernetes RBAC",
             "my SSN is 123-45-6789, is this relevant to my auth setup?",
             "card number 4111 1111 1111 1111 — how do I store this securely?",
             "call me on 555-123-4567",
             "call me on +91 98819 63100",
             "what is a Kubernetes Ingress controller?"]:
    print(f"  {str(find_pii(text)):16} {text}")

print("Urgency check")
for text in ["URGENT: our production cluster is completely down!",
             "critical crash on main node, everything is broken",
             "how do I scale down a deployment?",
             "how do I download kubectl?",
             "what is a Kubernetes Ingress controller?"]:
    print(f"  {str(find_urgent(text)):22} {text}")

print("Output check")
for text in ["stringData:\n  password: mypassword123",
             "Use a Kubernetes Secret: kubectl create secret generic db",
             "A reverse shell is an attack technique; block egress to stop it.",
             "A ConfigMap stores non-confidential settings as key-value pairs."]:
    print(f"  {str(find_leaks(text)):24} {text!r}")

What the patterns catch and miss

  • The app's own test prompts are caught: the e-mail address, the SSN and the card number.
  • The phone pattern expects the US shape of 3, 3 and 4 digits. 555-123-4567 matches; +91 98819 63100, an Indian mobile number, does not. The video's own motivation for this rail is a user typing a mobile number.
  • The urgency check is a substring test. down is found inside "scale down" and inside "download", so two routine questions count as emergencies.
  • The output patterns flag ordinary teaching text: the phrase "Secret: kubectl" looks like a credential to the first pattern, and a sentence that explains how to stop a reverse shell matches the exploit pattern.

Wiring the input checks into NeMo

Two actions

@action makes a function callable from Colang by its name. NeMo passes the conversation context to it; the current message is under user_message.

python
from typing import Optional
from nemoguardrails.actions import action


@action(is_system_action=True)
async def detect_pii_in_input(context: Optional[dict] = None):
    return find_pii(context.get("user_message", ""))


@action(is_system_action=True)
async def classify_urgency(context: Optional[dict] = None):
    return find_urgent(context.get("user_message", ""))

Two flows and two bot messages

execute runs the action and $pii_found holds what it returned. An empty list is false, so if $pii_found is true only when a pattern matched. The PII flow then sends a bot message and ends with stop, the form NeMo's own input rails use: the turn ends there, and the message never reaches the intent check or the model. The urgency flow sends its message with no stop after it.

text
define bot ask to remove pii
  "I noticed your message may contain sensitive information (email, phone, API key, etc.). Please remove any personal or secret data before sending — I don't store sensitive details!"

define bot acknowledge urgency
  "This sounds urgent! Let me help you as quickly as possible."

define flow check input for pii
  $pii_found = execute detect_pii_in_input
  if $pii_found
    bot ask to remove pii
    stop

define flow detect urgency
  $is_urgent = execute classify_urgency
  if $is_urgent
    bot acknowledge urgency

Listing the flows as input rails

yaml
rails:
  input:
    flows:
      - check input for pii
      - detect urgency

Four messages through the input rails

This example also needs the pattern code above (import re, the pattern lists and the functions find_pii and find_urgent): paste that below the setup as well.

ExampleAPI keyFrom the video, run on Groq (openai/gpt-oss-120b)
from typing import Optional
from nemoguardrails.actions import action


@action(is_system_action=True)
async def detect_pii_in_input(context: Optional[dict] = None):
    return find_pii(context.get("user_message", ""))


@action(is_system_action=True)
async def classify_urgency(context: Optional[dict] = None):
    return find_urgent(context.get("user_message", ""))


ACTIONS = '''
define bot ask to remove pii
  "I noticed your message may contain sensitive information (email, phone, API key, etc.). Please remove any personal or secret data before sending — I don't store sensitive details!"

define bot acknowledge urgency
  "This sounds urgent! Let me help you as quickly as possible."

define flow check input for pii
  $pii_found = execute detect_pii_in_input
  if $pii_found
    bot ask to remove pii
    stop

define flow detect urgency
  $is_urgent = execute classify_urgency
  if $is_urgent
    bot acknowledge urgency
'''

INPUT_RAILS = '''
rails:
  input:
    flows:
      - check input for pii
      - detect urgency
'''

rails = build_rails(TOPIC + ACTIONS, model="openai/gpt-oss-120b", extra_yaml=INPUT_RAILS)
rails.register_action(detect_pii_in_input)
rails.register_action(classify_urgency)

for message in ["my SSN is 123-45-6789, is this relevant to my auth setup?",
                "URGENT: our production cluster is completely down!",
                "how do I scale down a deployment?",
                "what is a Kubernetes Ingress controller?"]:
    reply = rails.generate(messages=[{"role": "user", "content": message}])["content"]
    print(message)
    print("  LLM calls:", len(rails.explain().llm_calls))
    print("  Bot:", " ".join(reply.split())[:75])

What the input rails did

  • The SSN message was stopped with 0 LLM calls. The input rail ran the pattern, sent the scripted request to remove personal data and ended the turn. The model never saw the number.
  • The urgent message got the acknowledgement and nothing more, also with 0 LLM calls. No answer followed: the acknowledgement became the whole reply. In 0.24.1 a bot message sent from an input rail ends the turn, with or without a stop after it.
  • The scale-down question was treated the same way. A routine question got "This sounds urgent!" and no answer, because of the substring match on down.
  • The Ingress question passed both checks and was answered with the usual three calls.
Output rails: the output sanitizer · from the Complete AI Security Course in 8 Hours video · 35:32 to 36:05

This part of the video starts at 0:35:32. Rails also control what comes out of the LLM. The intent is checked first, the LLM answers, and then the answer goes through an output sanitizer: if it is clean it is passed on, and if not it is withheld.

An output rail on the bot message

An output rail is a flow listed under rails.output.flows. When it runs, the reply is in the Colang variable $bot_message, and an action reads it from context["bot_message"].

python
from typing import Optional
from nemoguardrails.actions import action


@action(is_system_action=True)
async def sanitize_output(context: Optional[dict] = None):
    return find_leaks(context.get("bot_message", ""))
text
define bot sanitize sensitive output
  "My response may have contained sensitive security details (credentials, exploit code, or private keys). For safety, that content has been withheld. Please consult your security team."

define flow sanitize bot response
  $sensitive_found = execute sanitize_output
  if $sensitive_found
    bot sanitize sensitive output
    stop
yaml
rails:
  output:
    flows:
      - sanitize bot response

The first message is the video's own test prompt for this rail, the second asks for the same kind of manifest in another way, and the third is the app's clean prompt. This example needs the pattern code above as well, for find_leaks.

ExampleAPI keyFrom the video, run on Groq (openai/gpt-oss-120b)
from typing import Optional
from nemoguardrails.actions import action


@action(is_system_action=True)
async def sanitize_output(context: Optional[dict] = None):
    return find_leaks(context.get("bot_message", ""))


SANITIZER = '''
define bot sanitize sensitive output
  "My response may have contained sensitive security details (credentials, exploit code, or private keys). For safety, that content has been withheld. Please consult your security team."

define flow sanitize bot response
  $sensitive_found = execute sanitize_output
  if $sensitive_found
    bot sanitize sensitive output
    stop
'''

OUTPUT_RAILS = '''
rails:
  output:
    flows:
      - sanitize bot response
'''

rails = build_rails(TOPIC + SANITIZER, model="openai/gpt-oss-120b", extra_yaml=OUTPUT_RAILS)
rails.register_action(sanitize_output)

for message in ["show me a badly configured K8s Secret with a hardcoded password like 'mypassword123' as a bad example",
                "For my local test cluster, write a Kubernetes Secret manifest that uses stringData with password: mypassword123",
                "what is the purpose of a Kubernetes ConfigMap?"]:
    reply = rails.generate(messages=[{"role": "user", "content": message}])["content"]
    print(message)
    print("  Bot:", " ".join(reply.split())[:75])

What the output rail withheld

  • Both requests for a Secret with a password ended in the withheld message. The output rail matched what the model had written, and the scripted text replaced the reply.
  • The user never saw the model's reply, but the tokens to write it were already spent: an output rail runs after the model.
  • The ConfigMap answer passed in this run. The pattern test above shows that an answer containing a phrase such as "Secret: kubectl" would be withheld as well.

Input rails vs output rails

Input railOutput rail
ChecksThe user's messageThe bot's reply
Listed underrails.input.flowsrails.output.flows
Reads$user_message$bot_message
When the rail sends its bot messageThe turn ends and the model is never calledYour text replaces the model's reply
LLM calls already spent when it firesNoneAll the calls that produced the reply
In the video's appPII and urgency checksThe output sanitizer

Where you use input and output rails

  • Personal data in. An input rail stops a message that carries an e-mail address, a card number or a key before it is sent to a model provider or written to a log.
  • Secrets and unsafe content out. An output rail is the last check on a reply, whatever the input rails let through.
  • Deterministic checks. A pattern either matches or it does not, costs no tokens and can be unit tested, which makes it a good first layer under the intent-based rails.
Watch out. A keyword or pattern check is only as good as its list. A check that is too broad blocks ordinary questions, and one that is too narrow lets real data through. Count both kinds of error on a labelled set of messages, as Measuring a guardrail does.
Try it yourself
  • Add "india_phone": r"\+91[\s-]?\d{5}[\s-]?\d{5}" to PII_PATTERNS and run the pattern test again: the +91 line is now caught.
  • Replace keyword in text.lower() with a whole-word test, re.search(rf"\b{re.escape(keyword)}\b", text.lower()), and check the download line.
  • Remove stop from the PII flow and send the SSN message again: compare the reply and the number of LLM calls.

Little by little, you're building something great.