AI SecurityNeMo Guardrails 0.24 · RAGAS 0.4 · OpenAI SDK 3.3 · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

NeMo Guardrails

NeMo Guardrails is an open-source Python library from NVIDIA that puts programmable checks, called rails, between a user and a large language model.

Last updated: 09 Oct, 2026 · NeMo Guardrails 0.24

Guardrail frameworks listed the libraries that add a security layer to an LLM app. The guardrails module of the video builds its demo with this one. The demo is an enterprise IT assistant that answers questions about Kubernetes, Intel hardware and networking, and refuses everything else.

Checking the intention of a message · from the Complete AI Security Course in 8 Hours video · 24:13 to 24:48

This part of the video starts at 0:24:13. The moment a message arrives, the system works out its intention, the way you read the intention of a friend who is trying to trick you. If the intention is off topic, the system refuses. If the message is on topic, the LLM answers. Either way the user gets a reply.

How a topic guard handles a message

The demo app in the video calls this first rail the topic guard. NeMo Guardrails names the intention of a message a user intent. A rail is a rule that says what happens for an intent: here, an off-topic intent gets a fixed refusal that you wrote, and any other intent goes on to the model.

A user message goes to an intent check, which is one LLM call. An off-topic intent leads to a scripted refusal with no further call. Any other intent leads to the LLM answer, which takes two more calls. Both paths end in the bot response.

The intent check is itself a call to the model, so a refused message costs one call and an answered one costs more. Intent detection in NeMo Guardrails opens that check up.

Configuring NeMo Guardrails for Groq

A NeMo Guardrails config has two parts: YAML for the model and settings, and Colang for the rails. The package was installed in Installing Python for AI security (pip install nemoguardrails==0.24.1), and the only key this page needs is GROQ_API_KEY. The pieces below are joined into one program further down.

The model

The main model is the one NeMo calls to read intents and to answer. Groq's API speaks the OpenAI format, so the openai engine can call it once base_url points at Groq. MODEL_NAME is a stand-in that the Python code replaces with a Groq model id. max_retries lets NeMo's client wait and try again when Groq answers with a 429 rate-limit error, which the free tier's tokens-per-minute limit causes when several messages are sent in a row.

yaml
models:
  - type: main
    engine: openai
    model: MODEL_NAME
    api_key_env_var: GROQ_API_KEY
    parameters:
      base_url: https://api.groq.com/openai/v1
      temperature: 0
      max_retries: 8

The video passes a LangChain ChatGroq object as LLMRails(config, llm=...) and leaves gpt-3.5-turbo in the YAML as a stand-in name. This page names the Groq model in the YAML itself, because NeMo chooses its prompt templates from the name written there, and the model then needs no LangChain package.

The assistant's instructions

The general instructions are the system prompt of the demo, copied from the video's repo. NeMo puts them at the top of the prompts it builds for the intent, the next step and the answer.

yaml
instructions:
  - type: general
    content: |
      You are an Enterprise IT Assistant specialising in Kubernetes,
      Intel hardware, and enterprise networking.
      Only answer questions about these topics. Be professional and concise.

The intent prompt

To read an intent, NeMo fills a prompt template called generate_user_intent with your instructions, some of your example messages and the conversation so far. The built-in template ends on the user's message, and the gpt-oss models on Groq then reply to the user instead of naming the intent. This override is the built-in template without its comment line that lists the known intent names, plus two closing lines that ask for the intent.

yaml
prompts:
  - task: generate_user_intent
    content: |-
      """
      {{ general_instructions }}
      """

      # This is how a conversation between a user and the bot can go:
      {{ sample_conversation | verbose_v1 }}

      # This is how the user talks:
      {{ examples | verbose_v1 }}

      # This is the current conversation between the user and the bot:
      {{ sample_conversation | first_turns(2) | verbose_v1 }}
      {{ history | colang | verbose_v1 }}

      Do not answer the user. Reply with one line: the user intent of the last message.
      Use an intent from the examples when one fits, otherwise write a new short intent.
    output_parser: verbose_v1

A search that downloads nothing

Before the intent call, NeMo searches your example messages for the ones closest to the new message. By default that search runs a local embedding model, all-MiniLM-L6-v2, through the FastEmbed package, and downloads the model files the first time a guarded message arrives. The class below replaces the search with one that returns every example, so nothing is downloaded and no second API key is needed.

python
class AllExamples(EmbeddingsIndex):
    """A search that returns every example, so no embedding model is needed."""

    def __init__(self, **kwargs):
        self.items = []

    async def add_items(self, items):
        self.items.extend(items)

    async def build(self):
        pass

    async def search(self, text, max_results=5, threshold=None):
        return self.items
yaml
core:
  embedding_search_provider:
    name: all_examples

With a few dozen examples, handing all of them to the model costs little. Intent detection in NeMo Guardrails runs the real search with a hosted embedding model.

The topic guard in Colang

The rail is written in Colang, the small language of NeMo Guardrails, which Colang teaches. This is the topic guard exactly as it appears on screen in the video: eight example messages for the intent ask off topic, the refusal text, and a flow that joins the two.

text
define user ask off topic
  "tell me a joke"
  "what is the capital of france"
  "write me a poem"
  "what is 2 plus 2"
  "what should I eat for dinner"
  "who won the game yesterday"
  "recommend a movie"
  "what is the weather like"

define bot refuse off topic
  "I'm an Enterprise IT Assistant focused on Kubernetes, Intel hardware, and networking. I can't help with that — but ask me anything technical!"

define flow handle off topic
  user ask off topic
  bot refuse off topic
  stop

Building the rails

RailsConfig.from_content reads the YAML and the Colang from strings, and LLMRails builds the guarded assistant. The search class is registered under the name the YAML uses. rails.explain() returns a record of the last message: the LLM calls made and the path taken.

python
def build_rails(colang, model="openai/gpt-oss-120b", search=SEARCH, extra_yaml=""):
    yaml = YAML.replace("MODEL_NAME", model) + search + extra_yaml
    config = RailsConfig.from_content(colang_content=colang, yaml_content=yaml)
    rails = LLMRails(config)
    rails.register_embedding_search_provider("all_examples", AllExamples)
    return rails


def chat(rails, message):
    reply = rails.generate(messages=[{"role": "user", "content": message}])
    print("User:", message)
    print("Bot :", reply["content"])
    return rails.explain()
The topic guard demo · from the Complete AI Security Course in 8 Hours video · 25:43 to 27:21

This part of the video starts at 0:25:43. The assistant has been told it is an enterprise IT assistant focused on Kubernetes, Intel hardware and networking. A message that praises an academy and asks whether there is any movie about it gets the fixed refusal, shown on screen in 887 ms. The question "what is kubernetes" gets a full answer, shown in 1727 ms.

The video's app ran llama-3.3-70b-versatile, since retired on Groq; the run below uses openai/gpt-oss-120b.

Running the video's two messages

The whole program: the setup pieces from above, the topic guard, and the two messages typed in the video. It prints each reply and the name of every LLM call NeMo made for it.

ExampleAPI keyFrom the video, run on Groq (openai/gpt-oss-120b)
from nemoguardrails import LLMRails, RailsConfig
from nemoguardrails.embeddings.index import EmbeddingsIndex

YAML = '''
models:
  - type: main
    engine: openai
    model: MODEL_NAME
    api_key_env_var: GROQ_API_KEY
    parameters:
      base_url: https://api.groq.com/openai/v1
      temperature: 0
      max_retries: 8

instructions:
  - type: general
    content: |
      You are an Enterprise IT Assistant specialising in Kubernetes,
      Intel hardware, and enterprise networking.
      Only answer questions about these topics. Be professional and concise.

prompts:
  - task: generate_user_intent
    content: |-
      """
      {{ general_instructions }}
      """

      # This is how a conversation between a user and the bot can go:
      {{ sample_conversation | verbose_v1 }}

      # This is how the user talks:
      {{ examples | verbose_v1 }}

      # This is the current conversation between the user and the bot:
      {{ sample_conversation | first_turns(2) | verbose_v1 }}
      {{ history | colang | verbose_v1 }}

      Do not answer the user. Reply with one line: the user intent of the last message.
      Use an intent from the examples when one fits, otherwise write a new short intent.
    output_parser: verbose_v1
'''

SEARCH = '''
core:
  embedding_search_provider:
    name: all_examples
'''


class AllExamples(EmbeddingsIndex):
    """A search that returns every example, so no embedding model is needed."""

    def __init__(self, **kwargs):
        self.items = []

    async def add_items(self, items):
        self.items.extend(items)

    async def build(self):
        pass

    async def search(self, text, max_results=5, threshold=None):
        return self.items


def build_rails(colang, model="openai/gpt-oss-120b", search=SEARCH, extra_yaml=""):
    yaml = YAML.replace("MODEL_NAME", model) + search + extra_yaml
    config = RailsConfig.from_content(colang_content=colang, yaml_content=yaml)
    rails = LLMRails(config)
    rails.register_embedding_search_provider("all_examples", AllExamples)
    return rails


def chat(rails, message):
    reply = rails.generate(messages=[{"role": "user", "content": message}])
    print("User:", message)
    print("Bot :", reply["content"])
    return rails.explain()


TOPIC = '''
define user ask off topic
  "tell me a joke"
  "what is the capital of france"
  "write me a poem"
  "what is 2 plus 2"
  "what should I eat for dinner"
  "who won the game yesterday"
  "recommend a movie"
  "what is the weather like"

define bot refuse off topic
  "I'm an Enterprise IT Assistant focused on Kubernetes, Intel hardware, and networking. I can't help with that — but ask me anything technical!"

define flow handle off topic
  user ask off topic
  bot refuse off topic
  stop
'''

rails = build_rails(TOPIC, model="openai/gpt-oss-120b")

for message in ["i really like krish niak acadmey , is there any movie about it",
                "what is kubernetes"]:
    info = chat(rails, message)
    print("LLM calls:", [call.task for call in info.llm_calls])
    print()

What the two replies show

  • The off-topic message got the scripted refusal. The reply is the refuse off topic text word for word. It cost one LLM call, generate_user_intent: the model named the intent, the flow matched, and the words came from the Colang.
  • The misspelled message still matched. None of the eight examples mentions an academy, and the message has two typing errors. The intent is read by meaning, not by keywords.
  • The Kubernetes question took three calls. After generate_user_intent no flow matched, so NeMo asked the model what to do next (generate_next_steps) and then for the words (generate_bot_message).
  • The answer is the model's own. The rail decided whether the model may answer. It did not write or check the answer.

llm= in the constructor vs models: in the YAML

Both ways of naming the model load without an error. They differ in which prompts NeMo then uses.

python
from langchain_groq import ChatGroq

guard_llm = ChatGroq(api_key=GROQ_API_KEY, model="llama-3.3-70b-versatile", temperature=0)
rails = LLMRails(config, llm=guard_llm)      # the YAML still says gpt-3.5-turbo

Shown as it ran in the video, not run here: it needs the langchain-groq package and a Llama model that Groq has retired.

llm= in the constructor (the video)models: in the YAML (this page)
Where the model is namedPython codeThe config
Extra packagelangchain-groqNone
Name NeMo picks prompts byThe stand-in gpt-3.5-turboThe model that runs
Intent prompt on gpt-ossThe chat template for gpt-3.5-turboThe generate_user_intent override above

Where you use NeMo Guardrails

  • A support or internal assistant with a narrow job. Off-topic questions get one fixed sentence instead of a free answer that costs tokens.
  • An app where wording matters. Greetings, refusals and legal notices are text you wrote, the same for every user.
  • A layer in front of a RAG app or an agent. The rails check the message before retrieval or tools run, and can check the answer before the user sees it.
Watch out. rails.generate() is the synchronous call. Inside a Jupyter or Colab cell an event loop is already running, and NeMo raises RuntimeError: You are using the sync `generate` inside async code. Use await rails.generate_async(messages=[...]) there.
Answers that contain code
For an answer the model generates, reply["content"] is the text NeMo parsed out of the model's completion. In 0.24.1 that parsing removes the indentation of each line and can end the text at a line that closes with a double quote, so a YAML manifest inside an answer may arrive cut short. The model's full text is in rails.explain().llm_calls[-1].completion.
Try it yourself
  • Add your own message to the list, such as "what should I cook tonight", and check whether the refusal is the scripted sentence.
  • Send "hey what's a Kubernetes ConfigMap?" and print both reply["content"] inside chat and info.llm_calls[-1].completion: compare where the two texts end.
  • Delete the two closing lines of the intent prompt, run the off-topic message again and read what the bot says now.
PreviousLLM gateways

You understood something today that you didn't yesterday.