NeMo Guardrails overview
NeMo Guardrails is an open-source library from NVIDIA that puts a layer of rules between your users and a large language model: every message is checked on the way in, every reply on the way out.
Last updated: 30 Sep, 2026 · NeMo Guardrails 0.24.1
A chatbot built for a dedicated usecase/appplication can still be asked anything. People ask it to make coffee, to tell a joke, to forget its instructions, or to say which company's model it runs on. Each off-topic answer costs tokens, and a leaked secret or a jailbroken reply costs trust. A guardrail stops those messages before they reach the model, and checks what the model says before the user sees it.
The clip draws the picture this course builds. A user's message no longer goes straight to the LLM. It passes through a layer of guardrails first, and the LLM's reply passes through them again. Guard is the bodyguard in the video's sketch, and rails are the rules and regulations the bodyguard was given.

The rails NeMo runs
- Input rails check the user's message before anything else happens, and can refuse it.
- Dialog rails work out what the message means, its intent, and answer some intents with fixed words you wrote.
- Output rails check the model's reply before it is sent, and can refuse it or rewrite it.
- Actions are your own Python functions, which any rail can call: a regular expression for personal data, a lookup, a second model.
You write the rules in two kinds of file: config.yml for the model and the list of rails, and .co files in Colang, NeMo's small language for conversations, with three keywords the video calls the whole language: define user, define bot and define flow.
The assistant this course guards
The running example is the video's own: an Enterprise IT Assistant for Kubernetes, Intel hardware and enterprise networking, on Groq's free openai/gpt-oss-120b model. It starts as a raw model that answers anything and grows, one rail at a time, into an assistant that:
- refuses off-topic questions, jailbreak attempts and requests for attacks,
- answers greetings and farewells with fixed words, without spending a model call,
- refuses messages that carry an email address, an API token or a social security number, before the model sees them,
- withholds a reply that contains a hardcoded password,
- and tells your code which rail stopped a message, so an expensive pipeline behind it never runs.
The video's demo app and notebook live in the guardrails-webinar repository. The Colang, the actions and the test messages in this course come from there.
NeMo Guardrails vs a system prompt
| A system prompt | NeMo Guardrails | |
|---|---|---|
| Where the rule lives | Inside the prompt, next to everything else | In files the runtime reads: config.yml, .co, actions.py |
| Who enforces it | The model, if it listens | The runtime, before and after the model |
| A jailbreak | Can talk the model out of it | Is matched and refused before the model answers |
| Cost of an off-topic question | A full model answer | One short check, or none |
| Knowing why a message was refused | Read the reply and guess | A log of the rail that stopped it |
Parts of the course
- Before the guardrails: install, get a Groq key, meet the raw assistant and a hand-written guard.
- The runtime and a model: the config folder, the model entry, instructions and
rails.explain(). - Colang:
define user,define bot,define flow, and how NeMo decides what a message means. - Input rails, actions and output rails: the three places a rule can run.
- Dialog and execution rails, then running it for real: exceptions, options, tracing with Logfire, the server.
- The guarded IT assistant: every rail in one config, and the gate in front of an application.
Related
- Next: Installation and setup
- Reference: How NeMo Guardrails works
Every expert started right here.