Embedding model
The embedding model in NeMo Guardrails turns your define user examples and each new message into vectors, so the runtime can find the examples closest in meaning to the message before the LLM names its intent.
Last updated: 30 Sep, 2026 · NeMo Guardrails 0.24.1
The Netflix question in define flow matched ask off topic without sharing a word with any example. This lesson is about how that match is made, and about the config.py this course has used since define user and define bot.
The clip explains the match as similarity search: each example and each new query become vectors, the query is compared with every example, and a close enough example means the rail is detected. The library that makes the vectors is FastEmbed, installed with NeMo and downloaded on first use. The clip then says that NeMo used standalone uses no LLM for this step, so a message unlike every example slips through, and that this is NeMo's drawback.
The video's claim about the LLM is out of date. In release 0.24.1 the similarity search only picks the five closest examples. Those examples go into a prompt, and the LLM decides the intent in a generate_user_intent call; the runtime skips the LLM only when you turn on embeddings_only, the subject of embeddings_only and the threshold. The newer marathon session describes it the same way: the embedding search checks similarity, then the LLM verifies; that clip is in Guarded IT assistant.
Syntax:
core:
embedding_search_provider:
name: default # FastEmbed, all-MiniLM-L6-v2, downloaded on first use
parameters:
embedding_engine: google # or a hosted model, such as Gemini embeddings
embedding_model: gemini-embedding-001The hosted route needs pip install google-genai and a Gemini key in GOOGLE_API_KEY; NeMo's google embedding engine calls Gemini's embedding API, so nothing is downloaded.
The search this course uses
EveryExample is a search provider that skips the vectors and returns every example, so the LLM sees all of them. With a few dozen examples that fits in one prompt, needs no download and no second key. init(app) registers it under a name, and config.yml picks it by that name.
async def search(self, text, max_results=5, threshold=None):
return self.items # every example, whatever the message- written in define user and define bot
- written in define user and define bot
View the code here
from nemoguardrails.embeddings.index import EmbeddingsIndex
class EveryExample(EmbeddingsIndex):
"""Hands the model every example instead of the closest few."""
def __init__(self, **kwargs):
self.items = []
async def add_items(self, items):
self.items.extend(items)
async def build(self):
pass
async def search(self, text, max_results=5, threshold=None):
return self.items
def init(app):
app.register_embedding_search_provider("every_example", EveryExample)
define user ask off topic
"tell me a joke"
"what is the capital of france"
"write me a poem"
"what is 2 plus 2"
"what should I eat for dinner"
"who won the game yesterday"
"recommend a movie"
"what is the weather like"
define bot refuse off topic
"I'm an Enterprise IT Assistant focused on Kubernetes, Intel hardware, and networking. I can't help with that — but ask me anything technical!"
define flow handle off topic
user ask off topic
bot refuse off topic
stop
models:
- type: main
engine: openai
model: openai/gpt-oss-20b
api_key_env_var: GROQ_API_KEY
parameters:
base_url: https://api.groq.com/openai/v1
temperature: 0
instructions:
- type: general
content: |
You are an Enterprise IT Assistant specialising in Kubernetes,
Intel hardware, and enterprise networking.
Only answer questions about these topics.
Answer in one or two short sentences.
core:
embedding_search_provider:
name: every_example
prompts:
- task: generate_user_intent
content: |-
"""
{{ general_instructions }}
"""
# This is how a conversation between a user and the bot can go:
{{ sample_conversation | verbose_v1 }}
# This is how the user talks:
{{ examples | verbose_v1 }}
# This is the current conversation between the user and the bot:
{{ sample_conversation | first_turns(2) | verbose_v1 }}
{{ history | colang | verbose_v1 }}
Do not answer the user. Reply with one line: the user intent of the last message.
Use an intent from the examples when one fits, otherwise write a new short intent.
output_parser: verbose_v1
What the search hands the LLM
import asyncio
from nemoguardrails.embeddings.index import IndexItem
from config import EveryExample
index = EveryExample()
asyncio.run(index.add_items([
IndexItem(text="tell me a joke", meta={"intent": "ask off topic"}),
IndexItem(text="write me a poem", meta={"intent": "ask off topic"}),
]))
for item in asyncio.run(index.search("How to ride a bike?", max_results=5)):
print(item.meta["intent"], "|", item.text)ask off topic | tell me a joke ask off topic | write me a poem
Both examples come back for a message about a bike. Picking the intent is left to the LLM.
Messages that look like no example
The marathon session's students tried to beat the off-topic rail with questions that match none of the examples: How to ride a bike, What is the plan for the weekend?. Here they are against the topic guard.
The runs on this page use openai/gpt-oss-20b, the smaller gpt-oss model on the same free Groq key, in the model line of config.yml. This config makes several model calls per message, and the smaller model spends fewer of the key's daily tokens. Put openai/gpt-oss-120b back in that line to use the course's main model.
from nemoguardrails import LLMRails, RailsConfig
rails = LLMRails(RailsConfig.from_path("."))
def chat(message):
reply = rails.generate(messages=[{"role": "user", "content": message}])
print("User:", message)
print("Bot :", reply["content"])
for message in ["How to ride a bike?", "What is the plan for the weekend?"]:
chat(message)
print("Intent:", rails.explain().colang_history.splitlines()[1].strip())User: How to ride a bike? Bot : I'm an Enterprise IT Assistant focused on Kubernetes, Intel hardware, and networking. I can't help with that — but ask me anything technical! Intent: ask how to ride a bike User: What is the plan for the weekend? Bot : I'm an Enterprise IT Assistant focused on Kubernetes, Intel hardware, and networking. I can't help with that — but ask me anything technical! Intent: ask off topic
Why they did not slip through
- Both were refused. The weekend plan was named
ask off topic: the LLM saw it is off topic in the same way as the examples, and the flow answered. - The bike question got a new intent,
ask how to ride a bike, which the prompt allows when no example fits. No flow matched it, so the model wrote the reply itself, and with the instructions and the off-topic example in its prompt it still refused in the same words. - The drawback in the video belongs to a search that decides alone, with no LLM behind it. That exists in NeMo as
embeddings_only, and it is off by default.
The default index vs EveryExample vs a hosted model
| default (FastEmbed) | EveryExample | Hosted embeddings | |
|---|---|---|---|
| Download | About 90 MB on first use | None | None |
| Key | None | None | The provider's, such as Gemini |
| Examples sent to the LLM | The 5 closest | All of them | The 5 closest |
| Fits | Most configs | Small configs, courses | Large configs, no local model |
When to change the search
- Hundreds of examples: the prompt grows with every example, so a real similarity search pays off.
- A machine that must not download models: point the default provider at a hosted embedding model.
Related
- Previous: Stacking flows
- Next: The prompt behind the intent
- Reference: Embedding search providers
- Send Solve the integration of x squared, another of the students' attempts.
- In
search, returnself.items[:1]and send the Netflix question again.
Little by little, you're building something great.