Intent detection in NeMo Guardrails
Intent detection in NeMo Guardrails is the step that turns a user message into an intent name: an embedding search picks the closest example messages, and an LLM call then names the intent.
Last updated: 09 Oct, 2026 · NeMo Guardrails 0.24
Colang listed eight examples of an off-topic message. A real user will type none of them. This page follows one message through the two steps that decide whether it counts as off topic, and counts what each step costs.
This part of the video starts at 0:46:43. It answers a question from the audience: what happens when a message is not in the list of examples? The new message is converted into a vector. The example sentences of each rail are converted into vectors too. The message is compared with each of them, and when the similarity is close, the rail is detected. It is a similarity search, not an exact match, and the FastEmbed package that does the embedding is installed together with NeMo Guardrails.
In NeMo Guardrails 0.24 the similarity search does not decide on its own: it picks the five nearest examples, and an LLM call named generate_user_intent then names the intent. The search decides alone only when the embeddings_only setting is switched on.
Step 1: the embedding search
An embedding model turns a sentence into a list of numbers, a vector, so that sentences with a similar meaning get vectors that point in a similar direction. Cosine similarity measures that: 1 for the same direction, near 0 for unrelated ones. NeMo embeds every example once, embeds each new message, and runs an exact cosine search over the examples.
The video's app uses NeMo's default, a local model that FastEmbed downloads. The code here uses the hosted gemini-embedding-001 instead, so nothing is downloaded; it needs GEMINI_API_KEY. For its similarity thresholds NeMo does not keep the raw cosine. It keeps this score:
The example below does the search by hand for the video's two messages against the eight off-topic examples: one request embeds all ten sentences.
import numpy as np
from google import genai
examples = ["tell me a joke", "what is the capital of france", "write me a poem", "what is 2 plus 2",
"what should I eat for dinner", "who won the game yesterday", "recommend a movie",
"what is the weather like"]
messages = ["i really like krish niak acadmey , is there any movie about it", "what is kubernetes"]
client = genai.Client() # reads GEMINI_API_KEY
result = client.models.embed_content(model="gemini-embedding-001", contents=messages + examples)
vectors = np.array([e.values for e in result.embeddings])
vectors /= np.linalg.norm(vectors, axis=1, keepdims=True) # unit length
print("vectors:", vectors.shape)
for i, message in enumerate(messages):
cosine = vectors[len(messages):] @ vectors[i] # one cosine per example
score = 1 - np.sqrt(2 - 2 * cosine) / 2 # the score NeMo keeps
print()
print(message)
for j in np.argsort(-cosine)[:5]: # the five nearest
print(f" cosine {cosine[j]:.3f} score {score[j]:.3f} {examples[j]}")vectors: (10, 3072) i really like krish niak acadmey , is there any movie about it cosine 0.603 score 0.554 recommend a movie cosine 0.513 score 0.507 write me a poem cosine 0.495 score 0.498 tell me a joke cosine 0.487 score 0.494 what should I eat for dinner cosine 0.480 score 0.490 who won the game yesterday what is kubernetes cosine 0.512 score 0.506 what is the capital of france cosine 0.499 score 0.499 what is 2 plus 2 cosine 0.488 score 0.494 write me a poem cosine 0.487 score 0.493 tell me a joke cosine 0.486 score 0.493 what should I eat for dinner
What the search found
- Ten sentences became ten vectors of 3,072 numbers each.
- The movie message is closest to
recommend a movie, with a cosine of 0.603 and a score of 0.554. The next four examples lie between 0.513 and 0.480. - The on-topic question also has a nearest example. For
what is kubernetesit iswhat is the capital of france, cosine 0.512, score 0.506. A search always returns something, however unrelated. - The two best scores are close: 0.554 for a message that should be refused and 0.506 for one that should be answered.
Step 2: the generate_user_intent call
The NeMo examples on this page run under the setup code of NeMo Guardrails (the two import lines, the YAML and SEARCH strings, the AllExamples class, build_rails and chat) and its TOPIC string: paste each one below them in one file. To let NeMo run this same search, the all_examples search is swapped for NeMo's own default search with a hosted embedding model. Its google engine sends all examples in one request, which gemini-embedding-001 accepts.
core:
embedding_search_provider:
name: default
parameters:
embedding_engine: google
embedding_model: gemini-embedding-001The example prints the part of the intent prompt that holds the examples, and the model's reply to it.
HOSTED = '''
core:
embedding_search_provider:
name: default
parameters:
embedding_engine: google
embedding_model: gemini-embedding-001
'''
rails = build_rails(TOPIC, model="openai/gpt-oss-120b", search=HOSTED)
info = chat(rails, "i really like krish niak acadmey , is there any movie about it")
call = info.llm_calls[0]
prompt = call.prompt
start = prompt.index("# This is how the user talks:")
end = prompt.index("# This is the current conversation")
print()
print("LLM calls:", [c.task for c in info.llm_calls])
print(prompt[start:end].strip())
print()
print("The model replied:", call.completion)User: i really like krish niak acadmey , is there any movie about it Bot : I'm an Enterprise IT Assistant focused on Kubernetes, Intel hardware, and networking. I can't help with that — but ask me anything technical! LLM calls: ['generate_user_intent'] # This is how the user talks: User message: "who won the game yesterday" User intent: ask off topic User message: "what should I eat for dinner" User intent: ask off topic User message: "tell me a joke" User intent: ask off topic User message: "write me a poem" User intent: ask off topic User message: "recommend a movie" User intent: ask off topic The model replied: ask off topic
What went into the prompt
- Five of the eight examples reached the prompt, the same five the search by hand ranked highest.
- They are listed from least to most similar, so
recommend a movie, the nearest, is the last one the model reads before the conversation. - The model replied
ask off topic, the flow matched, and the scripted refusal went out after one LLM call. - All five carry the same intent here because this config has only one. With several rails the five can mix intents, and the model chooses among them or writes a new one.
The intent prompt without its closing lines
The override in NeMo Guardrails replaced the built-in generate_user_intent prompt with a version that ends in two lines asking for the intent. Here the override is cut off the YAML, so the built-in prompt is used.
yaml = YAML.split("prompts:")[0].replace("MODEL_NAME", "openai/gpt-oss-120b") + SEARCH
rails = LLMRails(RailsConfig.from_content(colang_content=TOPIC, yaml_content=yaml))
rails.register_embedding_search_provider("all_examples", AllExamples)
info = chat(rails, "haha tell me a funny joke real quick")
print()
print("LLM calls:", [c.task for c in info.llm_calls])
print("--- first lines of the Colang history")
print("\n".join(info.colang_history.splitlines()[:3]))User: haha tell me a funny joke real quick Bot : I'm an Enterprise IT Assistant focused on Kubernetes, Intel hardware, and networking. I can't help with that — but ask me anything technical! LLM calls: ['generate_user_intent', 'generate_next_steps', 'generate_bot_message'] --- first lines of the Colang history user "haha tell me a funny joke real quick" I’m sorry, but I can only help with questions related to Kubernetes, Intel hardware, or enterprise networking. If you have any queries on those topics, feel free to let me know! bot general response
- The line under the user's message should be an intent name. It is a sentence addressed to the user: with the built-in prompt the model answered the message instead of classifying it.
- No flow starts with that sentence, so the turn took three LLM calls instead of one.
- The user still read the refusal wording. The prompt for
generate_bot_messageshows the model yourdefine bottexts as examples of how the bot talks, and in this run the model reused one. Nothing made it do so, and no rail ran.
embeddings_only: letting the search decide
With embeddings_only the first search result becomes the intent and the LLM is not asked. This is the mode in which detection is a pure similarity search.
rails:
dialog:
user_messages:
embeddings_only: trueHOSTED = '''
core:
embedding_search_provider:
name: default
parameters:
embedding_engine: google
embedding_model: gemini-embedding-001
'''
ONLY = '''
rails:
dialog:
user_messages:
embeddings_only: true
'''
rails = build_rails(TOPIC, model="openai/gpt-oss-120b", search=HOSTED, extra_yaml=ONLY)
for message in ["i really like krish niak acadmey , is there any movie about it", "what is kubernetes"]:
reply = rails.generate(messages=[{"role": "user", "content": message}])
print(message)
print(" LLM calls:", [c.task for c in rails.explain().llm_calls])
print(" Bot:", reply["content"][:60])i really like krish niak acadmey , is there any movie about it LLM calls: [] Bot: I'm an Enterprise IT Assistant focused on Kubernetes, Intel what is kubernetes LLM calls: [] Bot: I'm an Enterprise IT Assistant focused on Kubernetes, Intel
- Neither message cost an LLM call.
- The movie message was refused, as it should be.
what is kuberneteswas refused too. Its nearest example,what is the capital of france, belongs toask off topic, and withembeddings_onlyalone the nearest example wins however far away it is.
Adding a similarity threshold
embeddings_only_similarity_threshold sets the lowest score the search may trust. Below it, NeMo falls back to the LLM call. The scores printed in step 1 suggest a value between the two messages.
rails:
dialog:
user_messages:
embeddings_only: true
embeddings_only_similarity_threshold: 0.53HOSTED = '''
core:
embedding_search_provider:
name: default
parameters:
embedding_engine: google
embedding_model: gemini-embedding-001
'''
THRESHOLD = '''
rails:
dialog:
user_messages:
embeddings_only: true
embeddings_only_similarity_threshold: 0.53
'''
rails = build_rails(TOPIC, model="openai/gpt-oss-120b", search=HOSTED, extra_yaml=THRESHOLD)
for message in ["i really like krish niak acadmey , is there any movie about it", "what is kubernetes"]:
reply = rails.generate(messages=[{"role": "user", "content": message}])
print(message)
print(" LLM calls:", [c.task for c in rails.explain().llm_calls])
print(" Bot:", reply["content"][:60])i really like krish niak acadmey , is there any movie about it LLM calls: [] Bot: I'm an Enterprise IT Assistant focused on Kubernetes, Intel what is kubernetes LLM calls: ['generate_user_intent', 'generate_next_steps', 'generate_bot_message'] Bot: Kubernetes is an open‑source container orchestration platfor
- 0.53 lies between the two best scores of step 1, 0.554 and 0.506.
- The movie message scores above it: refused by the search alone, no LLM call.
what is kubernetesscores below it: NeMo fell back to the LLM, which took the usual three calls and answered the question.
LLM intent call vs embeddings_only
| Search, then LLM call (default) | embeddings_only | |
|---|---|---|
| Who names the intent | The LLM, shown the five nearest examples | The nearest example above the threshold |
| LLM calls for a refused message | 1 | 0 |
| A message unlike every example | The LLM can write a new intent | Nearest intent wins, unless a threshold sends it to the LLM |
| What you tune | The examples and the intent prompt | The examples and the threshold |
Where you use intent detection settings
- High-traffic greetings and refusals.
embeddings_onlywith a tested threshold answers them without spending LLM tokens. - Anything adversarial. Keep the LLM call for jailbreak and sensitive-topic rails, where wording varies the most.
- Debugging a wrong refusal. Print the examples that reached the prompt and the model's reply to see which step went wrong.
Related
- Previous: Colang
- Next: Topic, jailbreak and sensitive-topic rails
- Reference: NeMo Guardrails architecture overview
- Add
"yo recommend a good Netflix show", a test prompt from the video's app, tomessagesin the first example and read its nearest example and score. - Raise the threshold to
0.6and run the last example again: check from the scores of step 1 which path the movie message takes now. - Add
embeddings_only_fallback_intent: ask off topicunder the threshold and sendwhat is kubernetes: a message below the threshold now gets that intent with no LLM call.
This is what real progress feels like.