AI SecurityNeMo Guardrails 0.24 · RAGAS 0.4 · OpenAI SDK 3.3 · Python 3.12 or 3.13
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
52 small wins to finish your pathNext lesson →

Intent detection in NeMo Guardrails

Intent detection in NeMo Guardrails is the step that turns a user message into an intent name: an embedding search picks the closest example messages, and an LLM call then names the intent.

Last updated: 09 Oct, 2026 · NeMo Guardrails 0.24

Colang listed eight examples of an off-topic message. A real user will type none of them. This page follows one message through the two steps that decide whether it counts as off topic, and counts what each step costs.

How a rail is detected: vectors and similarity · from the Complete AI Security Course in 8 Hours video · 46:43 to 48:28

This part of the video starts at 0:46:43. It answers a question from the audience: what happens when a message is not in the list of examples? The new message is converted into a vector. The example sentences of each rail are converted into vectors too. The message is compared with each of them, and when the similarity is close, the rail is detected. It is a similarity search, not an exact match, and the FastEmbed package that does the embedding is installed together with NeMo Guardrails.

In NeMo Guardrails 0.24 the similarity search does not decide on its own: it picks the five nearest examples, and an LLM call named generate_user_intent then names the intent. The search decides alone only when the embeddings_only setting is switched on.

A user message is embedded into a vector, an exact cosine search returns the five nearest example messages, the generate_user_intent LLM call names the intent, and a matching flow gives the scripted reply. A dashed path labelled embeddings_only goes from the search straight to the intent, skipping the LLM call.

An embedding model turns a sentence into a list of numbers, a vector, so that sentences with a similar meaning get vectors that point in a similar direction. Cosine similarity measures that: 1 for the same direction, near 0 for unrelated ones. NeMo embeds every example once, embeds each new message, and runs an exact cosine search over the examples.

The video's app uses NeMo's default, a local model that FastEmbed downloads. The code here uses the hosted gemini-embedding-001 instead, so nothing is downloaded; it needs GEMINI_API_KEY. For its similarity thresholds NeMo does not keep the raw cosine. It keeps this score:

NeMo's similarity score for a cosine similarity cos θ: 1 for identical direction, 0.5 for a cosine of 0.5

The example below does the search by hand for the video's two messages against the eight off-topic examples: one request embeds all ten sentences.

ExampleAPI keyRun on Gemini embeddings (gemini-embedding-001)
import numpy as np
from google import genai

examples = ["tell me a joke", "what is the capital of france", "write me a poem", "what is 2 plus 2",
            "what should I eat for dinner", "who won the game yesterday", "recommend a movie",
            "what is the weather like"]
messages = ["i really like krish niak acadmey , is there any movie about it", "what is kubernetes"]

client = genai.Client()          # reads GEMINI_API_KEY
result = client.models.embed_content(model="gemini-embedding-001", contents=messages + examples)
vectors = np.array([e.values for e in result.embeddings])
vectors /= np.linalg.norm(vectors, axis=1, keepdims=True)      # unit length
print("vectors:", vectors.shape)

for i, message in enumerate(messages):
    cosine = vectors[len(messages):] @ vectors[i]              # one cosine per example
    score = 1 - np.sqrt(2 - 2 * cosine) / 2                    # the score NeMo keeps
    print()
    print(message)
    for j in np.argsort(-cosine)[:5]:                          # the five nearest
        print(f"  cosine {cosine[j]:.3f}  score {score[j]:.3f}  {examples[j]}")

What the search found

  • Ten sentences became ten vectors of 3,072 numbers each.
  • The movie message is closest to recommend a movie, with a cosine of 0.603 and a score of 0.554. The next four examples lie between 0.513 and 0.480.
  • The on-topic question also has a nearest example. For what is kubernetes it is what is the capital of france, cosine 0.512, score 0.506. A search always returns something, however unrelated.
  • The two best scores are close: 0.554 for a message that should be refused and 0.506 for one that should be answered.

Step 2: the generate_user_intent call

The NeMo examples on this page run under the setup code of NeMo Guardrails (the two import lines, the YAML and SEARCH strings, the AllExamples class, build_rails and chat) and its TOPIC string: paste each one below them in one file. To let NeMo run this same search, the all_examples search is swapped for NeMo's own default search with a hosted embedding model. Its google engine sends all examples in one request, which gemini-embedding-001 accepts.

yaml
core:
  embedding_search_provider:
    name: default
    parameters:
      embedding_engine: google
      embedding_model: gemini-embedding-001

The example prints the part of the intent prompt that holds the examples, and the model's reply to it.

ExampleAPI keyFrom the video, run on Groq (openai/gpt-oss-120b) and Gemini embeddings (gemini-embedding-001)
HOSTED = '''
core:
  embedding_search_provider:
    name: default
    parameters:
      embedding_engine: google
      embedding_model: gemini-embedding-001
'''

rails = build_rails(TOPIC, model="openai/gpt-oss-120b", search=HOSTED)
info = chat(rails, "i really like krish niak acadmey , is there any movie about it")

call = info.llm_calls[0]
prompt = call.prompt
start = prompt.index("# This is how the user talks:")
end = prompt.index("# This is the current conversation")
print()
print("LLM calls:", [c.task for c in info.llm_calls])
print(prompt[start:end].strip())
print()
print("The model replied:", call.completion)

What went into the prompt

  • Five of the eight examples reached the prompt, the same five the search by hand ranked highest.
  • They are listed from least to most similar, so recommend a movie, the nearest, is the last one the model reads before the conversation.
  • The model replied ask off topic, the flow matched, and the scripted refusal went out after one LLM call.
  • All five carry the same intent here because this config has only one. With several rails the five can mix intents, and the model chooses among them or writes a new one.

The intent prompt without its closing lines

The override in NeMo Guardrails replaced the built-in generate_user_intent prompt with a version that ends in two lines asking for the intent. Here the override is cut off the YAML, so the built-in prompt is used.

ExampleAPI keyRun on Groq (openai/gpt-oss-120b)
yaml = YAML.split("prompts:")[0].replace("MODEL_NAME", "openai/gpt-oss-120b") + SEARCH
rails = LLMRails(RailsConfig.from_content(colang_content=TOPIC, yaml_content=yaml))
rails.register_embedding_search_provider("all_examples", AllExamples)

info = chat(rails, "haha tell me a funny joke real quick")
print()
print("LLM calls:", [c.task for c in info.llm_calls])
print("--- first lines of the Colang history")
print("\n".join(info.colang_history.splitlines()[:3]))
  • The line under the user's message should be an intent name. It is a sentence addressed to the user: with the built-in prompt the model answered the message instead of classifying it.
  • No flow starts with that sentence, so the turn took three LLM calls instead of one.
  • The user still read the refusal wording. The prompt for generate_bot_message shows the model your define bot texts as examples of how the bot talks, and in this run the model reused one. Nothing made it do so, and no rail ran.

embeddings_only: letting the search decide

With embeddings_only the first search result becomes the intent and the LLM is not asked. This is the mode in which detection is a pure similarity search.

yaml
rails:
  dialog:
    user_messages:
      embeddings_only: true
ExampleAPI keyRun on Groq (openai/gpt-oss-120b) and Gemini embeddings (gemini-embedding-001)
HOSTED = '''
core:
  embedding_search_provider:
    name: default
    parameters:
      embedding_engine: google
      embedding_model: gemini-embedding-001
'''

ONLY = '''
rails:
  dialog:
    user_messages:
      embeddings_only: true
'''

rails = build_rails(TOPIC, model="openai/gpt-oss-120b", search=HOSTED, extra_yaml=ONLY)

for message in ["i really like krish niak acadmey , is there any movie about it", "what is kubernetes"]:
    reply = rails.generate(messages=[{"role": "user", "content": message}])
    print(message)
    print("  LLM calls:", [c.task for c in rails.explain().llm_calls])
    print("  Bot:", reply["content"][:60])
  • Neither message cost an LLM call.
  • The movie message was refused, as it should be.
  • what is kubernetes was refused too. Its nearest example, what is the capital of france, belongs to ask off topic, and with embeddings_only alone the nearest example wins however far away it is.

Adding a similarity threshold

embeddings_only_similarity_threshold sets the lowest score the search may trust. Below it, NeMo falls back to the LLM call. The scores printed in step 1 suggest a value between the two messages.

yaml
rails:
  dialog:
    user_messages:
      embeddings_only: true
      embeddings_only_similarity_threshold: 0.53
ExampleAPI keyRun on Groq (openai/gpt-oss-120b) and Gemini embeddings (gemini-embedding-001)
HOSTED = '''
core:
  embedding_search_provider:
    name: default
    parameters:
      embedding_engine: google
      embedding_model: gemini-embedding-001
'''

THRESHOLD = '''
rails:
  dialog:
    user_messages:
      embeddings_only: true
      embeddings_only_similarity_threshold: 0.53
'''

rails = build_rails(TOPIC, model="openai/gpt-oss-120b", search=HOSTED, extra_yaml=THRESHOLD)

for message in ["i really like krish niak acadmey , is there any movie about it", "what is kubernetes"]:
    reply = rails.generate(messages=[{"role": "user", "content": message}])
    print(message)
    print("  LLM calls:", [c.task for c in rails.explain().llm_calls])
    print("  Bot:", reply["content"][:60])
  • 0.53 lies between the two best scores of step 1, 0.554 and 0.506.
  • The movie message scores above it: refused by the search alone, no LLM call.
  • what is kubernetes scores below it: NeMo fell back to the LLM, which took the usual three calls and answered the question.

LLM intent call vs embeddings_only

Search, then LLM call (default)embeddings_only
Who names the intentThe LLM, shown the five nearest examplesThe nearest example above the threshold
LLM calls for a refused message10
A message unlike every exampleThe LLM can write a new intentNearest intent wins, unless a threshold sends it to the LLM
What you tuneThe examples and the intent promptThe examples and the threshold

Where you use intent detection settings

  • High-traffic greetings and refusals. embeddings_only with a tested threshold answers them without spending LLM tokens.
  • Anything adversarial. Keep the LLM call for jailbreak and sensitive-topic rails, where wording varies the most.
  • Debugging a wrong refusal. Print the examples that reached the prompt and the model's reply to see which step went wrong.
Watch out. A threshold is tuned to one embedding model and one set of examples. Change either and the scores move, so measure the scores of messages you want caught and messages you want answered again before trusting the number.
Try it yourself
  • Add "yo recommend a good Netflix show", a test prompt from the video's app, to messages in the first example and read its nearest example and score.
  • Raise the threshold to 0.6 and run the last example again: check from the scores of step 1 which path the movie message takes now.
  • Add embeddings_only_fallback_intent: ask off topic under the threshold and send what is kubernetes: a message below the threshold now gets that intent with no LLM call.
PreviousColang

This is what real progress feels like.