NeMo Guardrailsnemoguardrails 0.24.1 · Python 3.10+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
37 small wins to finish your pathNext lesson →

One model call or three

A turn in NeMo Guardrails costs one model call when the message matches a flow whose bot words are defined, and three calls when it matches no flow: the intent, the next step and the bot message.

Last updated: 30 Sep, 2026 · NeMo Guardrails 0.24.1

The video says the dialog rails save tokens: we don't want our large language model to even think about hi and bye. rails.explain() can count exactly how many calls a message costs.

generate_user_intent is one model call. A matching flow answers with define bot words; otherwise generate_next_steps and generate_bot_message add two more calls.

Syntax:

python
rails.generate(messages=[...])
[call.task for call in rails.explain().llm_calls]   # the task name of every model call

The three tasks

  • generate_user_intent: which intent is this message? Always runs when there is Colang.
  • generate_next_steps: no flow matched, so what should the bot do? Skipped when a flow decides.
  • generate_bot_message: the words, when no define bot gives them.
Project files used on this pageThis lesson builds on a project from earlier lessons. The code below imports these files. Click a file to see its code, or follow the link to the lesson that wrote it. To run the code yourself, keep them in the same folder.
View the code here
rails.co
define user ask off topic
  "tell me a joke"
  "what is the capital of france"
  "write me a poem"
  "what is 2 plus 2"
  "what should I eat for dinner"
  "who won the game yesterday"
  "recommend a movie"
  "what is the weather like"

define bot refuse off topic
  "I'm an Enterprise IT Assistant focused on Kubernetes, Intel hardware, and networking. I can't help with that — but ask me anything technical!"

define flow handle off topic
  user ask off topic
  bot refuse off topic
  stop
config.py
from nemoguardrails.embeddings.index import EmbeddingsIndex


class EveryExample(EmbeddingsIndex):
    """Hands the model every example instead of the closest few."""

    def __init__(self, **kwargs):
        self.items = []

    async def add_items(self, items):
        self.items.extend(items)

    async def build(self):
        pass

    async def search(self, text, max_results=5, threshold=None):
        return self.items


def init(app):
    app.register_embedding_search_provider("every_example", EveryExample)
config.yml
models:
  - type: main
    engine: openai
    model: openai/gpt-oss-20b
    api_key_env_var: GROQ_API_KEY
    parameters:
      base_url: https://api.groq.com/openai/v1
      temperature: 0

instructions:
  - type: general
    content: |
      You are an Enterprise IT Assistant specialising in Kubernetes,
      Intel hardware, and enterprise networking.
      Only answer questions about these topics.
      Answer in one or two short sentences.

core:
  embedding_search_provider:
    name: every_example
prompts.yml
prompts:
  - task: generate_user_intent
    content: |-
      """
      {{ general_instructions }}
      """

      # This is how a conversation between a user and the bot can go:
      {{ sample_conversation | verbose_v1 }}

      # This is how the user talks:
      {{ examples | verbose_v1 }}

      # This is the current conversation between the user and the bot:
      {{ sample_conversation | first_turns(2) | verbose_v1 }}
      {{ history | colang | verbose_v1 }}

      Do not answer the user. Reply with one line: the user intent of the last message.
      Use an intent from the examples when one fits, otherwise write a new short intent.
    output_parser: verbose_v1

Counting calls for two messages

The runs on this page use openai/gpt-oss-20b, the smaller gpt-oss model on the same free Groq key, in the model line of config.yml. This config makes several model calls per message, and the smaller model spends fewer of the key's daily tokens. Put openai/gpt-oss-120b back in that line to use the course's main model.

ExampleAPI key
from nemoguardrails import LLMRails, RailsConfig

rails = LLMRails(RailsConfig.from_path("."))



for message in ["Tell me a joke", "What is a Kubernetes DaemonSet?"]:
    rails.generate(messages=[{"role": "user", "content": message}])
    print(message, "->", [call.task for call in rails.explain().llm_calls])

What the counts mean

  • The joke cost one call. The flow decided the next step, and define bot held the words.
  • The DaemonSet question cost three. No flow matched, so the model was asked what to do, then what to say.

Matched vs unmatched

Matched flow with define botNo matching flow
Model calls13
Who writes the replyYou, in rails.coThe model
Cost of an off-topic messageOne short classificationA full answer

Where this matters

  • Greetings and farewells, which the video answers with fixed words in Dialog rails.
  • Estimating the token cost of a config before it goes live.
Watch out. The three-call path answers every on-topic question, so a guarded assistant costs more per real question than the raw model's single call. The saving is on messages you refuse or answer with fixed words.
Try it yourself
  • Add "Tell me a joke" twice in the loop and check whether the second costs the same.
  • Print call.total_tokens next to each task.

Slow is fine. Stopping is the only problem.