Real model: swapping in a hosted LLM
A hosted model is a provider's model that CrewAI reaches through its LLM class, chosen by the prefix of the model string, and it takes the place of the stand-in in a single argument.
Last updated: 28 Sep, 2026 · CrewAI 1.15
Every lesson so far ran on ShopLLM, the model you wrote, so no key was needed. Going to production means one change: hand each agent an LLM instead. Every agent, task, crew and flow around it stays as it is.
The one line that swaps the model
from crewai import LLM
llm = LLM(model="gpt-4o-mini") # the prefix picks the providerThe string's prefix picks the provider. A bare name like gpt-4o-mini is OpenAI; openrouter/... is OpenRouter. A hosted model writes different text each run, so read its answers as one run of many.
Putting it on the agents
from crewai import LLM
model = LLM(model="gpt-4o-mini")
clerk.llm = model
writer.llm = modelThe clerk and the writer now call the provider. The tools, the guardrail and the flow do not change, because they never depended on which model was behind the agent.
View the code here
from crewai.tools import tool
ORDERS = {"A17": "shipped on 3 March", "C40": "waiting for stock"}
@tool
def lookup_order(order_id: str) -> str:
"""Look up an order's shipping status by its id, such as A17."""
status = ORDERS.get(order_id)
return f"{order_id} {status}." if status else f"{order_id} is not an order we have."
import json
import os
import re
from crewai import BaseLLM
os.environ["OTEL_SDK_DISABLED"] = "true"
os.environ["CREWAI_DISABLE_TELEMETRY"] = "true"
os.environ["CREWAI_TRACING_ENABLED"] = "false"
os.environ["CREWAI_DISABLE_VERSION_CHECK"] = "true"
class ShopLLM(BaseLLM):
script: list = []
def supports_function_calling(self):
return True
def call(self, messages, tools=None, **kwargs):
if isinstance(messages, str):
messages = [{"role": "user", "content": messages}]
if self.script:
return self.script.pop(0)
return self.decide(messages, tools or [])
def decide(self, messages, tools):
last = messages[-1]
if last["role"] == "tool":
return last["content"]
text = last["content"]
orders = re.findall(r"\b[A-Z]\d+\b", text)
want_refund = "refund" in text.lower()
chosen = None
for t in tools:
fn = t["function"]
label = (fn["name"] + " " + (fn.get("description") or "")).lower()
is_refund = "refund" in label
if want_refund and is_refund:
chosen = fn["name"]
break
if not want_refund and not is_refund and ("look up" in label or "status" in label):
chosen = fn["name"]
break
if orders and chosen:
args = json.dumps({"order_id": orders[0]})
return [{"id": f"call_{orders[0]}", "type": "function",
"function": {"name": chosen, "arguments": args}}]
if "working with:" in text:
context = text.split("working with:")[1].strip().split("\n\n")[0]
return f"Dear customer, {context}"
if orders:
return f"I have no way to look up {orders[0]} yet."
return "Hello. Which order is this about?"
The key a provider asks for
import os
os.environ["OTEL_SDK_DISABLED"] = "true"
os.environ["CREWAI_DISABLE_TELEMETRY"] = "true"
os.environ.pop("OPENROUTER_API_KEY", None)
from crewai import LLM
try:
LLM(model="openrouter/openai/gpt-oss-120b")
except Exception as error:
print(type(error).__name__)
print(error)ImportError
Error importing native provider: 1 validation error for OpenAICompatibleCompletion
Value error, API key required for openrouter. Set OPENROUTER_API_KEY environment variable or pass api_key parameter. [type=value_error, input_value={'model': 'openai/gpt-oss...provider': 'openrouter'}, input_type=dict]
For further information visit https://errors.pydantic.dev/2.12/v/value_errorReading the refusal
- The model is refused when it is created, before any call, because the provider needs its key up front.
- The message names the variable to set,
OPENROUTER_API_KEY; set it and the same line runs. - The error type is ImportError here because CrewAI loads the provider's client lazily and it validates the key as it loads.
The stand-in against a hosted model
| ShopLLM stand-in | A hosted model | |
|---|---|---|
| Needs a key | No | Yes |
| Output | Fixed, so tests are stable | Different each run |
| Token counts | Zero | Real, in result.token_usage |
| Cost | None | Per call |
When to swap in a real model
- Going to production, where answers must be written, not scripted.
- Judging answer quality, which a fixed stand-in cannot show.
- Keep the stand-in for tests and teaching, where a stable output matters more.
Related
- Previous: MCP server tools for an agent
- Next: Integrations: from stand-in to production
- Reference: CrewAI docs, LLMs
- Set
OPENROUTER_API_KEYand run the block with a real key. - Give the writer a hosted model and keep
ShopLLMfor the clerk. - Print
result.token_usageafter a run with a hosted model.
This is what real progress feels like.