The desk on a hosted model
Putting the desk on a hosted model is one argument in desk.py: DeskModel() becomes the init_chat_model call you have used since the setup lesson. The tools, guards and checkpointer around it do not change.
Last updated: 27 Sep, 2026 · LangChain 1.4
Every provider is a package
As invoke, batch and stream showed, each provider lives in its own package with its own chat model class, and init_chat_model picks the class from the prefix. Install the package, set its key, and the rest of your code does not change.
Back to the shop. You saw the desk's tools run on Groq in the desk-tools lesson. Here the whole desk, guardrails and checkpointer included, runs on the Groq model: desk.py builds it on init_chat_model instead of DeskModel for this lesson.
Naming a model with init_chat_model
from langchain.chat_models import init_chat_model
model = init_chat_model("groq:openai/gpt-oss-120b", temperature=0) # uses your GROQ_API_KEYThe one-argument model swap
In desk.py, that model goes where DeskModel() was, and the import of DeskModel becomes an import of init_chat_model. The tools, the four middleware, the context and the checkpointer stay where they are, because each was written against the base class. create_agent accepts the string directly too, as create_agent("groq:openai/gpt-oss-120b", ...).
agent = create_agent(
init_chat_model("groq:openai/gpt-oss-120b", temperature=0), # was DeskModel()
system_prompt="You are the support assistant for a small online shop. Answer in one or two short sentences, using only what the tools returned. If a tool says an order is not the customer's, say exactly that. Add nothing the tools did not say.",
tools=[lookup_order, refund_order, search_policies],
context_schema=Customer,
middleware=[
no_passwords,
PIIMiddleware("credit_card", strategy="mask"),
ModelCallLimitMiddleware(run_limit=6),
HumanInTheLoopMiddleware(interrupt_on={"refund_order": True}),
],
checkpointer=InMemorySaver(),
)- written in Several tool calls at once
- written in Three tools and a model that picks
- written in Documents and splitting
- written in Embeddings and a vector store
- written in Retrieval as a tool
- written in Three tools and a model that picks
- written in The desk's guardrails
- written in The desk's guardrails
- written in The desk's guardrails
View the code here
import re
from langchain.chat_models import BaseChatModel
from langchain.messages import AIMessage, ToolMessage
from langchain_core.outputs import ChatGeneration, ChatResult
class ShopModel(BaseChatModel):
tools: list = []
@property
def _llm_type(self):
return "shop"
def bind_tools(self, tools, **kwargs):
return self.model_copy(update={"tools": tools}) # a copy holding the tools
def _generate(self, messages, stop=None, run_manager=None, **kwargs):
message = self.decide(messages) # the reply comes from decide
return ChatResult(generations=[ChatGeneration(message=message)])
def decide(self, messages):
results = [] # the tool results at the end
for m in reversed(messages):
if not isinstance(m, ToolMessage):
break
results.insert(0, m.text)
if results: # results are back: answer with them
return AIMessage(" ".join(results))
text = messages[-1].text
orders = re.findall(r"\b[A-Z]\d+\b", text)
tool = "refund_order" if "refund" in text.lower() else "lookup_order"
if orders and tool in [t.name for t in self.tools]: # one call per order id
calls = [{"name": tool, "args": {"order_id": o}, "id": f"call_{o}"}
for o in orders]
return AIMessage("", tool_calls=calls)
if orders: # that tool is not bound
return AIMessage(f"I have no way to look up {orders[0]} yet.")
return AIMessage("Hello. Which order is this about?")
from dataclasses import dataclass
from langchain.tools import ToolRuntime, tool
ORDERS = {"A17": ("ravi", "shipped on 3 March"), "C40": ("mei", "waiting for stock")}
@dataclass
class Customer:
name: str
@tool
def lookup_order(order_id: str, runtime: ToolRuntime[Customer]) -> str:
"""Look up one of the customer's orders by its id, such as A17."""
owner, status = ORDERS.get(order_id, (None, None))
if owner != runtime.context.name:
return f"{order_id} is not one of your orders."
return f"{order_id} {status}."
@tool
def refund_order(order_id: str, runtime: ToolRuntime[Customer]) -> str:
"""Refund one of the customer's orders in full. This cannot be undone."""
owner, _ = ORDERS.get(order_id, (None, None))
if owner != runtime.context.name:
return f"{order_id} is not one of your orders, so it cannot be refunded."
return f"Refunded {order_id}."
from langchain_core.documents import Document
from langchain_text_splitters import RecursiveCharacterTextSplitter
POLICIES = {
"refunds.md": "Refunds go back to the card you paid with. They take up to 5 working days to arrive."
"\n\nYou can ask for a refund within 30 days of delivery. Opened items can be refunded if they are faulty.",
"shipping.md": "Standard shipping takes 3 to 5 working days. Shipping is free on orders over 50 euros."
"\n\nExpress shipping arrives the next working day and costs 9 euros.",
"accounts.md": "To reset your password, use the reset link on the sign-in page. Support staff never ask for your password.",
}
docs = [Document(page_content=text, metadata={"source": name}) for name, text in POLICIES.items()]
splitter = RecursiveCharacterTextSplitter(chunk_size=120, chunk_overlap=0, add_start_index=True)
chunks = splitter.split_documents(docs)
import re
import zlib
from langchain_core.embeddings import Embeddings
COMMON = {"a", "an", "and", "are", "can", "do", "does", "for", "how", "i",
"if", "is", "it", "my", "of", "on", "the", "to", "what", "with", "you", "your"}
class WordEmbeddings(Embeddings):
def embed_query(self, text):
vector = [0.0] * 256
for word in re.findall(r"[a-z]+", text.lower()):
if word not in COMMON:
vector[zlib.crc32(word.rstrip("s").encode()) % 256] += 1.0
return vector
def embed_documents(self, texts):
return [self.embed_query(text) for text in texts]
from langchain.tools import tool
from langchain_core.vectorstores import InMemoryVectorStore
from policies import chunks
from word_embeddings import WordEmbeddings
store = InMemoryVectorStore(WordEmbeddings())
store.add_documents(chunks)
@tool
def search_policies(query: str) -> str:
"""Search the shop's policies on refunds, shipping and accounts.
Pass the customer's question, word for word, as the query."""
found = [doc for doc, score in store.similarity_search_with_score(query, k=2) if score >= 0.3]
if not found:
return "No policy covers this."
return "\n".join(f"[{doc.metadata['source']}] {doc.page_content}" for doc in found)
import re
from langchain.messages import AIMessage
from shop_model import ShopModel
class DeskModel(ShopModel):
def decide(self, messages):
last = messages[-1]
if last.type == "tool" and last.text == "No policy covers this.":
return AIMessage("Our policies do not cover that. A person will reply.")
if last.type == "tool" or re.findall(r"\b[A-Z]\d+\b", last.text):
return super().decide(messages)
query = {"name": "search_policies", "args": {"query": last.text}, "id": "call_p"}
return AIMessage("", tool_calls=[query])
from langchain.agents.middleware import before_agent
from langchain.messages import AIMessage
@before_agent(can_jump_to=["end"])
def no_passwords(state, runtime):
if "password" in state["messages"][-1].text.lower():
answer = AIMessage("I cannot help with passwords. Please use the reset link.")
return {"messages": [answer], "jump_to": "end"}
from langchain.agents import create_agent
from langchain.chat_models import init_chat_model
from langchain.agents.middleware import HumanInTheLoopMiddleware, ModelCallLimitMiddleware, PIIMiddleware
from langgraph.checkpoint.memory import InMemorySaver
from password_check import no_passwords
from search import search_policies
from desk_tools import Customer, lookup_order, refund_order
agent = create_agent(
init_chat_model("groq:openai/gpt-oss-120b", temperature=0), # was DeskModel()
system_prompt="You are the support assistant for a small online shop. Answer in one or two short sentences, using only what the tools returned. If a tool says an order is not the customer's, say exactly that. Add nothing the tools did not say.",
tools=[lookup_order, refund_order, search_policies],
context_schema=Customer,
middleware=[
no_passwords,
PIIMiddleware("credit_card", strategy="mask"),
ModelCallLimitMiddleware(run_limit=6),
HumanInTheLoopMiddleware(interrupt_on={"refund_order": True}),
],
checkpointer=InMemorySaver(),
)
from langgraph.types import Command
from desk import Customer, agent
def say(who, text, thread):
config = {"configurable": {"thread_id": thread}}
result = agent.invoke({"messages": [{"role": "user", "content": text}]}, config,
context=Customer(who), version="v2")
if result.interrupts:
print(f"{who}: {text}\n paused for approval: {result.interrupts[0].value['action_requests'][0]['args']}")
result = agent.invoke(Command(resume={"decisions": [{"type": "approve"}]}), config,
context=Customer(who), version="v2")
text = "(approved)"
print(f"{who}: {text}\n desk: {result.value['messages'][-1].text}")
The desk on Groq
Two messages from Ravi through say from chat.py, now with Groq deciding: a policy question, then a refund of his own order, which still pauses for approval.
from chat import say
say("ravi", "How long does a refund take?", "ravi-1")
say("ravi", "Please refund A17", "ravi-1")ravi: How long does a refund take?
desk: Refunds take up to 5 working days to arrive.
ravi: Please refund A17
paused for approval: {'order_id': 'A17'}
ravi: (approved)
desk: Refunded A17.Groq answered the policy question from refunds.md, and the refund of A17 still paused for approval before it ran, because the guards sit around the model whichever model it is. The replies are Groq's own wording rather than DeskModel's fixed sentences: the policy answer restates the passage instead of quoting it with its file name. Wording that varies is what the key-free desk tests would trip over (test_an_uncovered_question_is_refused looks for DeskModel's exact "A person will reply"), which is why they stay on DeskModel.
Checking a real model's answers
A test with exact words cannot check a model that rewords its answers. An evaluation can. The evaluation crash course starts from the questions a chatbot raises: which LLM to use (OpenAI, Google Gemini, or open-source models on Groq), where accuracy for the use case matters more than cost, and how to decide that a model's output is right for that use case. That needs a ground truth to compare against, and it sets out four steps:
- Gather data points: each input with the output it should get, the ground truth. This lesson calls them goldens.
- Use an LLM as a judge: a model, given a prompt, compares each generated output with the expected one.
- Apply evaluation metrics to those comparisons.
- Compare several LLM models on the same data points and keep the one with the best metric results.
The video runs these steps with LangSmith, which tracks every evaluation in its cloud. Here is the smallest version for the desk, without LangSmith: three goldens, the Groq desk answering each one, and a second Groq call judging the answer against the expected one.
from uuid import uuid4
from langchain.chat_models import init_chat_model
from desk import agent
from desk_tools import Customer
GOLDENS = [ # each question with the answer you expect
("How long does a refund take?", "Refunds take up to 5 working days."),
("Is shipping free?", "Shipping is free on orders over 50 euros."),
("Do you sell gift cards?", "The policies do not cover this."),
]
judge = init_chat_model("groq:openai/gpt-oss-120b", temperature=0) # a second model grades
for question, expected in GOLDENS:
config = {"configurable": {"thread_id": str(uuid4())}}
result = agent.invoke({"messages": [{"role": "user", "content": question}]}, config,
context=Customer("ravi"), version="v2")
answer = result.value["messages"][-1].text
verdict = judge.invoke(f"Expected answer: {expected}\nActual answer: {answer}\n"
"Does the actual answer say the same thing as the expected one? Reply PASS or FAIL only.")
print(verdict.text.strip(), "|", question, "->", answer)PASS | How long does a refund take? -> Refunds take up to 5 working days to arrive. PASS | Is shipping free? -> Shipping is free on orders over 50 euros. PASS | Do you sell gift cards? -> No policy covers this.
GOLDENS holds each question with the answer you expect. The desk answers, and the judge compares the two and replies PASS or FAIL. A judge reads meaning, not exact words, so a reworded answer can still pass, which an exact-match test cannot allow.
What happens with no key set
Load the model with no key in the environment and it fails before any request, naming the variable to fill in. If you set the key with export, run this in a new terminal without it, or run unset GROQ_API_KEY first.
from langchain.chat_models import init_chat_model
init_chat_model("groq:openai/gpt-oss-120b")Traceback (most recent call last):
File "main.py", line 3, in <module>
init_chat_model("groq:openai/gpt-oss-120b")
groq.GroqError: The api_key client option must be set either by passing api_key to the client or by setting the GROQ_API_KEY environment variableThe model cannot be built, and the error names the variable to fill in. Set it and the same desk runs against a hosted model.
What the string does, and what it needs
- The string names provider and model.
groq:openai/gpt-oss-120btellsinit_chat_modelwhich package to import and which model to request. - No key, no model. With
GROQ_API_KEYunset the call raises before any request, and the message names the variable to set. - The desk is unchanged. The tools, the four middleware, the context and the checkpointer were each written against the base class, so only the model line moves.
When you move to a hosted model
- Moving the desk from the stand-in to a hosted model once the tests pass.
- Switching providers by editing the string, from Groq to Google or OpenAI, with no other change.
DeskModel() back in desk.py when you finish here. The next two lessons swap the store and the saver, and keep DeskModel so their outputs and the five desk tests stay exact.Related
- Previous: Integrations: real base classes
- Next: A persistent vector store with Chroma
- Reference: Chat models
- Send the three messages from A morning at the desk to the Groq desk and compare its replies with
DeskModel's. - Change the string to
"google_genai:gemini-2.5-flash"after installinglangchain-google-genai. - Watch the call limit from the call-limits lesson with a hosted model and count the model calls.
Slow is fine. Stopping is the only problem.