LangChain (YT style)LangChain 1.4 · Python 3.12+
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
46 small wins to finish your pathNext lesson →

BaseChatModel: a model of your own

A custom chat model is a subclass of BaseChatModel that supplies a type name and a _generate method, and LangChain then treats it like any hosted model.

Last updated: 27 Sep, 2026 · LangChain 1.4

Every chat model in LangChain, from OpenAI's to Groq's, is a subclass of BaseChatModel. Why write one yourself? It shows what every chat model does under the hood, and it gives you a model that replies the same way every time, which is handy in tests. A subclass has to supply two things: a name for its type, and _generate, which turns messages into a reply. Everything else, including invoke, comes from the base class.

Subclassing BaseChatModel

python
from langchain.chat_models import BaseChatModel

class MyModel(BaseChatModel):
    @property
    def _llm_type(self):        # a label for logs and traces
        return "my-model"

    def _generate(self, messages, stop=None, run_manager=None, **kwargs):
        ...                     # return a ChatResult holding one AIMessage

Setting up the model class

Start with the imports and a name for the model's type.

python
import re

from langchain.chat_models import BaseChatModel
from langchain.messages import AIMessage
from langchain_core.outputs import ChatGeneration, ChatResult


class ShopModel(BaseChatModel):
    @property
    def _llm_type(self):
        return "shop"

_llm_type is a label LangChain uses in logs and traces. The imports are the base class, the message class the model returns, and two small wrappers that _generate has to put its reply in.

Writing the _generate method

Now the reply. _generate reads the messages and returns one AIMessage, wrapped in the result type every chat model returns.

python
    def _generate(self, messages, stop=None, run_manager=None, **kwargs):
        text = messages[-1].text
        orders = re.findall(r"\b[A-Z]\d+\b", text)
        if orders:
            reply = f"I have no way to look up {orders[0]} yet."
        else:
            reply = "Hello. Which order is this about?"
        message = AIMessage(reply)
        return ChatResult(generations=[ChatGeneration(message=message)])

The model reads the last message and looks for an order id, a capital letter followed by digits. With one, it admits it cannot look the order up yet. Without one, it asks which order the customer means.

This model only matches a pattern, so it cannot help a customer. That is why the course runs on a real Groq model, and uses a stand-in like this only when a lesson needs a reply a real model will not give on cue.

Running the whole model

Create the model and call it: add these lines to the end of shop_model.py.

Example
model = ShopModel()

reply = model.invoke("Where is my order A17?")
print(type(reply).__name__)
print(reply.text)

invoke came from the base class. It turned the string into a HumanMessage, called your _generate, and returned the AIMessage inside the result.

A list of messages works too

Example
model = ShopModel()

reply = model.invoke([
    {"role": "system", "content": "You help customers of a small online shop."},
    {"role": "user", "content": "Hello"},
])
print(reply.text)

A list of dictionaries works too, converted to message objects before _generate sees them. The model reads only the last one, the customer's "Hello", so it asks which order this is about.

A hosted model works the same way
A hosted model's _generate sends the messages to the provider's servers and wraps what comes back. Yours decides with a few lines of Python. The rest of LangChain cannot tell the two apart, which is why a written stand-in can take a real model's place in a test or in a lesson that needs a scripted reply.

Your model vs a hosted model

Your modelHosted model
Where the reply comes fromA few lines of PythonThe provider's servers
Needs an API keyNoYes
usage_metadataNoneFilled with token counts
Rest of LangChainTreats it the sameTreats it the same

When to write your own model

  • Scripting a reply a real model will not give on cue, such as a deliberately bad value to test error handling.
  • A deterministic stand-in in tests, so a run's output never drifts.
  • A canned model while you build and check the code around it.
Watch out. A subclass must define both _llm_type and _generate. Leave either out and creating the model raises TypeError for the missing abstract method before it ever runs.
Try it yourself
  • Invoke it with "Hi, is B22 on its way?" and check which order it names.
  • Change the reply for an order id so it includes every id it found.
  • Delete the _llm_type property and read the error when you create the model.

This is what real progress feels like.