LangChain (YT style)LangChain 1.4 · Python 3.12+
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
46 small wins to finish your pathNext lesson →

invoke, batch and stream

A chat model can be called three ways: invoke returns one reply, batch runs many independent requests at once, and stream returns a reply in pieces.

Last updated: 27 Sep, 2026 · LangChain 1.4

Loading models with init_chat_model · from the Updated LangChain Version V1 Crash Course · 37:13 to 42:23

One model, three providers

Every company ships its own model API. If your code talks to one of them directly, switching providers means rewriting it. init_chat_model loads a chat model from any provider by name. The video loads the OpenAI, Google and Groq keys from .env, passes "gpt-4.1" to init_chat_model and gets a ChatOpenAI model back. model.invoke("Hello How are you?") sends that text as a human message and returns an AI message, and .content holds its words. For Gemini only the string changes: google_genai: followed by the model name.

Here is that first call. The video uses OpenAI's gpt-4.1; here the same example runs on Groq:

ExampleAPI keyFrom the video, run on Groq
from langchain.chat_models import init_chat_model

model = init_chat_model("groq:openai/gpt-oss-120b")
response = model.invoke("Hello How are you?")
print(response.content)

invoke returns an AIMessage. The video prints .content; .text gives the same words as plain text, and the course uses .text from here on. For well-known OpenAI names such as gpt-4.1 the openai: prefix can be left off. The Gemini version reads GOOGLE_API_KEY and the Groq version reads GROQ_API_KEY; the call itself does not change.

Provider classes: ChatOpenAI, ChatGoogleGenerativeAI and ChatGroq · from the Updated LangChain Version V1 Crash Course · 42:22 to 47:29

Each provider package also has its own class: ChatOpenAI from langchain-openai, ChatGoogleGenerativeAI from langchain-google-genai and ChatGroq from langchain-groq, each needing its package installed. A model from init_chat_model is the same class underneath: the gpt-4.1 model above was a ChatOpenAI. For Groq the string starts with groq:. Both ways give the same kind of object, so the code after that line does not change:

from langchain.chat_models import init_chat_model
from langchain_openai import ChatOpenAI

model = init_chat_model("openai:gpt-4.1")          # or
model = ChatOpenAI(model="gpt-4.1")
Streaming a reply with stream · from the Updated LangChain Version V1 Crash Course · 47:47 to 52:49

Seeing the chunks arrive

With invoke you wait until the whole reply is written. stream returns an iterator that yields chunks while the model is still writing, which makes long replies much nicer to read. The video loops over model.stream("Write me a 200 words paragraph on Artificial Intelligence") and prints chunk.text with an end delimiter and flush=True, so each piece shows the moment it arrives. Printing a | after each chunk makes the pieces visible:

ExampleAPI keyFrom the video, run on Groq
from langchain.chat_models import init_chat_model

model = init_chat_model("groq:openai/gpt-oss-120b")

for chunk in model.stream("Write me a 200 words paragraph on Artificial Intelligence"):
    print(chunk.text, end="|", flush=True)

Each | is where one chunk ended. Groq sends almost one token per chunk, so this 200-word answer arrived in over a thousand pieces; other providers send larger ones. The || at the end are empty chunks, which is why the shop loop further down skips pieces with no text. For a chatbot, stream is the call you use most, because the user starts reading before the whole answer exists.

Sending questions in parallel with batch · from the Updated LangChain Version V1 Crash Course · 52:44 to 55:12

batch is the other case: a collection of independent requests, processed in parallel, which can improve performance and reduce cost. The video sends three questions in one model.batch call and gets all three replies back at once, as a list in the order asked. The max_concurrency setting in config caps how many calls run at the same time: with 10 questions and a limit of 5, they go 5 at a time.

python
model.batch(
    ["Why do parrots have colorful feathers?",
     "How do airplanes fly?",
     "What is quantum computing?"],
    config={"max_concurrency": 5},   # at most 5 calls at the same time
)

Now the shop. invoke, batch and stream all come from BaseChatModel, so they work on any chat model: the ShopModel you wrote and a hosted model alike. With a Groq key set, the runs below call a real model.

The invoke, batch and stream methods

python
model.invoke("one question")             # -> a single AIMessage
model.batch(["q1", "q2", "q3"])          # -> a list of replies, same order
for chunk in model.stream("a question"):  # -> pieces to print as they arrive
    print(chunk.text, end="")

Making the model

Make a model. temperature=0 makes the model pick its most likely words, so reruns give nearly the same reply. The ShopModel you wrote in the last lesson drops in unchanged.

python
from langchain.chat_models import init_chat_model

model = init_chat_model("groq:openai/gpt-oss-120b", temperature=0)  # uses your GROQ_API_KEY

Calling invoke

invoke sends one input and returns one AIMessage.

python
reply = model.invoke("Where is my order A17?")
print(reply.type, reply.text)

What invoke returns

ExampleAPI key
from langchain.messages import SystemMessage, HumanMessage
from langchain.chat_models import init_chat_model

model = init_chat_model("groq:openai/gpt-oss-120b", temperature=0)
reply = model.invoke([
    SystemMessage("You are the support assistant for a small online shop. Reply in one short sentence that states only the facts in the message. Do not add anything else."),
    HumanMessage("Order A17 shipped on 3 March. A customer asks where it is. Reply to them."),
])

print(reply.type)
print(reply.text)
print(reply.usage_metadata)

The reply is an AIMessage. usage_metadata holds the token counts the provider reports: input and output tokens, which is how you track what a conversation costs. The reasoning entry is part of the output count: gpt-oss thinks before it writes the answer, and those thinking tokens are billed too. The ShopModel from the last lesson counts no tokens, so there it is None.

Many requests at once

ExampleAPI key
from langchain.messages import SystemMessage, HumanMessage
from langchain.chat_models import init_chat_model

model = init_chat_model("groq:openai/gpt-oss-120b", temperature=0)
system = SystemMessage("You are the support assistant for a small online shop. Answer in one short sentence.")
replies = model.batch([
    [system, HumanMessage("Greet the customer.")],
    [system, HumanMessage("Order B22 was delivered on 5 March. A customer asks where it is. Reply to them.")],
    [system, HumanMessage("Order C40 is waiting for stock. A customer asks if it is in stock. Reply to them.")],
])

for reply in replies:
    print(reply.text)

The replies come back in the order asked. The calls run in parallel on your side, which saves time with a hosted model. This is separate from the batch APIs some providers sell at a discount.

A reply in pieces

ExampleAPI key
from langchain.messages import SystemMessage, HumanMessage
from langchain.chat_models import init_chat_model

model = init_chat_model("groq:openai/gpt-oss-120b", temperature=0)
messages = [
    SystemMessage("You are the support assistant for a small online shop. Answer in one short sentence."),
    HumanMessage("Greet the customer."),
]
for chunk in model.stream(messages):
    if chunk.text:
        print(type(chunk).__name__, repr(chunk.text))

A hosted model streams many small AIMessageChunk pieces, which is what lets a chat window show text as it is written; you can see them arrive in the output above. Some pieces carry no text, which is why the loop checks chunk.text. A model that supplies only _generate, like ShopModel, yields the finished reply as a single piece instead. Code that loops over a stream works with both.

invoke vs batch vs stream

MethodInputReturns
invokeOne inputOne AIMessage
batchA list of inputsA list of replies, same order
streamOne inputPieces you loop over

When to use each call

  • invoke for a single question and answer.
  • batch to reply to or grade many items at once.
  • stream to show a chat reply as it is written.
Watch out. batch runs the inputs in parallel, so the requests finish in any order and must not depend on each other. The replies still come back in the order you asked, but code that updates a shared list or counter while the calls run can see them out of order.
Try it yourself
  • Batch five questions and check the replies come back in the order you asked.
  • Collect the streamed pieces' text in a list, join it into one string, and compare it with an invoke reply.
  • Print reply.id for two calls and compare them.

Every expert started right here.