LangChainLangChain 1.4 · Python 3.10+
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
46 small wins to finish your pathNext lesson →

invoke, batch and stream

A chat model can be called three ways: invoke returns one reply, batch runs many independent requests at once, and stream returns a reply in pieces.

Last updated: 27 Sep, 2026 · LangChain 1.4

Lesson 3's ShopModel supplied only _generate. All three calling methods come from BaseChatModel and work on it unchanged.

The invoke, batch and stream methods

python
model.invoke("one question")             # -> a single AIMessage
model.batch(["q1", "q2", "q3"])          # -> a list of replies, same order
for chunk in model.stream("a question"):  # -> pieces to print as they arrive
    print(chunk.text)

Making the model

Import the model you wrote in the last lesson and make one.

python
from shop_model import ShopModel

model = ShopModel()

Calling invoke

invoke sends one input and returns one AIMessage.

python
reply = model.invoke("Where is my order A17?")
print(reply.type, reply.text)

Calling batch

batch takes a list of inputs and returns the replies in the same order.

python
replies = model.batch(["Hello", "Where is B22?"])
for reply in replies:
    print(reply.text)

Calling stream

stream yields the reply in pieces you loop over.

python
for chunk in model.stream("Hello"):
    print(chunk.text)

What invoke returns

Example
from shop_model import ShopModel

reply = ShopModel().invoke("Where is my order A17?")

print(reply.type)
print(reply.text)
print(reply.usage_metadata)

The reply is an AIMessage. usage_metadata holds token counts when a provider reports them; your model counts no tokens, so it is None. A hosted model fills it in, which is how you track what a conversation costs.

Many requests at once

Example
from shop_model import ShopModel

replies = ShopModel().batch(["Hello", "Where is B22?", "Is C40 in stock?"])

for reply in replies:
    print(reply.text)

batch takes a list of inputs and returns the replies in the same order. The calls run in parallel on your side, which saves time with a hosted model. This is separate from the batch APIs some providers sell at a discount.

A reply in pieces

Example
from shop_model import ShopModel

for chunk in ShopModel().stream("Hello"):
    print(type(chunk).__name__, repr(chunk.text))

A hosted model streams many small AIMessageChunk pieces, which is what lets a chat window show text as it is written. Your model supplies only _generate, so stream falls back to calling it once and yields the finished AIMessage as a single piece. Code that loops over a stream works with both.

What the three calls returned

  • invoke returns one AIMessage; usage_metadata is None because your model counts no tokens.
  • batch returns the replies in the order you asked, running the calls in parallel on your side.
  • stream falls back to one piece here, because your model supplies only _generate; a hosted model yields many small chunks.

invoke vs batch vs stream

MethodInputReturns
invokeOne inputOne AIMessage
batchA list of inputsA list of replies, same order
streamOne inputPieces you loop over

When to use each call

  • invoke for a single question and answer.
  • batch to reply to or grade many items at once.
  • stream to show a chat reply as it is written.
Watch out. batch runs the inputs in parallel, so a shared list or counter touched inside a tool can update out of order. Keep each call's work independent, or the results race.
Try it yourself
  • Batch five questions and check the replies come back in the order you asked.
  • Create two AIMessageChunk objects, add them with +, and print the text of the result.
  • Print reply.id for two calls and compare them.

Every expert started right here.