invoke, batch and stream
A chat model can be called three ways: invoke returns one reply, batch runs many independent requests at once, and stream returns a reply in pieces.
Last updated: 27 Sep, 2026 · LangChain 1.4
Lesson 3's ShopModel supplied only _generate. All three calling methods come from BaseChatModel and work on it unchanged.
The invoke, batch and stream methods
model.invoke("one question") # -> a single AIMessage
model.batch(["q1", "q2", "q3"]) # -> a list of replies, same order
for chunk in model.stream("a question"): # -> pieces to print as they arrive
print(chunk.text)Making the model
Import the model you wrote in the last lesson and make one.
from shop_model import ShopModel
model = ShopModel()Calling invoke
invoke sends one input and returns one AIMessage.
reply = model.invoke("Where is my order A17?")
print(reply.type, reply.text)Calling batch
batch takes a list of inputs and returns the replies in the same order.
replies = model.batch(["Hello", "Where is B22?"])
for reply in replies:
print(reply.text)Calling stream
stream yields the reply in pieces you loop over.
for chunk in model.stream("Hello"):
print(chunk.text)What invoke returns
from shop_model import ShopModel
reply = ShopModel().invoke("Where is my order A17?")
print(reply.type)
print(reply.text)
print(reply.usage_metadata)The reply is an AIMessage. usage_metadata holds token counts when a provider reports them; your model counts no tokens, so it is None. A hosted model fills it in, which is how you track what a conversation costs.
Many requests at once
from shop_model import ShopModel
replies = ShopModel().batch(["Hello", "Where is B22?", "Is C40 in stock?"])
for reply in replies:
print(reply.text)batch takes a list of inputs and returns the replies in the same order. The calls run in parallel on your side, which saves time with a hosted model. This is separate from the batch APIs some providers sell at a discount.
A reply in pieces
from shop_model import ShopModel
for chunk in ShopModel().stream("Hello"):
print(type(chunk).__name__, repr(chunk.text))A hosted model streams many small AIMessageChunk pieces, which is what lets a chat window show text as it is written. Your model supplies only _generate, so stream falls back to calling it once and yields the finished AIMessage as a single piece. Code that loops over a stream works with both.
What the three calls returned
- invoke returns one
AIMessage;usage_metadataisNonebecause your model counts no tokens. - batch returns the replies in the order you asked, running the calls in parallel on your side.
- stream falls back to one piece here, because your model supplies only
_generate; a hosted model yields many small chunks.
invoke vs batch vs stream
| Method | Input | Returns |
|---|---|---|
invoke | One input | One AIMessage |
batch | A list of inputs | A list of replies, same order |
stream | One input | Pieces you loop over |
When to use each call
invokefor a single question and answer.batchto reply to or grade many items at once.streamto show a chat reply as it is written.
batch runs the inputs in parallel, so a shared list or counter touched inside a tool can update out of order. Keep each call's work independent, or the results race.Related
- Previous: BaseChatModel: a model of your own
- Next: Tools: a function the model can call
- Reference: Chat models
- Batch five questions and check the replies come back in the order you asked.
- Create two
AIMessageChunkobjects, add them with+, and print the text of the result. - Print
reply.idfor two calls and compare them.
Every expert started right here.