invoke, batch and stream
A chat model can be called three ways: invoke returns one reply, batch runs many independent requests at once, and stream returns a reply in pieces.
Last updated: 27 Sep, 2026 · LangChain 1.4
One model, three providers
Every company ships its own model API. If your code talks to one of them directly, switching providers means rewriting it. init_chat_model loads a chat model from any provider by name. The video loads the OpenAI, Google and Groq keys from .env, passes "gpt-4.1" to init_chat_model and gets a ChatOpenAI model back. model.invoke("Hello How are you?") sends that text as a human message and returns an AI message, and .content holds its words. For Gemini only the string changes: google_genai: followed by the model name.
Here is that first call. The video uses OpenAI's gpt-4.1; here the same example runs on Groq:
from langchain.chat_models import init_chat_model
model = init_chat_model("groq:openai/gpt-oss-120b")
response = model.invoke("Hello How are you?")
print(response.content)Hello! I'm doing great, thank you for asking. How can I assist you today?
invoke returns an AIMessage. The video prints .content; .text gives the same words as plain text, and the course uses .text from here on. For well-known OpenAI names such as gpt-4.1 the openai: prefix can be left off. The Gemini version reads GOOGLE_API_KEY and the Groq version reads GROQ_API_KEY; the call itself does not change.
Each provider package also has its own class: ChatOpenAI from langchain-openai, ChatGoogleGenerativeAI from langchain-google-genai and ChatGroq from langchain-groq, each needing its package installed. A model from init_chat_model is the same class underneath: the gpt-4.1 model above was a ChatOpenAI. For Groq the string starts with groq:. Both ways give the same kind of object, so the code after that line does not change:
from langchain.chat_models import init_chat_model
from langchain_openai import ChatOpenAI
model = init_chat_model("openai:gpt-4.1") # or
model = ChatOpenAI(model="gpt-4.1")Seeing the chunks arrive
With invoke you wait until the whole reply is written. stream returns an iterator that yields chunks while the model is still writing, which makes long replies much nicer to read. The video loops over model.stream("Write me a 200 words paragraph on Artificial Intelligence") and prints chunk.text with an end delimiter and flush=True, so each piece shows the moment it arrives. Printing a | after each chunk makes the pieces visible:
from langchain.chat_models import init_chat_model
model = init_chat_model("groq:openai/gpt-oss-120b")
for chunk in model.stream("Write me a 200 words paragraph on Artificial Intelligence"):
print(chunk.text, end="|", flush=True)||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||||Artificial| intelligence|,| a| multidisciplinary| field| that| merges| computer| science|,| mathematics|,| neuroscience|,| and| philosophy|,| seeks| to| create| machines| capable| of| performing| tasks| that| traditionally| required| human| cognition|.| By| leveraging| algorithms| that| learn| from| data|,| such| as| neural| networks|,| decision| trees|,| and| reinforcement| learning| models|,| AI| systems| can| recognize| patterns|,| make| predictions|,| and| adapt| to| new| information| with| remarkable| speed| and| accuracy|.| Applications| span| from| natural| language| processing|,| where| chat|bots| and| translation| services| interpret| and| generate| human|‑|like| text|,| to| computer| vision|,| enabling| autonomous| vehicles| to| perceive| their| surroundings| and| diagnose| medical| images| with| expert|‑|level| precision|.| The| rapid| evolution| of| deep| learning| has| amplified| these| capabilities|,| allowing| models| to| process| massive| datasets| and| uncover| subtle| relationships| that| el|ude| conventional| statistical| methods|.| Yet|,| alongside| these| breakthroughs| arise| ethical| considerations| concerning| bias|,| privacy|,| and| the| displacement| of| labor|,| prompting| scholars| and| policymakers| to| devise| frameworks| that| ensure| transparency|,| accountability|,| and| equitable| benefit| distribution|.| As| research| progresses| toward| artificial| general| intelligence|—|systems| that| exhibit| flexible|,| human|‑|level| reasoning|—|soc|iety| must| balance| innovation| with| responsible| stewardship|,| fostering| collaborations| that| harness| AI|’s| transformative| potential| while| safeguarding| fundamental| human| values|.| In| education|,| AI| personal|izes| learning| paths|,| while| in| climate| science| it| optim|izes| resource| management|,| illustrating| its| pervasive| influence| across| sectors| and| its| promise| for| future| societal| advancement|.|||
Each | is where one chunk ended. Groq sends almost one token per chunk, so this 200-word answer arrived in over a thousand pieces; other providers send larger ones. The || at the end are empty chunks, which is why the shop loop further down skips pieces with no text. For a chatbot, stream is the call you use most, because the user starts reading before the whole answer exists.
batch is the other case: a collection of independent requests, processed in parallel, which can improve performance and reduce cost. The video sends three questions in one model.batch call and gets all three replies back at once, as a list in the order asked. The max_concurrency setting in config caps how many calls run at the same time: with 10 questions and a limit of 5, they go 5 at a time.
model.batch(
["Why do parrots have colorful feathers?",
"How do airplanes fly?",
"What is quantum computing?"],
config={"max_concurrency": 5}, # at most 5 calls at the same time
)Now the shop. invoke, batch and stream all come from BaseChatModel, so they work on any chat model: the ShopModel you wrote and a hosted model alike. With a Groq key set, the runs below call a real model.
The invoke, batch and stream methods
model.invoke("one question") # -> a single AIMessage
model.batch(["q1", "q2", "q3"]) # -> a list of replies, same order
for chunk in model.stream("a question"): # -> pieces to print as they arrive
print(chunk.text, end="")Making the model
Make a model. temperature=0 makes the model pick its most likely words, so reruns give nearly the same reply. The ShopModel you wrote in the last lesson drops in unchanged.
from langchain.chat_models import init_chat_model
model = init_chat_model("groq:openai/gpt-oss-120b", temperature=0) # uses your GROQ_API_KEYCalling invoke
invoke sends one input and returns one AIMessage.
reply = model.invoke("Where is my order A17?")
print(reply.type, reply.text)What invoke returns
from langchain.messages import SystemMessage, HumanMessage
from langchain.chat_models import init_chat_model
model = init_chat_model("groq:openai/gpt-oss-120b", temperature=0)
reply = model.invoke([
SystemMessage("You are the support assistant for a small online shop. Reply in one short sentence that states only the facts in the message. Do not add anything else."),
HumanMessage("Order A17 shipped on 3 March. A customer asks where it is. Reply to them."),
])
print(reply.type)
print(reply.text)
print(reply.usage_metadata)ai
Your order A17 was shipped on March 3.
{'input_tokens': 125, 'output_tokens': 125, 'total_tokens': 250, 'output_token_details': {'reasoning': 105}}The reply is an AIMessage. usage_metadata holds the token counts the provider reports: input and output tokens, which is how you track what a conversation costs. The reasoning entry is part of the output count: gpt-oss thinks before it writes the answer, and those thinking tokens are billed too. The ShopModel from the last lesson counts no tokens, so there it is None.
Many requests at once
from langchain.messages import SystemMessage, HumanMessage
from langchain.chat_models import init_chat_model
model = init_chat_model("groq:openai/gpt-oss-120b", temperature=0)
system = SystemMessage("You are the support assistant for a small online shop. Answer in one short sentence.")
replies = model.batch([
[system, HumanMessage("Greet the customer.")],
[system, HumanMessage("Order B22 was delivered on 5 March. A customer asks where it is. Reply to them.")],
[system, HumanMessage("Order C40 is waiting for stock. A customer asks if it is in stock. Reply to them.")],
])
for reply in replies:
print(reply.text)Hello! How can I help you today? Your order B22 was delivered on March 5. We’re still waiting for stock on order C40, so it’s not in stock yet.
The replies come back in the order asked. The calls run in parallel on your side, which saves time with a hosted model. This is separate from the batch APIs some providers sell at a discount.
A reply in pieces
from langchain.messages import SystemMessage, HumanMessage
from langchain.chat_models import init_chat_model
model = init_chat_model("groq:openai/gpt-oss-120b", temperature=0)
messages = [
SystemMessage("You are the support assistant for a small online shop. Answer in one short sentence."),
HumanMessage("Greet the customer."),
]
for chunk in model.stream(messages):
if chunk.text:
print(type(chunk).__name__, repr(chunk.text))AIMessageChunk 'Hello' AIMessageChunk '!' AIMessageChunk ' How' AIMessageChunk ' can' AIMessageChunk ' I' AIMessageChunk ' help' AIMessageChunk ' you' AIMessageChunk ' today' AIMessageChunk '?'
A hosted model streams many small AIMessageChunk pieces, which is what lets a chat window show text as it is written; you can see them arrive in the output above. Some pieces carry no text, which is why the loop checks chunk.text. A model that supplies only _generate, like ShopModel, yields the finished reply as a single piece instead. Code that loops over a stream works with both.
invoke vs batch vs stream
| Method | Input | Returns |
|---|---|---|
invoke | One input | One AIMessage |
batch | A list of inputs | A list of replies, same order |
stream | One input | Pieces you loop over |
When to use each call
invokefor a single question and answer.batchto reply to or grade many items at once.streamto show a chat reply as it is written.
batch runs the inputs in parallel, so the requests finish in any order and must not depend on each other. The replies still come back in the order you asked, but code that updates a shared list or counter while the calls run can see them out of order.Related
- Previous: BaseChatModel: a model of your own
- Next: Tools: a function the model can call
- Reference: Chat models
- Batch five questions and check the replies come back in the order you asked.
- Collect the streamed pieces' text in a list, join it into one string, and compare it with an
invokereply. - Print
reply.idfor two calls and compare them.
Every expert started right here.