Streaming
graph.stream() yields a run as it happens. stream_mode picks what you see: "updates" for each node's change, "values" for the whole state, "messages" for model tokens.
Last updated: 29 Sep, 2026 · LangGraph 1.2
invoke waits for the whole run and returns the final state. stream hands you each step as it finishes, which is how a UI shows progress and tokens as they arrive.
updates and values on the SuperBot graph
The crash course streams a one-node graph, START to SuperBot to END, compiled with a MemorySaver. graph.stream takes a stream_mode. With "updates", each chunk holds only what the node that ran returned: after SuperBot, the new AI message. With "values", each chunk holds the whole state after the step, so the conversation so far comes with it.
The run below is on Groq. The video prints each chunk whole, where the message metadata hides the difference; this version prints the node and the type of each message in the chunk. Its second call sends the video's later message, "I also like football", which the video sends on a new thread, "4", after this clip; here it goes on the same thread so the history shows:
from typing import Annotated
from typing_extensions import TypedDict
from langgraph.graph import StateGraph,START,END
from langgraph.graph.message import add_messages
from langgraph.checkpoint.memory import MemorySaver
from langchain.chat_models import init_chat_model
class State(TypedDict):
messages:Annotated[list,add_messages]
llm=init_chat_model("groq:openai/gpt-oss-120b")
memory=MemorySaver()
def superbot(state:State):
return {"messages":[llm.invoke(state['messages'])]}
graph=StateGraph(State)
graph.add_node("SuperBot",superbot)
graph.add_edge(START,"SuperBot")
graph.add_edge("SuperBot",END)
graph_builder=graph.compile(checkpointer=memory)
config = {"configurable": {"thread_id": "3"}}
for chunk in graph_builder.stream({'messages':"Hi,My name is Krish And I like cricket"},config,stream_mode="updates"):
for node, update in chunk.items():
print("updates:", node, [m.type for m in update["messages"]])
for chunk in graph_builder.stream({'messages':"I also like football"},config,stream_mode="values"):
print("values: ", [m.type for m in chunk["messages"]])updates: SuperBot ['ai'] values: ['human', 'ai', 'human'] values: ['human', 'ai', 'human', 'ai']
updates gave one chunk, SuperBot's new AI message. values gave the state twice on the same thread: once when the new human message arrived, with the first turn already in it, and once more after SuperBot added its reply. The shop's version below streams a joke node by node, then token by token.
The graph.stream API
for chunk in graph.stream(inputs, stream_mode="updates"):
print(chunk) # {node_name: what_that_node_changed}The example has one node that asks the model for a joke. Watch the run arrive two ways: node by node, then token by token. Build it in small steps.
The state
from typing_extensions import TypedDict
from langgraph.graph import StateGraph, START, END
class State(TypedDict):
topic: str # what to joke about
joke: str # the node fills this inThe state has two fields: topic is given at the start, and joke is what the node writes.
The make node
from langchain.chat_models import init_chat_model
model = init_chat_model("groq:openai/gpt-oss-120b", temperature=0) # uses your GROQ_API_KEY
def make(state):
reply = model.invoke(f"Tell one short, clean joke about {state['topic']}. One line.")
return {"joke": reply.content} # writes the joke fieldThe make node asks the model for a joke about the topic and writes the reply into joke. That is the one change the run produces.
Building the graph
builder = StateGraph(State)
builder.add_node("make", make)
builder.add_edge(START, "make") # START -> make -> END
builder.add_edge("make", END)One node again, running from START to make to END.
Streaming the run
graph = builder.compile()
for chunk in graph.stream({"topic": "cats", "joke": ""}, stream_mode="updates"): # each node's change
print(chunk)- stream hands you each step as it finishes instead of waiting for the whole run.
- In "updates" mode each chunk is {node_name: the_change}.
- There is one node, so the loop prints one chunk: the make node's whole joke at once.
Streaming the model's tokens
Switch to stream_mode="messages" and the same run yields the model's reply as it is written, one token at a time, with a small metadata dict beside each.
for token, meta in graph.stream({"topic": "cats", "joke": ""}, stream_mode="messages"):
if token.content: # skip empty pieces
print(repr(token.content), end=" ")Streaming a run end to end
Both ways in one file: the node update first, then the same joke as tokens.
from typing_extensions import TypedDict
from langgraph.graph import StateGraph, START, END
from langchain.chat_models import init_chat_model
model = init_chat_model("groq:openai/gpt-oss-120b", temperature=0) # uses your GROQ_API_KEY
class State(TypedDict):
topic: str
joke: str
def make(state):
reply = model.invoke(f"Tell one short, clean joke about {state['topic']}. One line.")
return {"joke": reply.content}
builder = StateGraph(State)
builder.add_node("make", make)
builder.add_edge(START, "make")
builder.add_edge("make", END)
graph = builder.compile()
for chunk in graph.stream({"topic": "cats", "joke": ""}, stream_mode="updates"):
print(chunk)
print("---")
for token, meta in graph.stream({"topic": "cats", "joke": ""}, stream_mode="messages"):
if token.content:
print(repr(token.content), end=" ")
print(){'make': {'joke': 'Why did the cat sit on the computer? It wanted to keep an eye on the mouse!'}}
---
'Why' ' did' ' the' ' cat' ' sit' ' on' ' the' ' computer' '?' ' It' ' wanted' ' to' ' keep' ' an' ' eye' ' on' ' the' ' mouse' '!' What each chunk showed
- In "updates" mode the chunk is {node_name: the_change}: the make node's whole joke arrives at once, when the node finishes.
- In "messages" mode each item is a (token, metadata) pair from the model call inside make, so the same joke arrives word by word. This is how a chat UI shows a reply as it is typed.
- The SuperBot run above showed "values": the whole state after each step.
Stream modes
| stream_mode | Yields |
|---|---|
"updates" | Each node's change, keyed by node name |
"values" | The full state after each step |
"messages" | Model tokens as they arrive |
When to stream a run
- Showing progress in a UI while a multi-step run works.
- Streaming a model's reply token by token instead of waiting for the whole answer.
stream_mode=["updates", "messages"], to get both; each chunk then says which mode it came from.Related
- Previous: Review and edit tool calls
- Next: Time travel
- Reference: Streaming
- Stream with
stream_mode=["updates", "messages"]and read which mode each chunk came from. - Add a second node and watch two update chunks arrive.
Slow is fine. Stopping is the only problem.