Streaming: showing the answer as it is written
run_stream is an async context manager that gives you the answer while the model is still writing it, so a chat window can show words as they arrive instead of after the whole reply.
Last updated: 28 Sep, 2026 · Pydantic AI 2.51
A full reply can take several seconds. Streaming fills the screen as the text comes, which is what makes a chat feel quick even when the model is not.
A model that yields word by word
A streaming stand-in is an async generator: it yields the text in pieces, the way a provider sends them, and is passed as stream_function=.
import asyncio
from pydantic_ai import Agent
from pydantic_ai.models.function import FunctionModel
async def slow_reply(messages, info):
for word in ["Your ", "refund ", "was ", "sent ", "today."]:
yield word
agent = Agent(FunctionModel(stream_function=slow_reply))Reading the growing answer
stream_text() yields the answer so far, longer each time.
async def main():
async with agent.run_stream("Where is my refund?") as result:
async for text in result.stream_text(debounce_by=None):
print(repr(text))
asyncio.run(main())'Your ' 'Your refund ' 'Your refund was ' 'Your refund was sent ' 'Your refund was sent today.'
debounce_by=None yields on every piece. The default, 0.1 seconds, groups pieces that arrive close together, which is kinder to a browser redrawing a page.
Reading only the new text
delta=True yields only what was added since last time, for code that appends to what is already on screen.
async def main():
async with agent.run_stream("Where is my refund?") as result:
async for piece in result.stream_text(delta=True, debounce_by=None):
print(piece, end="|")
print()
print(result.usage)
asyncio.run(main())Your |refund |was |sent |today.| RunUsage(input_tokens=50, output_tokens=6, requests=1)
Once the stream is done, result.usage and the messages are there, the same as for a normal run. With an output_type, result.stream_output() yields partial objects as fields arrive, and validators run on each.
Whole answer vs delta
| Mode | Each yield is | Use it when |
|---|---|---|
stream_text() | The full answer so far | You replace the shown text each time |
stream_text(delta=True) | Only the new piece | You append to what is already shown |
When you reach for streaming
- A chat interface where the reader should see words appear.
- A long answer where waiting for the whole thing feels slow.
- A live view of a structured result filling in field by field.
run_stream ends the run at the first answer that matches the output type. Tool calls the model sent after it in the same response are not run. For an agent that uses tools, agent.run_stream_events() streams every step of the run instead.run_stream returns a context manager, so it needs async with and an async function. Call agent.run_sync on a stream-only stand-in and you get an error, because there is no non-streaming response to return.Related
- Previous: Saving messages to JSON and back
- Next: Tool approval: approve, deny or edit a refund
- Reference: Streaming output
- Remove
debounce_by=Noneand count the lines printed. - Add
await asyncio.sleep(0.2)after eachyieldand try the default again. - Call
agent.run_syncon this agent and read the error.
Slow is fine. Stopping is the only problem.