Pydantic AIPydantic AI 2.51 · Python 3.10+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
29 small wins to finish your pathNext lesson →

Streaming: showing the answer as it is written

run_stream is an async context manager that gives you the answer while the model is still writing it, so a chat window can show words as they arrive instead of after the whole reply.

Last updated: 28 Sep, 2026 · Pydantic AI 2.51

A full reply can take several seconds. Streaming fills the screen as the text comes, which is what makes a chat feel quick even when the model is not.

A model that yields word by word

A streaming stand-in is an async generator: it yields the text in pieces, the way a provider sends them, and is passed as stream_function=.

python
import asyncio

from pydantic_ai import Agent
from pydantic_ai.models.function import FunctionModel


async def slow_reply(messages, info):
    for word in ["Your ", "refund ", "was ", "sent ", "today."]:
        yield word


agent = Agent(FunctionModel(stream_function=slow_reply))

Reading the growing answer

stream_text() yields the answer so far, longer each time.

Example
async def main():
    async with agent.run_stream("Where is my refund?") as result:
        async for text in result.stream_text(debounce_by=None):
            print(repr(text))


asyncio.run(main())

debounce_by=None yields on every piece. The default, 0.1 seconds, groups pieces that arrive close together, which is kinder to a browser redrawing a page.

Reading only the new text

delta=True yields only what was added since last time, for code that appends to what is already on screen.

Example
async def main():
    async with agent.run_stream("Where is my refund?") as result:
        async for piece in result.stream_text(delta=True, debounce_by=None):
            print(piece, end="|")
        print()
        print(result.usage)


asyncio.run(main())

Once the stream is done, result.usage and the messages are there, the same as for a normal run. With an output_type, result.stream_output() yields partial objects as fields arrive, and validators run on each.

Whole answer vs delta

ModeEach yield isUse it when
stream_text()The full answer so farYou replace the shown text each time
stream_text(delta=True)Only the new pieceYou append to what is already shown

When you reach for streaming

  • A chat interface where the reader should see words appear.
  • A long answer where waiting for the whole thing feels slow.
  • A live view of a structured result filling in field by field.
run_stream stops at the first output
run_stream ends the run at the first answer that matches the output type. Tool calls the model sent after it in the same response are not run. For an agent that uses tools, agent.run_stream_events() streams every step of the run instead.
Watch out. run_stream returns a context manager, so it needs async with and an async function. Call agent.run_sync on a stream-only stand-in and you get an error, because there is no non-streaming response to return.
Try it yourself
  • Remove debounce_by=None and count the lines printed.
  • Add await asyncio.sleep(0.2) after each yield and try the default again.
  • Call agent.run_sync on this agent and read the error.

Slow is fine. Stopping is the only problem.