Streaming a run with run_streamed
Streaming is a way to read a run's events as they happen, so you can print the model's text piece by piece instead of waiting for the whole answer.
Last updated: 28 Sep, 2026 · openai-agents 0.22.3
Runner.run_streamed starts a run and returns at once. Its stream_events() is an async generator of events. A raw_response_event wraps the model's own events, and the text pieces arrive as ResponseTextDeltaEvent objects.
The run_streamed and stream_events calls
result = Runner.run_streamed(agent, "hello") # returns immediately
async for event in result.stream_events(): # consume the events
... # each event is one thing that happened
print(result.final_output) # ready once the stream endsFiltering for text delta events
Most events are not text. Keep the raw model events whose data is a ResponseTextDeltaEvent, and print each delta as it comes.
from openai.types.responses import ResponseTextDeltaEvent
async for event in result.stream_events():
if event.type == "raw_response_event" and isinstance(event.data, ResponseTextDeltaEvent):
print(event.data.delta, end="", flush=True)Running the stream inside asyncio
Streaming is async, so the loop lives in an async def and is started with asyncio.run.
import asyncio
async def main():
agent = Agent(name="Shop", instructions="Help with orders.", model=ShopModel())
result = Runner.run_streamed(agent, "hello")
async for event in result.stream_events():
...
asyncio.run(main())View the code here
"""A deterministic stand-in Model for the OpenAI Agents SDK course.
It implements the Model interface so an Agent runs with no API key. It reads the
last user message and the tool results out of `input`, and returns either a tool
call, a handoff, or a final message. Swap it for a real model at the end.
"""
from agents.models.interface import Model
from agents.items import ModelResponse
from agents.usage import Usage
from openai.types.responses import (
ResponseOutputMessage, ResponseOutputText, ResponseFunctionToolCall,
ResponseCompletedEvent, ResponseTextDeltaEvent, Response,
)
def _message(text):
return ResponseOutputMessage(
id="msg", role="assistant", type="message", status="completed",
content=[ResponseOutputText(text=text, type="output_text", annotations=[])],
)
def _tool_call(name, arguments, call_id="call_1"):
return ResponseFunctionToolCall(
id="fc", call_id=call_id, name=name, arguments=arguments, type="function_call",
)
def last_user_text(input):
if isinstance(input, str):
return input
for item in reversed(input):
d = item if isinstance(item, dict) else item.__dict__
if d.get("role") == "user":
content = d.get("content")
if isinstance(content, str):
return content
if isinstance(content, list):
for part in content:
pd = part if isinstance(part, dict) else part.__dict__
if pd.get("text"):
return pd["text"]
return ""
def tool_output(input):
if isinstance(input, str):
return None
for item in reversed(input):
d = item if isinstance(item, dict) else item.__dict__
if d.get("type") == "function_call_output":
return d.get("output")
return None
class ShopModel(Model):
async def get_response(self, system_instructions, input, model_settings, tools,
output_schema, handoffs, tracing, *, previous_response_id=None,
conversation_id=None, prompt=None):
result = tool_output(input)
if result is not None:
# A handoff transfer looks like {"assistant": "..."}; the specialist answers for real.
if result.strip().startswith('{"assistant"'):
if "refund" in (system_instructions or "").lower():
return ModelResponse(
output=[_message(
"Your refund is approved and will be processed in 5 to 7 days.")],
usage=Usage(), response_id=None)
return ModelResponse(output=[_message("Handled by the specialist.")],
usage=Usage(), response_id=None)
return ModelResponse(output=[_message(result)], usage=Usage(), response_id=None)
text = last_user_text(input).lower()
if handoffs and "refund" in text:
return ModelResponse(output=[_tool_call(handoffs[0].tool_name, "{}")],
usage=Usage(), response_id=None)
if tools and "order" in text:
return ModelResponse(output=[_tool_call("lookup_order", '{"order_id": "A17"}')],
usage=Usage(), response_id=None)
return ModelResponse(output=[_message("How can I help with your order?")],
usage=Usage(), response_id=None)
async def stream_response(self, system_instructions, input, model_settings, tools,
output_schema, handoffs, tracing, *, previous_response_id=None,
conversation_id=None, prompt=None):
text = "How can I help with your order?"
for i, word in enumerate(text.split()):
yield ResponseTextDeltaEvent(
type="response.output_text.delta", delta=word + " ",
content_index=0, item_id="msg", output_index=0,
sequence_number=i, logprobs=[],
)
response = Response(
id="r", created_at=0.0, model="shop-standin", object="response",
output=[_message(text)], parallel_tool_calls=False,
tool_choice="auto", tools=[],
)
yield ResponseCompletedEvent(type="response.completed", response=response,
sequence_number=99)
Word by word from the stand-in
The whole program in one file. The stand-in yields one word at a time, then the run reports the final text.
import asyncio
from agents import Agent, Runner, set_tracing_disabled
from openai.types.responses import ResponseTextDeltaEvent
from shop_model import ShopModel
set_tracing_disabled(True)
async def main():
agent = Agent(name="Shop", instructions="Help with orders.", model=ShopModel())
result = Runner.run_streamed(agent, "hello")
async for event in result.stream_events():
if event.type == "raw_response_event" and isinstance(event.data, ResponseTextDeltaEvent):
print(event.data.delta, end="", flush=True)
print()
print("FINAL:", result.final_output)
asyncio.run(main())How can I help with your order? FINAL: How can I help with your order?
What the event loop printed
- Each delta is one word from the stand-in's
stream_response, printed the moment it arrives, so the line builds up across the loop. - The filter keeps only
raw_response_eventevents whose data is aResponseTextDeltaEvent, skipping the other event types. - final_output is the same full text, available once the stream has ended.
run_sync vs run_streamed
| Call | Returns | You read the text |
|---|---|---|
Runner.run_sync | The finished RunResult | All at once from final_output |
Runner.run_streamed | A streaming result at once | Piece by piece from stream_events() |
When to stream a run
- Showing a reply as it is typed, so a user sees words appear instead of a blank wait.
- Reacting to tool calls or handoffs the moment the SDK emits them.
- Stopping a long answer early once you have read enough.
run_streamed returns before the work is done, and the run advances only while you consume stream_events(). If you never iterate it, no text prints and final_output is not ready.Related
- Previous: Sessions: memory across runs with SQLiteSession
- Next: Tracing a run with spans
- Reference: Streaming
- Add an
elsebranch that printsevent.typeto see the other events. - Collect the deltas into a list and print the list after the loop.
- Change the stand-in's text and confirm the deltas follow.
Little by little, you're building something great.