Model: a stand-in you can run without a key
A Model is the object the SDK asks for a reply, and a stand-in Model is one you write yourself that returns fixed replies instead of calling a service, so an agent runs with no API key.
Last updated: 28 Sep, 2026 · openai-agents 0.22.3
The SDK reaches a model through the Model interface: it calls get_response with the conversation and the tools, and expects a ModelResponse back. By subclassing Model and writing those methods yourself, you decide the reply with plain Python. The result is ShopModel, used by every following lesson.
The Model interface
You implement two async methods. get_response returns one whole reply; stream_response yields it in pieces (used later, in the streaming lesson).
from agents.models.interface import Model
class ShopModel(Model):
async def get_response(self, ...): # returns a ModelResponse
...
async def stream_response(self, ...): # yields events (used later)
...The imports
Model is the base class. ModelResponse and Usage are what you return. The Response... types describe the message and tool-call shapes the SDK understands. The streaming method also imports ResponseTextDeltaEvent, ResponseCompletedEvent and Response.
from agents.models.interface import Model # the base class
from agents.items import ModelResponse # what get_response returns
from agents.usage import Usage # a token count (zeros here)
from openai.types.responses import ( # reply shapes
ResponseOutputMessage, ResponseOutputText,
ResponseFunctionToolCall,
)Reading the last user message
Helper functions do the reading. last_user_text finds the most recent user message, whether the input is a plain string (the first turn) or a list of items (later turns).
def last_user_text(input):
if isinstance(input, str): # first turn: a plain string
return input
for item in reversed(input): # later: newest item first
d = item if isinstance(item, dict) else item.__dict__
if d.get("role") == "user": # find the user's message
content = d.get("content")
if isinstance(content, str):
return content
return "" # full file also reads list contentFinding a tool result
tool_output scans the input for the most recent tool result. When one is present, it means a tool ran and the model now needs to answer using its output.
def tool_output(input):
if isinstance(input, str): # no tools have run yet
return None
for item in reversed(input): # scan newest first
d = item if isinstance(item, dict) else item.__dict__
if d.get("type") == "function_call_output":
return d.get("output") # the string the tool returned
return NoneChoosing the reply in get_response
The method decides in order: if a tool result is present, answer from it; if the message asks for a refund and a specialist is listed, hand the run over; if it mentions an order and a tool exists, ask for that tool; otherwise greet. Each branch returns a ModelResponse.
async def get_response(self, system_instructions, input, model_settings,
tools, output_schema, handoffs, tracing, **kwargs):
result = tool_output(input) # a tool result came back?
if result is not None: # a tool or a handoff has run
... # compose the reply (next piece)
text = last_user_text(input).lower()
if handoffs and "refund" in text: # hand the run to the specialist
return ModelResponse(output=[_tool_call(handoffs[0].tool_name, "{}")],
usage=Usage(), response_id=None)
if tools and "order" in text: # ask the lookup tool
return ModelResponse(output=[_tool_call("lookup_order",
'{"order_id": "A17"}')], usage=Usage(), response_id=None)
return ModelResponse(output=[_message("How can I help with your order?")],
usage=Usage(), response_id=None)Composing a reply after a handoff
When the tool result is a handoff transfer record, which looks like {"assistant": ...}, the specialist that received the run composes its own reply. A plain tool result is passed straight back instead.
# inside get_response, where result is not None:
if result.strip().startswith('{"assistant"'): # a handoff transfer record
if "refund" in (system_instructions or "").lower():
return ModelResponse(output=[_message( # the specialist answers
"Your refund is approved and will be processed in 5 to 7 days.")],
usage=Usage(), response_id=None)
return ModelResponse(output=[_message("Handled by the specialist.")],
usage=Usage(), response_id=None)
return ModelResponse(output=[_message(result)], # a plain tool result
usage=Usage(), response_id=None)Streaming words one at a time
stream_response yields one delta per word, then a completed event. The streaming lesson later reads these events; the runs up to then use get_response.
async def stream_response(self, system_instructions, input, model_settings,
tools, output_schema, handoffs, tracing, **kwargs):
text = "How can I help with your order?"
for i, word in enumerate(text.split()): # one delta per word
yield ResponseTextDeltaEvent(
type="response.output_text.delta", delta=word + " ",
content_index=0, item_id="msg", output_index=0,
sequence_number=i, logprobs=[])
# a completed event then ends the stream (see the streaming lesson)Running the first agent on the stand-in
Save the class as shop_model.py; the following lessons import ShopModel from it. The last three lines are a quick test that runs one message through an agent, and are not part of the saved file.
from agents import Agent, Runner, set_tracing_disabled
from agents.models.interface import Model
from agents.items import ModelResponse
from agents.usage import Usage
from openai.types.responses import (
ResponseOutputMessage, ResponseOutputText, ResponseFunctionToolCall,
ResponseCompletedEvent, ResponseTextDeltaEvent, Response,
)
set_tracing_disabled(True)
def _message(text):
return ResponseOutputMessage(
id="msg", role="assistant", type="message", status="completed",
content=[ResponseOutputText(text=text, type="output_text", annotations=[])],
)
def _tool_call(name, arguments, call_id="call_1"):
return ResponseFunctionToolCall(
id="fc", call_id=call_id, name=name, arguments=arguments, type="function_call",
)
def last_user_text(input):
if isinstance(input, str):
return input
for item in reversed(input):
d = item if isinstance(item, dict) else item.__dict__
if d.get("role") == "user":
content = d.get("content")
if isinstance(content, str):
return content
if isinstance(content, list):
for part in content:
pd = part if isinstance(part, dict) else part.__dict__
if pd.get("text"):
return pd["text"]
return ""
def tool_output(input):
if isinstance(input, str):
return None
for item in reversed(input):
d = item if isinstance(item, dict) else item.__dict__
if d.get("type") == "function_call_output":
return d.get("output")
return None
class ShopModel(Model):
async def get_response(self, system_instructions, input, model_settings, tools,
output_schema, handoffs, tracing, *, previous_response_id=None,
conversation_id=None, prompt=None):
result = tool_output(input)
if result is not None:
# A handoff transfer looks like {"assistant": "..."}; the specialist answers for real.
if result.strip().startswith('{"assistant"'):
if "refund" in (system_instructions or "").lower():
return ModelResponse(output=[_message(
"Your refund is approved and will be processed in 5 to 7 days.")],
usage=Usage(), response_id=None)
return ModelResponse(output=[_message("Handled by the specialist.")],
usage=Usage(), response_id=None)
return ModelResponse(output=[_message(result)], usage=Usage(), response_id=None)
text = last_user_text(input).lower()
if handoffs and "refund" in text:
return ModelResponse(output=[_tool_call(handoffs[0].tool_name, "{}")],
usage=Usage(), response_id=None)
if tools and "order" in text:
return ModelResponse(output=[_tool_call("lookup_order", '{"order_id": "A17"}')],
usage=Usage(), response_id=None)
return ModelResponse(output=[_message("How can I help with your order?")],
usage=Usage(), response_id=None)
async def stream_response(self, system_instructions, input, model_settings, tools,
output_schema, handoffs, tracing, *, previous_response_id=None,
conversation_id=None, prompt=None):
text = "How can I help with your order?"
for i, word in enumerate(text.split()):
yield ResponseTextDeltaEvent(
type="response.output_text.delta", delta=word + " ",
content_index=0, item_id="msg", output_index=0,
sequence_number=i, logprobs=[],
)
response = Response(
id="r", created_at=0.0, model="shop-standin", object="response",
output=[_message(text)], parallel_tool_calls=False,
tool_choice="auto", tools=[],
)
yield ResponseCompletedEvent(type="response.completed", response=response,
sequence_number=99)
agent = Agent(name="Shop", instructions="Help with orders.", model=ShopModel())
result = Runner.run_sync(agent, "hello")
print(result.final_output)What the run produced
- The agent had no tools, so
get_responseskipped the order branch. - The message was "hello", which holds no tool result and no order, so the greeting branch ran.
- final_output is the greeting text, returned with no key and no network call.
A stand-in model vs a real model
| ShopModel stand-in | Real model | |
|---|---|---|
| Needs a key | No | Yes |
| Reply comes from | Your fixed rules | The model's generation |
| Same input, same output | Always | Not guaranteed |
| Good for | Learning and tests | Production answers |
Where a stand-in model helps
- Following a course or writing tests with no key and no cost.
- Reproducing a run exactly, because the reply is fixed.
- Checking your agent, tool and handoff wiring apart from the model.
get_response must be async and must return a ModelResponse. Returning a plain string, or forgetting async, makes the Runner fail before your agent ever answers.Related
- Previous: Agent: name, instructions and tools
- Next: Runner and the RunResult object
- See also: Models
- Reference: Model interface reference
- Change the greeting string and rerun; the new text is the final output.
- Send
"help me"instead of"hello"; the greeting branch still runs. - Remove
asyncfromget_responseand read the error the Runner raises.
Slow is fine. Stopping is the only problem.