0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
27 small wins to finish your pathNext lesson →

Model: a stand-in you can run without a key

A Model is the object the SDK asks for a reply, and a stand-in Model is one you write yourself that returns fixed replies instead of calling a service, so an agent runs with no API key.

Last updated: 28 Sep, 2026 · openai-agents 0.22.3

The SDK reaches a model through the Model interface: it calls get_response with the conversation and the tools, and expects a ModelResponse back. By subclassing Model and writing those methods yourself, you decide the reply with plain Python. The result is ShopModel, used by every following lesson.

The Model interface

You implement two async methods. get_response returns one whole reply; stream_response yields it in pieces (used later, in the streaming lesson).

python
from agents.models.interface import Model

class ShopModel(Model):
    async def get_response(self, ...):    # returns a ModelResponse
        ...
    async def stream_response(self, ...): # yields events (used later)
        ...

The imports

Model is the base class. ModelResponse and Usage are what you return. The Response... types describe the message and tool-call shapes the SDK understands. The streaming method also imports ResponseTextDeltaEvent, ResponseCompletedEvent and Response.

python
from agents.models.interface import Model      # the base class
from agents.items import ModelResponse          # what get_response returns
from agents.usage import Usage                   # a token count (zeros here)
from openai.types.responses import (             # reply shapes
    ResponseOutputMessage, ResponseOutputText,
    ResponseFunctionToolCall,
)

Reading the last user message

Helper functions do the reading. last_user_text finds the most recent user message, whether the input is a plain string (the first turn) or a list of items (later turns).

python
def last_user_text(input):
    if isinstance(input, str):        # first turn: a plain string
        return input
    for item in reversed(input):      # later: newest item first
        d = item if isinstance(item, dict) else item.__dict__
        if d.get("role") == "user":   # find the user's message
            content = d.get("content")
            if isinstance(content, str):
                return content
    return ""                         # full file also reads list content

Finding a tool result

tool_output scans the input for the most recent tool result. When one is present, it means a tool ran and the model now needs to answer using its output.

python
def tool_output(input):
    if isinstance(input, str):        # no tools have run yet
        return None
    for item in reversed(input):      # scan newest first
        d = item if isinstance(item, dict) else item.__dict__
        if d.get("type") == "function_call_output":
            return d.get("output")    # the string the tool returned
    return None

Choosing the reply in get_response

The method decides in order: if a tool result is present, answer from it; if the message asks for a refund and a specialist is listed, hand the run over; if it mentions an order and a tool exists, ask for that tool; otherwise greet. Each branch returns a ModelResponse.

python
async def get_response(self, system_instructions, input, model_settings,
                       tools, output_schema, handoffs, tracing, **kwargs):
    result = tool_output(input)            # a tool result came back?
    if result is not None:                 # a tool or a handoff has run
        ...                                # compose the reply (next piece)
    text = last_user_text(input).lower()
    if handoffs and "refund" in text:      # hand the run to the specialist
        return ModelResponse(output=[_tool_call(handoffs[0].tool_name, "{}")],
                             usage=Usage(), response_id=None)
    if tools and "order" in text:          # ask the lookup tool
        return ModelResponse(output=[_tool_call("lookup_order",
            '{"order_id": "A17"}')], usage=Usage(), response_id=None)
    return ModelResponse(output=[_message("How can I help with your order?")],
                         usage=Usage(), response_id=None)

Composing a reply after a handoff

When the tool result is a handoff transfer record, which looks like {"assistant": ...}, the specialist that received the run composes its own reply. A plain tool result is passed straight back instead.

python
# inside get_response, where result is not None:
if result.strip().startswith('{"assistant"'):    # a handoff transfer record
    if "refund" in (system_instructions or "").lower():
        return ModelResponse(output=[_message(    # the specialist answers
            "Your refund is approved and will be processed in 5 to 7 days.")],
            usage=Usage(), response_id=None)
    return ModelResponse(output=[_message("Handled by the specialist.")],
                         usage=Usage(), response_id=None)
return ModelResponse(output=[_message(result)],   # a plain tool result
                     usage=Usage(), response_id=None)

Streaming words one at a time

stream_response yields one delta per word, then a completed event. The streaming lesson later reads these events; the runs up to then use get_response.

python
async def stream_response(self, system_instructions, input, model_settings,
                          tools, output_schema, handoffs, tracing, **kwargs):
    text = "How can I help with your order?"
    for i, word in enumerate(text.split()):   # one delta per word
        yield ResponseTextDeltaEvent(
            type="response.output_text.delta", delta=word + " ",
            content_index=0, item_id="msg", output_index=0,
            sequence_number=i, logprobs=[])
    # a completed event then ends the stream (see the streaming lesson)

Running the first agent on the stand-in

Save the class as shop_model.py; the following lessons import ShopModel from it. The last three lines are a quick test that runs one message through an agent, and are not part of the saved file.

Example
from agents import Agent, Runner, set_tracing_disabled
from agents.models.interface import Model
from agents.items import ModelResponse
from agents.usage import Usage
from openai.types.responses import (
    ResponseOutputMessage, ResponseOutputText, ResponseFunctionToolCall,
    ResponseCompletedEvent, ResponseTextDeltaEvent, Response,
)

set_tracing_disabled(True)


def _message(text):
    return ResponseOutputMessage(
        id="msg", role="assistant", type="message", status="completed",
        content=[ResponseOutputText(text=text, type="output_text", annotations=[])],
    )


def _tool_call(name, arguments, call_id="call_1"):
    return ResponseFunctionToolCall(
        id="fc", call_id=call_id, name=name, arguments=arguments, type="function_call",
    )


def last_user_text(input):
    if isinstance(input, str):
        return input
    for item in reversed(input):
        d = item if isinstance(item, dict) else item.__dict__
        if d.get("role") == "user":
            content = d.get("content")
            if isinstance(content, str):
                return content
            if isinstance(content, list):
                for part in content:
                    pd = part if isinstance(part, dict) else part.__dict__
                    if pd.get("text"):
                        return pd["text"]
    return ""


def tool_output(input):
    if isinstance(input, str):
        return None
    for item in reversed(input):
        d = item if isinstance(item, dict) else item.__dict__
        if d.get("type") == "function_call_output":
            return d.get("output")
    return None


class ShopModel(Model):
    async def get_response(self, system_instructions, input, model_settings, tools,
                           output_schema, handoffs, tracing, *, previous_response_id=None,
                           conversation_id=None, prompt=None):
        result = tool_output(input)
        if result is not None:
            # A handoff transfer looks like {"assistant": "..."}; the specialist answers for real.
            if result.strip().startswith('{"assistant"'):
                if "refund" in (system_instructions or "").lower():
                    return ModelResponse(output=[_message(
                        "Your refund is approved and will be processed in 5 to 7 days.")],
                        usage=Usage(), response_id=None)
                return ModelResponse(output=[_message("Handled by the specialist.")],
                                     usage=Usage(), response_id=None)
            return ModelResponse(output=[_message(result)], usage=Usage(), response_id=None)
        text = last_user_text(input).lower()
        if handoffs and "refund" in text:
            return ModelResponse(output=[_tool_call(handoffs[0].tool_name, "{}")],
                                 usage=Usage(), response_id=None)
        if tools and "order" in text:
            return ModelResponse(output=[_tool_call("lookup_order", '{"order_id": "A17"}')],
                                 usage=Usage(), response_id=None)
        return ModelResponse(output=[_message("How can I help with your order?")],
                             usage=Usage(), response_id=None)

    async def stream_response(self, system_instructions, input, model_settings, tools,
                              output_schema, handoffs, tracing, *, previous_response_id=None,
                              conversation_id=None, prompt=None):
        text = "How can I help with your order?"
        for i, word in enumerate(text.split()):
            yield ResponseTextDeltaEvent(
                type="response.output_text.delta", delta=word + " ",
                content_index=0, item_id="msg", output_index=0,
                sequence_number=i, logprobs=[],
            )
        response = Response(
            id="r", created_at=0.0, model="shop-standin", object="response",
            output=[_message(text)], parallel_tool_calls=False,
            tool_choice="auto", tools=[],
        )
        yield ResponseCompletedEvent(type="response.completed", response=response,
                                     sequence_number=99)


agent = Agent(name="Shop", instructions="Help with orders.", model=ShopModel())
result = Runner.run_sync(agent, "hello")
print(result.final_output)

What the run produced

  • The agent had no tools, so get_response skipped the order branch.
  • The message was "hello", which holds no tool result and no order, so the greeting branch ran.
  • final_output is the greeting text, returned with no key and no network call.

A stand-in model vs a real model

ShopModel stand-inReal model
Needs a keyNoYes
Reply comes fromYour fixed rulesThe model's generation
Same input, same outputAlwaysNot guaranteed
Good forLearning and testsProduction answers

Where a stand-in model helps

  • Following a course or writing tests with no key and no cost.
  • Reproducing a run exactly, because the reply is fixed.
  • Checking your agent, tool and handoff wiring apart from the model.
Watch out. get_response must be async and must return a ModelResponse. Returning a plain string, or forgetting async, makes the Runner fail before your agent ever answers.
Try it yourself
  • Change the greeting string and rerun; the new text is the final output.
  • Send "help me" instead of "hello"; the greeting branch still runs.
  • Remove async from get_response and read the error the Runner raises.

Slow is fine. Stopping is the only problem.