0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
27 small wins to finish your pathNext lesson →

Shaping the model with ModelSettings

ModelSettings is a group of options like temperature and max_tokens that the SDK passes to the model on every call to shape its replies.

Last updated: 28 Sep, 2026 · openai-agents 0.22.3

You build a ModelSettings object and attach it to an agent with model_settings=. On each call the SDK forwards it to the model's get_response. A real model reads these values; the stand-in here reads temperature and puts it in the reply, so changing the setting changes the output.

The ModelSettings fields

python
from agents import ModelSettings

settings = ModelSettings(
    temperature=0.2,   # lower means steadier, less varied wording
    max_tokens=256,    # a cap on how long the reply may be
    top_p=1.0,         # another knob on how the model samples words
)

Building a settings object

Set only the fields you care about; the rest keep their defaults. Here two values are read back to show they are stored on the object.

python
settings = ModelSettings(temperature=0.2, max_tokens=256)
print("temperature:", settings.temperature)
print("max_tokens:", settings.max_tokens)

A model that reads the settings

The stand-in reads model_settings.temperature and puts it in the reply, so the value the SDK forwards is visible in the output.

python
class SettingsModel(Model):
    async def get_response(self, system_instructions, input, model_settings,
                           tools, output_schema, handoffs, tracing, **k):
        temp = getattr(model_settings, "temperature", None)   # the forwarded value
        return ModelResponse(output=[_message(f"Replying at temperature {temp}.")],
                             usage=Usage(), response_id=None)

Attaching settings to the agent

Pass the settings to the agent. The SDK forwards them on every model call, and the reply shows the temperature the agent was given.

python
steady = Agent(name="Shop", instructions="Help with orders.",
               model=SettingsModel(), model_settings=ModelSettings(temperature=0.2))
print(Runner.run_sync(steady, "hello").final_output)   # Replying at temperature 0.2.

Passing temperature through to the reply

The whole program in one file. It reads the two settings back, then runs the same message at two temperatures. The reply carries whichever temperature the agent was given, so the setting is visible in the output.

Example
from agents import Agent, Runner, ModelSettings, set_tracing_disabled
from agents.models.interface import Model
from agents.items import ModelResponse
from agents.usage import Usage
from openai.types.responses import ResponseOutputMessage, ResponseOutputText
set_tracing_disabled(True)


def _message(text):
    return ResponseOutputMessage(
        id="msg", role="assistant", type="message", status="completed",
        content=[ResponseOutputText(text=text, type="output_text", annotations=[])],
    )


class SettingsModel(Model):
    async def get_response(self, system_instructions, input, model_settings, tools,
                           output_schema, handoffs, tracing, **k):
        temp = getattr(model_settings, "temperature", None)
        return ModelResponse(output=[_message(f"Replying at temperature {temp}.")],
                             usage=Usage(), response_id=None)

    async def stream_response(self, *a, **k):
        raise NotImplementedError


settings = ModelSettings(temperature=0.2, max_tokens=256)
print("temperature:", settings.temperature)
print("max_tokens:", settings.max_tokens)

steady = Agent(name="Shop", instructions="Help with orders.",
               model=SettingsModel(), model_settings=settings)
varied = Agent(name="Shop", instructions="Help with orders.",
               model=SettingsModel(), model_settings=ModelSettings(temperature=0.9))
print("At 0.2:", Runner.run_sync(steady, "hello").final_output)
print("At 0.9:", Runner.run_sync(varied, "hello").final_output)

What each setting asks the model to do

  • temperature is the value this stand-in reads and prints; a low number keeps a real model's wording steady across runs.
  • max_tokens caps the length of a real model's reply, so a long answer is cut off at the limit.
  • top_p is a second sampling knob for a real model; leaving it at 1.0 lets the temperature do the shaping.

Default behaviour vs shaped behaviour

SettingLeft at defaultSet explicitly
temperatureThe model's own default varies wordingA low value keeps wording steady
max_tokensNo cap you choseThe reply stops at your limit

When to set ModelSettings

  • Lowering temperature so a support answer reads the same way each time.
  • Capping max_tokens so a reply stays short and cheap.
  • Preparing the agent before you swap in a real model that reads these values.
Watch out. This stand-in reads only temperature to show the value arrives. Fields like max_tokens and top_p shape a real model's reply, not this one, so set them before you swap a real model in.
Try it yourself
  • Change temperature to 0.7 and watch the printed reply change to match.
  • Print settings itself to see every field at once.
  • Add top_p=0.5 to the settings and confirm the reply still reports only the temperature.

Slow is fine. Stopping is the only problem.