Shaping the model with ModelSettings
ModelSettings is a group of options like temperature and max_tokens that the SDK passes to the model on every call to shape its replies.
Last updated: 28 Sep, 2026 · openai-agents 0.22.3
You build a ModelSettings object and attach it to an agent with model_settings=. On each call the SDK forwards it to the model's get_response. A real model reads these values; the stand-in here reads temperature and puts it in the reply, so changing the setting changes the output.
The ModelSettings fields
from agents import ModelSettings
settings = ModelSettings(
temperature=0.2, # lower means steadier, less varied wording
max_tokens=256, # a cap on how long the reply may be
top_p=1.0, # another knob on how the model samples words
)Building a settings object
Set only the fields you care about; the rest keep their defaults. Here two values are read back to show they are stored on the object.
settings = ModelSettings(temperature=0.2, max_tokens=256)
print("temperature:", settings.temperature)
print("max_tokens:", settings.max_tokens)A model that reads the settings
The stand-in reads model_settings.temperature and puts it in the reply, so the value the SDK forwards is visible in the output.
class SettingsModel(Model):
async def get_response(self, system_instructions, input, model_settings,
tools, output_schema, handoffs, tracing, **k):
temp = getattr(model_settings, "temperature", None) # the forwarded value
return ModelResponse(output=[_message(f"Replying at temperature {temp}.")],
usage=Usage(), response_id=None)Attaching settings to the agent
Pass the settings to the agent. The SDK forwards them on every model call, and the reply shows the temperature the agent was given.
steady = Agent(name="Shop", instructions="Help with orders.",
model=SettingsModel(), model_settings=ModelSettings(temperature=0.2))
print(Runner.run_sync(steady, "hello").final_output) # Replying at temperature 0.2.Passing temperature through to the reply
The whole program in one file. It reads the two settings back, then runs the same message at two temperatures. The reply carries whichever temperature the agent was given, so the setting is visible in the output.
from agents import Agent, Runner, ModelSettings, set_tracing_disabled
from agents.models.interface import Model
from agents.items import ModelResponse
from agents.usage import Usage
from openai.types.responses import ResponseOutputMessage, ResponseOutputText
set_tracing_disabled(True)
def _message(text):
return ResponseOutputMessage(
id="msg", role="assistant", type="message", status="completed",
content=[ResponseOutputText(text=text, type="output_text", annotations=[])],
)
class SettingsModel(Model):
async def get_response(self, system_instructions, input, model_settings, tools,
output_schema, handoffs, tracing, **k):
temp = getattr(model_settings, "temperature", None)
return ModelResponse(output=[_message(f"Replying at temperature {temp}.")],
usage=Usage(), response_id=None)
async def stream_response(self, *a, **k):
raise NotImplementedError
settings = ModelSettings(temperature=0.2, max_tokens=256)
print("temperature:", settings.temperature)
print("max_tokens:", settings.max_tokens)
steady = Agent(name="Shop", instructions="Help with orders.",
model=SettingsModel(), model_settings=settings)
varied = Agent(name="Shop", instructions="Help with orders.",
model=SettingsModel(), model_settings=ModelSettings(temperature=0.9))
print("At 0.2:", Runner.run_sync(steady, "hello").final_output)
print("At 0.9:", Runner.run_sync(varied, "hello").final_output)What each setting asks the model to do
- temperature is the value this stand-in reads and prints; a low number keeps a real model's wording steady across runs.
- max_tokens caps the length of a real model's reply, so a long answer is cut off at the limit.
- top_p is a second sampling knob for a real model; leaving it at 1.0 lets the temperature do the shaping.
Default behaviour vs shaped behaviour
| Setting | Left at default | Set explicitly |
|---|---|---|
temperature | The model's own default varies wording | A low value keeps wording steady |
max_tokens | No cap you chose | The reply stops at your limit |
When to set ModelSettings
- Lowering temperature so a support answer reads the same way each time.
- Capping max_tokens so a reply stays short and cheap.
- Preparing the agent before you swap in a real model that reads these values.
temperature to show the value arrives. Fields like max_tokens and top_p shape a real model's reply, not this one, so set them before you swap a real model in.Related
- Previous: Tracing a run with spans
- Next: Capping a run with max_turns
- Reference: ModelSettings
- Change
temperatureto 0.7 and watch the printed reply change to match. - Print
settingsitself to see every field at once. - Add
top_p=0.5to the settings and confirm the reply still reports only the temperature.
Slow is fine. Stopping is the only problem.