Trimming history with a history processor
A history processor is a function that rewrites the message list before each model request. Pydantic AI runs it through the ProcessHistory capability, so you can drop or shorten old messages before they are sent.
Last updated: 28 Sep, 2026 · Pydantic AI 2.51
The last lesson sent the whole conversation on every run, and warned that the cost grows with it. A history processor is where you cut it back: keep the recent messages, drop the old ones, or replace a long stretch with a summary, without changing what you store.
In this version of Pydantic AI the processor is attached with the ProcessHistory capability. Older code passed a history_processors= argument to Agent; that name is gone in 2.x, and the capability is the current way.
The model that counts its messages
So the trimming is visible, the stand-in answers with the number of messages the request carried.
from pydantic_ai import Agent, ModelResponse, TextPart
from pydantic_ai.models.function import FunctionModel
from pydantic_ai.capabilities import ProcessHistory
def report(messages, info):
# answer with the number of messages this request carried
return ModelResponse(parts=[TextPart(f"model saw {len(messages)} messages")])Writing the processor
A processor takes the message list and returns the list the model should receive. This one keeps only the last two messages.
def keep_recent(messages):
# the processor returns the list the model will receive
return messages[-2:] # only the two most recent messagesAttaching it with ProcessHistory
One agent runs plain, the other has the processor in its capabilities. Everything else is the same.
plain = Agent(FunctionModel(report))
trimmed = Agent(FunctionModel(report), capabilities=[ProcessHistory(keep_recent)])The processor cutting a request down
Build a real conversation with the plain agent, then send one more turn on that same history through each agent.
history = None
for text in ["My order is A-1001", "Where is it?", "Is it late?"]:
history = plain.run_sync(text, message_history=history).all_messages()
print("stored history:", len(history), "messages")
# one more turn on that same history, without and with the processor
print(plain.run_sync("And the refund?", message_history=history).output)
print(trimmed.run_sync("And the refund?", message_history=history).output)What the two agents saw
- The stored history holds six messages, and neither agent changes it.
- The plain agent sent all six plus the new prompt, so the model saw seven.
- The trimmed agent ran
keep_recentfirst, so the model saw only the last two messages of that request. - The processor changes what is sent, not what you keep: your stored list is untouched, so nothing is lost.
No processor vs ProcessHistory
| No processor | ProcessHistory(keep_recent) | |
|---|---|---|
| Messages sent to the model | The whole history, every run | Only what the processor returns |
| Cost as the chat grows | Rises with each turn | Held roughly flat |
| Stored history | Unchanged | Unchanged |
When you reach for a history processor
- A long chat where the token cost of resending everything has grown too high.
- Keeping the most recent turns and dropping the rest, when only recent context matters.
- Replacing an old stretch of the conversation with a short summary you compute.
Related
- Previous: Message history: continuing a conversation
- Next: Saving messages to JSON and back
- Reference: History processors
- Change
keep_recenttomessages[-4:]and read the new count. - Add a second processor that also drops any system prompt, and pass both to
capabilities. - Give
keep_recentactxfirst argument and printctx.usagefrom inside it.
Little by little, you're building something great.