Trimming and removing messages
Trimming limits what you send the model at call time with trim_messages, while RemoveMessage deletes messages from the saved state for good.
Last updated: 27 Sep, 2026 · LangChain 1.4
A long conversation gets slow and costly to send in full. Trimming controls what the model sees on a call; removing controls what the thread keeps.
Trimming what the model sees
trim_messages returns a shorter list and leaves the saved conversation alone. With token_counter=len it counts messages, so max_tokens is a message budget; here it keeps the last two.
kept = trim_messages(chat, strategy="last",
token_counter=len, max_tokens=2) # keep the last 2Trimming a four-message chat
The whole snippet, printing the count before and after.
from langchain_core.messages.utils import trim_messages
from langchain.messages import HumanMessage, AIMessage
chat = [HumanMessage("a"), AIMessage("b"), HumanMessage("c"), AIMessage("d")]
# token_counter=len counts messages, so max_tokens is a message budget
kept = trim_messages(chat, strategy="last", token_counter=len, max_tokens=2)
print(len(chat), "->", len(kept))
print([m.content for m in kept])What trimming did
- The chat had four messages; the model should see fewer.
strategy="last"kept the most recent two,candd.- The original
chatlist is untouched; only what you pass the model changed.
Removing a message from the saved state
Trimming does not delete anything. To drop a message from a thread for good, return a RemoveMessage with its id from a node; the messages reducer deletes it.
from langchain.messages import RemoveMessage
# in a node: delete every message but the last from the saved thread
return {"messages": [RemoveMessage(id=m.id) for m in state["messages"][:-1]]}Trim vs remove
| trim_messages | RemoveMessage | |
|---|---|---|
| Changes | What you send the model | The saved state |
| Permanent | No, only this call | Yes, the message is gone |
| Use for | Fitting a long chat in the window | Pruning a thread you keep |
When to trim or remove
- Trim when a conversation grows past what the model should read each turn.
- Remove when the saved thread itself should be shortened or reset.
- Summarize instead when older turns still matter but should cost less.
Related
- Previous: Short-term memory with a checkpointer
- Next: Memory that outlives the conversation
- Reference: LangChain docs
- Change
max_tokensto 3 and read which messages are kept. - Switch
strategyto "first" and compare. - Remove only the first message instead of all but the last.
Every expert started right here.