Trimming and removing messages
Trimming limits what you send the model at call time with trim_messages, while RemoveMessage deletes messages from the saved state for good.
Last updated: 27 Sep, 2026 · LangChain 1.4
A long conversation gets slow and costly to send in full. Trimming controls what the model sees on a call; removing controls what the thread keeps. This is the compressed context from Runtime context: who is asking.
Trimming what the model sees
trim_messages returns a shorter list and leaves the saved conversation alone. With token_counter=len it counts messages, so max_tokens is a message budget; here it keeps the last two. The snippet prints the count before and after.
from langchain_core.messages.utils import trim_messages
from langchain.messages import HumanMessage, AIMessage
chat = [HumanMessage("a"), AIMessage("b"), HumanMessage("c"), AIMessage("d")]
# token_counter=len counts messages, so max_tokens is a message budget
kept = trim_messages(chat, strategy="last", token_counter=len, max_tokens=2)
print(len(chat), "->", len(kept))
print([m.content for m in kept])4 -> 2 ['c', 'd']
What trimming did
- Four messages went in; the budget was two.
strategy="last"kept the most recent two,candd.- The original
chatlist is untouched; only what you pass the model changed.
In an agent you want this on every model call, not only once. Middleware does that; Changing the call: wrap_model_call shows how, later in the course.
Removing a message from the saved state
Trimming does not delete anything. To drop messages from a saved thread for good, send RemoveMessage objects, one per message id, to update_state. The agent here is the checkpointed one from the short-term memory lesson.
from langchain.messages import RemoveMessage
messages = agent.get_state(thread).values["messages"]
agent.update_state(thread, {"messages": [RemoveMessage(id=m.id) for m in messages[1:-1]]})Removing the lookup from a thread
Ask one question on a thread, then remove everything between the question and the answer: the tool call and its result go together, so the thread stays valid.
from langchain.messages import RemoveMessage
thread = {"configurable": {"thread_id": "ravi-2"}} # a new, empty thread
agent.invoke({"messages": [{"role": "user", "content": "Where is A17?"}]}, thread)
messages = agent.get_state(thread).values["messages"]
print([m.type for m in messages])
agent.update_state(thread, {"messages": [RemoveMessage(id=m.id) for m in messages[1:-1]]})
print([m.type for m in agent.get_state(thread).values["messages"]])['human', 'ai', 'tool', 'ai'] ['human', 'ai']
The saved thread went from the question, the tool call, the tool result and the answer down to the question and the answer. The next call on this thread sees those two before its new question.
Trim vs remove
| trim_messages | RemoveMessage | |
|---|---|---|
| Changes | What you send the model | The saved state |
| Permanent | No, only this call | Yes, the message is gone |
| Use for | Fitting a long chat into the context window, the most the model can read at once | Pruning a thread you keep |
When to trim or remove
- Trim when a conversation grows past what the model should read each turn.
- Remove when the saved thread itself should be shortened or reset.
- Summarize instead when older turns still matter but should cost less: Summarization with SummarizationMiddleware.
Related
- Previous: Short-term memory with a checkpointer
- Next: Memory that outlives the conversation
- Reference: LangChain docs
- Change
max_tokensto 3 and read which messages are kept. - Switch
strategyto "first" and compare. - Remove only the first message, the question, and print the types left in the thread.
Every expert started right here.