Trimming and removing messages
Two different jobs: trim_messages limits what you send to the model at call time, while RemoveMessage deletes messages from the saved state for good.
Last updated: 27 Sep, 2026 · LangGraph 1.2
A long conversation gets expensive and slow to send in full. Trimming controls what the model sees; removing controls what the thread keeps.
The trim_messages API
from langchain_core.messages.utils import trim_messages, count_tokens_approximately
sent = trim_messages(
state["messages"], strategy="last",
token_counter=count_tokens_approximately, max_tokens=128, start_on="human",
) # only changes what you pass to the modelThe example removes every message from the saved state except the last one. Build it in four small steps.
The imports
from langgraph.graph import MessagesState, StateGraph, START, END
from langchain.messages import HumanMessage, AIMessage, RemoveMessage # RemoveMessage deletes a messageThe prune node
The one new name is RemoveMessage. You give it the id of a message you want gone, and the reducer drops that message from the state.
def prune(state):
old = state["messages"][:-1] # every message except the last
return {"messages": [RemoveMessage(id=m.id) for m in old]} # one RemoveMessage per old messagestate["messages"][:-1]is the list of messages without the last one.- For each of those, the node returns a
RemoveMessagecarrying that message'sid. - The reducer reads those markers and deletes the matching messages from the state.
Building the graph
builder = StateGraph(MessagesState)
builder.add_node("prune", prune)
builder.add_edge(START, "prune") # START -> prune -> END
builder.add_edge("prune", END)The graph has one node. The run goes from START into prune and then to END.
Invoking the graph
out = builder.compile().invoke(
{"messages": [HumanMessage("a"), AIMessage("b"), HumanMessage("c")]} # three messages go in
)
print(len(out["messages"])) # how many are left after pruningThree messages go in. The prune node removes the first two and keeps the last, so the number printed is the count of messages left.
Removing messages end to end
The four steps in one file, ready to run.
from langgraph.graph import MessagesState, StateGraph, START, END
from langchain.messages import HumanMessage, AIMessage, RemoveMessage
def prune(state):
old = state["messages"][:-1] # keep only the last message
return {"messages": [RemoveMessage(id=m.id) for m in old]}
builder = StateGraph(MessagesState)
builder.add_node("prune", prune)
builder.add_edge(START, "prune")
builder.add_edge("prune", END)
out = builder.compile().invoke(
{"messages": [HumanMessage("a"), AIMessage("b"), HumanMessage("c")]}
)
print(len(out["messages"]))Why one message remained
trim_messagesreturns a shorter list to hand the model; the saved state is untouched.RemoveMessagewith a message's id tells the reducer to delete it from the state.- The example removed every message but the last, so the saved conversation drops to one.
Trim what the model sees
trim_messages returns a shorter list to hand the model and leaves the saved state alone. With token_counter=len it counts messages, so max_tokens is a message budget. Here it keeps the last two.
from langchain_core.messages.utils import trim_messages
from langchain.messages import HumanMessage, AIMessage
chat = [HumanMessage("a"), AIMessage("b"), HumanMessage("c"), AIMessage("d")]
# token_counter=len counts messages, so max_tokens is a message budget
kept = trim_messages(chat, strategy="last", token_counter=len, max_tokens=2)
print(len(chat), "->", len(kept))
print([m.content for m in kept])Trim vs remove
| trim_messages | RemoveMessage | |
|---|---|---|
| Changes | What you send to the model | The saved state |
| Permanent | No, only this call | Yes, the message is gone |
| Use for | Fitting a long chat in the context | Pruning a thread you keep |
When to trim and when to remove
- Trim when a conversation grows past what the model should read each turn.
- Remove when the saved thread itself should be shortened or reset.
Related
- Previous: Runtime context: per-run data with context_schema
- Next: interrupt: pause for a human
- Reference: Manage memory
- Change
pruneto keep the last two messages. - Print the ids of the messages before and after pruning.
Every expert started right here.