Chat engines: remembering the conversation
A chat engine is a query interface that keeps past turns in memory, so a follow-up like "what is wrong with it" is understood from the conversation.
Last updated: 28 Sep, 2026 · LlamaIndex 0.14
A query engine answers each question alone. A real conversation refers back: the customer names a product once, then says it. A chat engine adds a memory of the turns so far and passes it to the model with each new message.
Recalling the product from earlier turns
The stand-in reads the whole conversation, not the current message alone. It looks for the product code the customer mentioned, and refuses when none has been named yet, so its answer depends on memory.
def chat(self, messages, **kwargs):
said = " ".join(m.content or "" for m in messages if m.role != MessageRole.SYSTEM)
codes = re.findall(r"LMP-\d+", said)
question = messages[-1].content or ""
if not codes:
reply = "Tell me which product you mean."
elif "?" not in question:
reply = f"Noted: your {codes[-1]}."
else:
context = "".join(m.content or "" for m in messages if m.role == MessageRole.SYSTEM)
sentences = [s.strip() for s in re.split(r"(?<=[.!?])\s+|\n+", context) if s.strip()]
best = max(sentences, key=lambda s: len(stems(question) & stems(s)), default="")
reply = f"For your {codes[-1]}: {best}"
return ChatResponse(message=ChatMessage(role=MessageRole.ASSISTANT, content=reply))Storing the turns in a memory buffer
ChatMemoryBuffer holds the conversation. Its token_limit caps how much is kept; once the turns exceed it, the oldest are dropped, which is what makes this short-term memory.
from llama_index.core.memory import ChatMemoryBuffer
memory = ChatMemoryBuffer.from_defaults(token_limit=3000)Building a chat engine over the index
chat_mode="context" retrieves chunks for each message and puts them, with the memory, in front of the model. Calling chat again reuses the same memory.
chat = index.as_chat_engine(
chat_mode="context", llm=RecallLLM(), memory=memory, similarity_top_k=2
)A follow-up that leans on the first turn
from llama_index.core.memory import ChatMemoryBuffer
from recall_llm import RecallLLM
memory = ChatMemoryBuffer.from_defaults(token_limit=3000)
chat = index.as_chat_engine(chat_mode="context", llm=RecallLLM(), memory=memory, similarity_top_k=2)
print("turn 1:", chat.chat("I bought the LMP-204 desk lamp."))
print("turn 2:", chat.chat("What is wrong with it?"))
print("turns in memory:", len(memory.get_all()))
fresh_memory = ChatMemoryBuffer.from_defaults(token_limit=3000)
fresh = index.as_chat_engine(chat_mode="context", llm=RecallLLM(), memory=fresh_memory, similarity_top_k=2)
print("no memory:", fresh.chat("What is wrong with it?"))Reading the three replies
- Turn 1 names the product. The customer says LMP-204, and the engine notes it. The memory now holds two turns, the user's and the reply.
- Turn 2 says only "it". The engine finds LMP-204 in the remembered turns, retrieves the lamp chunk, and answers about the cable fault.
- The no-memory run fails to recall. A fresh engine asked the same follow-up has no earlier turn, so it cannot resolve "it" and asks which product you mean.
With memory vs without memory
| Chat engine with memory | Query engine per question | |
|---|---|---|
| Follow-up "it" | Resolved from earlier turns | Unresolved, no history |
| State kept | The turns, up to token_limit | None between questions |
| Answer to turn 2 | About the LMP-204 | Cannot tell which product |
When to keep chat memory
- The assistant holds a back-and-forth, where later questions refer to earlier ones.
- A customer gives a detail once, an order or a product, and expects it remembered.
- You want short-term context without re-sending the whole history yourself each turn.
token_limit the oldest turns are dropped, so a very early detail can fall out. Keep the limit sized to the turns that matter, and store anything durable elsewhere.Related
- Change turn 2 to
"Is it refundable?"and see the answer follow the remembered product. - Ask a fresh engine
"What is wrong with it?"first and confirm it cannot answer. - Set
token_limitvery low, add several turns, and watch the earliest drop.
Every expert started right here.