LlamaIndexllama-index-core 0.14 · Python 3.10+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
30 small wins to finish your pathNext lesson →

Chat engines: remembering the conversation

A chat engine is a query interface that keeps past turns in memory, so a follow-up like "what is wrong with it" is understood from the conversation.

Last updated: 28 Sep, 2026 · LlamaIndex 0.14

A query engine answers each question alone. A real conversation refers back: the customer names a product once, then says it. A chat engine adds a memory of the turns so far and passes it to the model with each new message.

Recalling the product from earlier turns

The stand-in reads the whole conversation, not the current message alone. It looks for the product code the customer mentioned, and refuses when none has been named yet, so its answer depends on memory.

python
def chat(self, messages, **kwargs):
    said = " ".join(m.content or "" for m in messages if m.role != MessageRole.SYSTEM)
    codes = re.findall(r"LMP-\d+", said)
    question = messages[-1].content or ""
    if not codes:
        reply = "Tell me which product you mean."
    elif "?" not in question:
        reply = f"Noted: your {codes[-1]}."
    else:
        context = "".join(m.content or "" for m in messages if m.role == MessageRole.SYSTEM)
        sentences = [s.strip() for s in re.split(r"(?<=[.!?])\s+|\n+", context) if s.strip()]
        best = max(sentences, key=lambda s: len(stems(question) & stems(s)), default="")
        reply = f"For your {codes[-1]}: {best}"
    return ChatResponse(message=ChatMessage(role=MessageRole.ASSISTANT, content=reply))

Storing the turns in a memory buffer

ChatMemoryBuffer holds the conversation. Its token_limit caps how much is kept; once the turns exceed it, the oldest are dropped, which is what makes this short-term memory.

python
from llama_index.core.memory import ChatMemoryBuffer

memory = ChatMemoryBuffer.from_defaults(token_limit=3000)

Building a chat engine over the index

chat_mode="context" retrieves chunks for each message and puts them, with the memory, in front of the model. Calling chat again reuses the same memory.

python
chat = index.as_chat_engine(
    chat_mode="context", llm=RecallLLM(), memory=memory, similarity_top_k=2
)
Project files used on this pageThis lesson builds on a project from earlier lessons. The code below imports these files. Click a file to see its code, or follow the link to the lesson that wrote it. To run the code yourself, keep them in the same folder.
View the code here
help/lamps.md
# Lamps

The LMP-204 desk lamp has a known cable fault. Stop using a lamp with a damaged cable and we will replace it free of charge.

All lamps come with a two year guarantee against electrical faults.

Bulbs are not covered by the refund policy once they have been used.

The LMP-310 floor lamp needs a bulb with an E27 fitting, which is sold separately.
recall_llm.py
import re

from llama_index.core.llms import (
    ChatMessage,
    ChatResponse,
    CompletionResponse,
    CustomLLM,
    LLMMetadata,
    MessageRole,
)
from llama_index.core.llms.callbacks import llm_chat_callback, llm_completion_callback


def stems(text):
    return {w[:5] for w in re.findall(r"[a-z0-9-]+", text.lower()) if len(w) > 3}


class RecallLLM(CustomLLM):
    """Answers about the product the user named earlier in the conversation."""

    @property
    def metadata(self):
        return LLMMetadata(model_name="recall", is_chat_model=True)

    @llm_chat_callback()
    def chat(self, messages, **kwargs):
        said = " ".join(m.content or "" for m in messages if m.role != MessageRole.SYSTEM)
        codes = re.findall(r"LMP-\d+", said)
        question = messages[-1].content or ""
        if not codes:
            reply = "Tell me which product you mean."
        elif "?" not in question:
            reply = f"Noted: your {codes[-1]}."
        else:
            context = "".join(m.content or "" for m in messages if m.role == MessageRole.SYSTEM)
            sentences = [s.strip() for s in re.split(r"(?<=[.!?])\s+|\n+", context) if s.strip()]
            best = max(sentences, key=lambda s: len(stems(question) & stems(s)), default="")
            reply = f"For your {codes[-1]}: {best}"
        return ChatResponse(message=ChatMessage(role=MessageRole.ASSISTANT, content=reply))

    @llm_chat_callback()
    def stream_chat(self, messages, **kwargs):
        yield self.chat(messages, **kwargs)

    @llm_completion_callback()
    def complete(self, prompt, formatted=False, **kwargs):
        return CompletionResponse(text="")

    @llm_completion_callback()
    def stream_complete(self, prompt, formatted=False, **kwargs):
        yield self.complete(prompt)
help/refunds.md
# Refunds

You can get a full refund within 30 days of delivery. The money goes back to the card you paid with within 5 working days of us receiving the item.

Items bought in a sale can be refunded too, but the delivery charge is not returned.

To start a refund, open the order in your account and choose Return an item. Print the label and drop the parcel at any post office.

Personalised items cannot be refunded unless they arrive damaged.
help/delivery.md
# Delivery

Standard delivery takes 3 to 5 working days and is free on orders over 40.

Express delivery arrives the next working day if you order before 2pm. It costs 6.

We deliver to the mainland only. Parcels to islands take 2 extra working days.

If a parcel has not arrived after 10 working days, contact us and we will send a replacement.

A follow-up that leans on the first turn

Example
from llama_index.core.memory import ChatMemoryBuffer
from recall_llm import RecallLLM

memory = ChatMemoryBuffer.from_defaults(token_limit=3000)
chat = index.as_chat_engine(chat_mode="context", llm=RecallLLM(), memory=memory, similarity_top_k=2)
print("turn 1:", chat.chat("I bought the LMP-204 desk lamp."))
print("turn 2:", chat.chat("What is wrong with it?"))
print("turns in memory:", len(memory.get_all()))

fresh_memory = ChatMemoryBuffer.from_defaults(token_limit=3000)
fresh = index.as_chat_engine(chat_mode="context", llm=RecallLLM(), memory=fresh_memory, similarity_top_k=2)
print("no memory:", fresh.chat("What is wrong with it?"))

Reading the three replies

  • Turn 1 names the product. The customer says LMP-204, and the engine notes it. The memory now holds two turns, the user's and the reply.
  • Turn 2 says only "it". The engine finds LMP-204 in the remembered turns, retrieves the lamp chunk, and answers about the cable fault.
  • The no-memory run fails to recall. A fresh engine asked the same follow-up has no earlier turn, so it cannot resolve "it" and asks which product you mean.

With memory vs without memory

Chat engine with memoryQuery engine per question
Follow-up "it"Resolved from earlier turnsUnresolved, no history
State keptThe turns, up to token_limitNone between questions
Answer to turn 2About the LMP-204Cannot tell which product

When to keep chat memory

  • The assistant holds a back-and-forth, where later questions refer to earlier ones.
  • A customer gives a detail once, an order or a product, and expects it remembered.
  • You want short-term context without re-sending the whole history yourself each turn.
Watch out
Memory is not free context: a long conversation grows the prompt and its cost, and past token_limit the oldest turns are dropped, so a very early detail can fall out. Keep the limit sized to the turns that matter, and store anything durable elsewhere.
Try it yourself
  • Change turn 2 to "Is it refundable?" and see the answer follow the remembered product.
  • Ask a fresh engine "What is wrong with it?" first and confirm it cannot answer.
  • Set token_limit very low, add several turns, and watch the earliest drop.

Every expert started right here.