LlamaIndexllama-index-core 0.14 · Python 3.10+
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
30 small wins to finish your pathNext lesson →

Observability: tracing what a run did

Instrumentation is a way to receive an event each time LlamaIndex embeds, retrieves or answers, so you can see what a run did.

Last updated: 28 Sep, 2026 · LlamaIndex 0.14

When an answer looks wrong, the question is usually which chunks were retrieved. Instead of adding prints everywhere, attach one event handler to the dispatcher and read the steps of every run.

An event handler on the dispatcher

A handler is a class with a handle method. The global dispatcher sends it every event; you keep the ones you care about.

python
from llama_index.core.instrumentation import get_dispatcher
from llama_index.core.instrumentation.event_handlers import BaseEventHandler
from llama_index.core.instrumentation.events.retrieval import RetrievalEndEvent

class Trace(BaseEventHandler):
    events: list = []
    def handle(self, event, **kwargs):
        if isinstance(event, RetrievalEndEvent):
            self.events.append([n.metadata["file_name"] for n in event.nodes])

Attaching the handler

One call registers the handler. From then on, every embed, retrieval and query in the process reaches it.

python
trace = Trace()
get_dispatcher().add_event_handler(trace)

Reading the trace after a query

Run a normal query, then read what the handler collected. Each event is a real step the library took, not a guess.

python
answer = engine.query("How long until my refund money reaches my card?")
for line in trace.events:
    print("  -", line)

A query with its steps captured

The whole program watches for three kinds of event, runs one query, and prints both the answer and the steps behind it.

Example
from llama_index.core import Settings, SimpleDirectoryReader, VectorStoreIndex
from llama_index.core.instrumentation import get_dispatcher
from llama_index.core.instrumentation.event_handlers import BaseEventHandler
from llama_index.core.instrumentation.events.embedding import EmbeddingEndEvent
from llama_index.core.instrumentation.events.retrieval import RetrievalEndEvent
from llama_index.core.instrumentation.events.query import QueryEndEvent
from llama_index.embeddings.huggingface import HuggingFaceEmbedding

from extractive_llm import ExtractiveLLM

Settings.embed_model = HuggingFaceEmbedding(model_name="sentence-transformers/all-MiniLM-L6-v2")


class Trace(BaseEventHandler):
    events: list = []

    def handle(self, event, **kwargs):
        if isinstance(event, EmbeddingEndEvent):
            self.events.append(f"embedded {len(event.embeddings)} text(s)")
        elif isinstance(event, RetrievalEndEvent):
            names = [n.metadata["file_name"] for n in event.nodes]
            self.events.append(f"retrieved {names}")
        elif isinstance(event, QueryEndEvent):
            self.events.append("query finished")


trace = Trace()
get_dispatcher().add_event_handler(trace)

docs = SimpleDirectoryReader("help").load_data()
index = VectorStoreIndex.from_documents(docs)
engine = index.as_query_engine(llm=ExtractiveLLM(), similarity_top_k=2)

trace.events.clear()
answer = engine.query("How long until my refund money reaches my card?")
print("answer:", answer)
print("captured events:")
for line in trace.events:
    print("  -", line)

What the captured events tell you

  • The embed event shows the question was turned into one vector before the search.
  • The retrieval event names the files that were pulled, which is where most wrong answers are decided.
  • The query event marks the end of the run, after the answer was built.
  • Change the question and the retrieved files change, so the trace reflects the real run.

Prints vs an event handler

ApproachCoversCost
Scattered printsOnly where you added themEdits across the code, removed later
Event handlerEvery embed, retrieve and queryOne class, attached once

When you reach for tracing

  • An answer is wrong and you need to see which chunks were retrieved.
  • Counting retrievals or embeddings to find a slow or costly step.
  • Sending spans to a hosted tracing tool through a provided handler.
Watch out. A handler stays attached for the whole process and receives events from every run, so a shared list grows across queries. Clear it, or filter by run, before you read one query's steps.
Try it yourself
  • Add a handler branch for EmbeddingEndEvent and print how many texts were embedded when the index is built.
  • Ask a lamp question and read which file the trace shows.
  • Remove trace.events.clear() and run two queries to see the list grow.

Every expert started right here.