Observability: tracing what a run did
Instrumentation is a way to receive an event each time LlamaIndex embeds, retrieves or answers, so you can see what a run did.
Last updated: 28 Sep, 2026 · LlamaIndex 0.14
When an answer looks wrong, the question is usually which chunks were retrieved. Instead of adding prints everywhere, attach one event handler to the dispatcher and read the steps of every run.
An event handler on the dispatcher
A handler is a class with a handle method. The global dispatcher sends it every event; you keep the ones you care about.
from llama_index.core.instrumentation import get_dispatcher
from llama_index.core.instrumentation.event_handlers import BaseEventHandler
from llama_index.core.instrumentation.events.retrieval import RetrievalEndEvent
class Trace(BaseEventHandler):
events: list = []
def handle(self, event, **kwargs):
if isinstance(event, RetrievalEndEvent):
self.events.append([n.metadata["file_name"] for n in event.nodes])Attaching the handler
One call registers the handler. From then on, every embed, retrieval and query in the process reaches it.
trace = Trace()
get_dispatcher().add_event_handler(trace)Reading the trace after a query
Run a normal query, then read what the handler collected. Each event is a real step the library took, not a guess.
answer = engine.query("How long until my refund money reaches my card?")
for line in trace.events:
print(" -", line)A query with its steps captured
The whole program watches for three kinds of event, runs one query, and prints both the answer and the steps behind it.
from llama_index.core import Settings, SimpleDirectoryReader, VectorStoreIndex
from llama_index.core.instrumentation import get_dispatcher
from llama_index.core.instrumentation.event_handlers import BaseEventHandler
from llama_index.core.instrumentation.events.embedding import EmbeddingEndEvent
from llama_index.core.instrumentation.events.retrieval import RetrievalEndEvent
from llama_index.core.instrumentation.events.query import QueryEndEvent
from llama_index.embeddings.huggingface import HuggingFaceEmbedding
from extractive_llm import ExtractiveLLM
Settings.embed_model = HuggingFaceEmbedding(model_name="sentence-transformers/all-MiniLM-L6-v2")
class Trace(BaseEventHandler):
events: list = []
def handle(self, event, **kwargs):
if isinstance(event, EmbeddingEndEvent):
self.events.append(f"embedded {len(event.embeddings)} text(s)")
elif isinstance(event, RetrievalEndEvent):
names = [n.metadata["file_name"] for n in event.nodes]
self.events.append(f"retrieved {names}")
elif isinstance(event, QueryEndEvent):
self.events.append("query finished")
trace = Trace()
get_dispatcher().add_event_handler(trace)
docs = SimpleDirectoryReader("help").load_data()
index = VectorStoreIndex.from_documents(docs)
engine = index.as_query_engine(llm=ExtractiveLLM(), similarity_top_k=2)
trace.events.clear()
answer = engine.query("How long until my refund money reaches my card?")
print("answer:", answer)
print("captured events:")
for line in trace.events:
print(" -", line)What the captured events tell you
- The embed event shows the question was turned into one vector before the search.
- The retrieval event names the files that were pulled, which is where most wrong answers are decided.
- The query event marks the end of the run, after the answer was built.
- Change the question and the retrieved files change, so the trace reflects the real run.
Prints vs an event handler
| Approach | Covers | Cost |
|---|---|---|
| Scattered prints | Only where you added them | Edits across the code, removed later |
| Event handler | Every embed, retrieve and query | One class, attached once |
When you reach for tracing
- An answer is wrong and you need to see which chunks were retrieved.
- Counting retrievals or embeddings to find a slow or costly step.
- Sending spans to a hosted tracing tool through a provided handler.
Related
- Previous: Multi-agent: routing between two agents
- Next: Persisting an index: not embedding twice
- See also: Query engines: the prompt the model receives
- Reference: Observability
- Add a handler branch for
EmbeddingEndEventand print how many texts were embedded when the index is built. - Ask a lamp question and read which file the trace shows.
- Remove
trace.events.clear()and run two queries to see the list grow.
Every expert started right here.