Observability: tracing what a run did
Instrumentation is a way to receive an event each time LlamaIndex embeds, retrieves or answers, so you can see what a run did.
Last updated: 28 Sep, 2026 · LlamaIndex 0.14
When an answer looks wrong, the question is usually which chunks were retrieved. Instead of adding prints everywhere, attach one event handler to the dispatcher and read the steps of every run.
An event handler on the dispatcher
A handler is a class with a handle method. The global dispatcher sends it every event; you keep the ones you care about.
from llama_index.core.instrumentation import get_dispatcher
from llama_index.core.instrumentation.event_handlers import BaseEventHandler
from llama_index.core.instrumentation.events.retrieval import RetrievalEndEvent
class Trace(BaseEventHandler):
events: list = []
def handle(self, event, **kwargs):
if isinstance(event, RetrievalEndEvent):
self.events.append([n.metadata["file_name"] for n in event.nodes])Attaching the handler
One call registers the handler. From then on, every embed, retrieval and query in the process reaches it.
trace = Trace()
get_dispatcher().add_event_handler(trace)Reading the trace after a query
Run a normal query, then read what the handler collected. Each event is a real step the library took, not a guess.
answer = engine.query("How long until my refund money reaches my card?")
for line in trace.events:
print(" -", line)View the code here
# Lamps
The LMP-204 desk lamp has a known cable fault. Stop using a lamp with a damaged cable and we will replace it free of charge.
All lamps come with a two year guarantee against electrical faults.
Bulbs are not covered by the refund policy once they have been used.
The LMP-310 floor lamp needs a bulb with an E27 fitting, which is sold separately.
# Refunds
You can get a full refund within 30 days of delivery. The money goes back to the card you paid with within 5 working days of us receiving the item.
Items bought in a sale can be refunded too, but the delivery charge is not returned.
To start a refund, open the order in your account and choose Return an item. Print the label and drop the parcel at any post office.
Personalised items cannot be refunded unless they arrive damaged.
# Delivery
Standard delivery takes 3 to 5 working days and is free on orders over 40.
Express delivery arrives the next working day if you order before 2pm. It costs 6.
We deliver to the mainland only. Parcels to islands take 2 extra working days.
If a parcel has not arrived after 10 working days, contact us and we will send a replacement.
import re
from llama_index.core.llms import CompletionResponse, CustomLLM, LLMMetadata
from llama_index.core.llms.callbacks import llm_completion_callback
def stems(text):
"""Words longer than three letters, cut to five letters, so refund and refunds match."""
return {w[:5] for w in re.findall(r"[a-z0-9-]+", text.lower()) if len(w) > 3}
class ExtractiveLLM(CustomLLM):
"""Answers with the context sentence that shares most words with the question."""
@property
def metadata(self):
return LLMMetadata(model_name="extractive")
@llm_completion_callback()
def complete(self, prompt, formatted=False, **kwargs):
context = prompt.split("---------------------")[1]
question = prompt.split("Query:")[1].split("Answer:")[0]
asked = stems(question)
sentences = [s.strip() for s in re.split(r"(?<=[.!?])\s+|\n+", context)]
sentences = [s for s in sentences if s and not s.startswith("#") and ": " not in s[:20]]
best = max(sentences, key=lambda s: len(asked & stems(s)), default="")
if len(asked & stems(best)) < 2:
return CompletionResponse(text="I could not find that in the documents.")
return CompletionResponse(text=best)
@llm_completion_callback()
def stream_complete(self, prompt, formatted=False, **kwargs):
yield self.complete(prompt)
A query with its steps captured
The whole program watches for three kinds of event, runs one query, and prints both the answer and the steps behind it.
from llama_index.core import Settings, SimpleDirectoryReader, VectorStoreIndex
from llama_index.core.instrumentation import get_dispatcher
from llama_index.core.instrumentation.event_handlers import BaseEventHandler
from llama_index.core.instrumentation.events.embedding import EmbeddingEndEvent
from llama_index.core.instrumentation.events.retrieval import RetrievalEndEvent
from llama_index.core.instrumentation.events.query import QueryEndEvent
from llama_index.embeddings.huggingface import HuggingFaceEmbedding
from extractive_llm import ExtractiveLLM
Settings.embed_model = HuggingFaceEmbedding(model_name="sentence-transformers/all-MiniLM-L6-v2")
class Trace(BaseEventHandler):
events: list = []
def handle(self, event, **kwargs):
if isinstance(event, EmbeddingEndEvent):
self.events.append(f"embedded {len(event.embeddings)} text(s)")
elif isinstance(event, RetrievalEndEvent):
names = [n.metadata["file_name"] for n in event.nodes]
self.events.append(f"retrieved {names}")
elif isinstance(event, QueryEndEvent):
self.events.append("query finished")
trace = Trace()
get_dispatcher().add_event_handler(trace)
docs = SimpleDirectoryReader("help").load_data()
index = VectorStoreIndex.from_documents(docs)
engine = index.as_query_engine(llm=ExtractiveLLM(), similarity_top_k=2)
trace.events.clear()
answer = engine.query("How long until my refund money reaches my card?")
print("answer:", answer)
print("captured events:")
for line in trace.events:
print(" -", line)answer: The money goes back to the card you paid with within 5 working days of us receiving the item. captured events: - embedded 1 text(s) - retrieved ['refunds.md', 'delivery.md'] - query finished
What the captured events tell you
- The embed event shows the question was turned into one vector before the search.
- The retrieval event names the files that were pulled, which is where most wrong answers are decided.
- The query event marks the end of the run, after the answer was built.
- Change the question and the retrieved files change, so the trace reflects the real run.
Prints vs an event handler
| Approach | Covers | Cost |
|---|---|---|
| Scattered prints | Only where you added them | Edits across the code, removed later |
| Event handler | Every embed, retrieve and query | One class, attached once |
When you reach for tracing
- An answer is wrong and you need to see which chunks were retrieved.
- Counting retrievals or embeddings to find a slow or costly step.
- Sending spans to a hosted tracing tool through a provided handler.
Related
- Previous: Multi-agent: routing between two agents
- Next: Persisting an index: not embedding twice
- See also: Query engines: the prompt the model receives
- Reference: Observability
- Add a handler branch for
EmbeddingEndEventand print how many texts were embedded when the index is built. - Ask a lamp question and read which file the trace shows.
- Remove
trace.events.clear()and run two queries to see the list grow.
Every expert started right here.