Measuring retrieval: hit rate on labelled questions
Hit rate is a retrieval metric: the share of labelled questions whose correct document appears anywhere in the retrieved results.
Last updated: 28 Sep, 2026 · LlamaIndex 0.14
Chunk size, top-k, keyword or hybrid: every choice so far has been a judgement call. Labelled questions turn each into a number, so you compare retrievers instead of guessing between them.
Labelling questions with their answer file
Each question is paired with the file that should answer it. Several are deliberately hard: no shared words, a part number, a bulb question that also mentions refunds.
questions = {
"How do I get my money back?": "refunds.md",
"Is the refund paid to my card?": "refunds.md",
"My parcel still has not come": "delivery.md",
"Can I get it tomorrow?": "delivery.md",
"LMP-204": "lamps.md",
"Are used bulbs refundable?": "lamps.md",
}Counting hits over the set
For each question, retrieve and check whether the expected file is among the results. Hit rate is the share that pass. It is one of the metrics LlamaIndex's evaluation guide describes, alongside MRR, which also rewards the right result being first.
def hit_rate(retriever):
hits = 0
for question, expected in questions.items():
found = [n.metadata["file_name"] for n in retriever.retrieve(question)]
hits += expected in found
return hits / len(questions)View the code here
# Lamps
The LMP-204 desk lamp has a known cable fault. Stop using a lamp with a damaged cable and we will replace it free of charge.
All lamps come with a two year guarantee against electrical faults.
Bulbs are not covered by the refund policy once they have been used.
The LMP-310 floor lamp needs a bulb with an E27 fitting, which is sold separately.
# Refunds
You can get a full refund within 30 days of delivery. The money goes back to the card you paid with within 5 working days of us receiving the item.
Items bought in a sale can be refunded too, but the delivery charge is not returned.
To start a refund, open the order in your account and choose Return an item. Print the label and drop the parcel at any post office.
Personalised items cannot be refunded unless they arrive damaged.
# Delivery
Standard delivery takes 3 to 5 working days and is free on orders over 40.
Express delivery arrives the next working day if you order before 2pm. It costs 6.
We deliver to the mainland only. Parcels to islands take 2 extra working days.
If a parcel has not arrived after 10 working days, contact us and we will send a replacement.
Scoring three retrievers at two depths
from llama_index.core.llms import MockLLM
from llama_index.core.retrievers import QueryFusionRetriever
for k in (1, 2):
semantic = index.as_retriever(similarity_top_k=k)
keyword = BM25Retriever.from_defaults(nodes=nodes, similarity_top_k=k)
hybrid = QueryFusionRetriever([semantic, keyword], llm=MockLLM(), similarity_top_k=k, num_queries=1, mode="reciprocal_rerank", use_async=False)
print(f"top {k}: semantic {hit_rate(semantic):.0%} keyword {hit_rate(keyword):.0%} hybrid {hit_rate(hybrid):.0%}")top 1: semantic 100% keyword 83% hybrid 100% top 2: semantic 100% keyword 100% hybrid 100%
Reading the hit rates
- Six numbers replace six opinions. Each retriever now has a score you can compare and track across changes.
- Keyword misses one at top 1. A reworded question with no shared words is missed when only one chunk is returned, and caught at top 2.
- Semantic and hybrid find every file either way. With three short files the test is easy; the same function runs unchanged on a real set, where the numbers spread apart.
Guessing vs measuring
| Guessing | Measuring | |
|---|---|---|
| Basis | It feels better | A hit rate on labelled questions |
| Comparing options | Opinion | Two numbers side by side |
| Catching a regression | By luck | The score drops |
When to measure retrieval
- You are choosing between chunk sizes, top-k values or retrievers and want evidence.
- A change to the index needs a check that retrieval did not get worse.
- You are setting a cutoff and need to see what it refuses on real questions.
RetrieverEvaluator computes hit rate and MRR over such a set.Related
- Previous: Reranking: a second, closer look at the top results
- Next: FunctionAgent: an agent that queries your index
- Add three questions of your own, including one no document answers, and decide how to score it.
- Rebuild
nodeswithchunk_size=40and rerun the comparison. - Write
mrr(retriever): add 1 divided by the position of the right file, or 0 if it is missing.
This is what real progress feels like.