LlamaIndexllama-index-core 0.14 · Python 3.10+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
30 small wins to finish your pathNext lesson →

Measuring retrieval: hit rate on labelled questions

Hit rate is a retrieval metric: the share of labelled questions whose correct document appears anywhere in the retrieved results.

Last updated: 28 Sep, 2026 · LlamaIndex 0.14

Chunk size, top-k, keyword or hybrid: every choice so far has been a judgement call. Labelled questions turn each into a number, so you compare retrievers instead of guessing between them.

Labelling questions with their answer file

Each question is paired with the file that should answer it. Several are deliberately hard: no shared words, a part number, a bulb question that also mentions refunds.

python
questions = {
    "How do I get my money back?": "refunds.md",
    "Is the refund paid to my card?": "refunds.md",
    "My parcel still has not come": "delivery.md",
    "Can I get it tomorrow?": "delivery.md",
    "LMP-204": "lamps.md",
    "Are used bulbs refundable?": "lamps.md",
}

Counting hits over the set

For each question, retrieve and check whether the expected file is among the results. Hit rate is the share that pass. It is one of the metrics LlamaIndex's evaluation guide describes, alongside MRR, which also rewards the right result being first.

python
def hit_rate(retriever):
    hits = 0
    for question, expected in questions.items():
        found = [n.metadata["file_name"] for n in retriever.retrieve(question)]
        hits += expected in found
    return hits / len(questions)
Project files used on this pageThis lesson builds on a project from earlier lessons. The code below imports these files. Click a file to see its code, or follow the link to the lesson that wrote it. To run the code yourself, keep them in the same folder.
View the code here
help/lamps.md
# Lamps

The LMP-204 desk lamp has a known cable fault. Stop using a lamp with a damaged cable and we will replace it free of charge.

All lamps come with a two year guarantee against electrical faults.

Bulbs are not covered by the refund policy once they have been used.

The LMP-310 floor lamp needs a bulb with an E27 fitting, which is sold separately.
help/refunds.md
# Refunds

You can get a full refund within 30 days of delivery. The money goes back to the card you paid with within 5 working days of us receiving the item.

Items bought in a sale can be refunded too, but the delivery charge is not returned.

To start a refund, open the order in your account and choose Return an item. Print the label and drop the parcel at any post office.

Personalised items cannot be refunded unless they arrive damaged.
help/delivery.md
# Delivery

Standard delivery takes 3 to 5 working days and is free on orders over 40.

Express delivery arrives the next working day if you order before 2pm. It costs 6.

We deliver to the mainland only. Parcels to islands take 2 extra working days.

If a parcel has not arrived after 10 working days, contact us and we will send a replacement.

Scoring three retrievers at two depths

Example
from llama_index.core.llms import MockLLM
from llama_index.core.retrievers import QueryFusionRetriever

for k in (1, 2):
    semantic = index.as_retriever(similarity_top_k=k)
    keyword = BM25Retriever.from_defaults(nodes=nodes, similarity_top_k=k)
    hybrid = QueryFusionRetriever([semantic, keyword], llm=MockLLM(), similarity_top_k=k, num_queries=1, mode="reciprocal_rerank", use_async=False)
    print(f"top {k}:  semantic {hit_rate(semantic):.0%}  keyword {hit_rate(keyword):.0%}  hybrid {hit_rate(hybrid):.0%}")

Reading the hit rates

  • Six numbers replace six opinions. Each retriever now has a score you can compare and track across changes.
  • Keyword misses one at top 1. A reworded question with no shared words is missed when only one chunk is returned, and caught at top 2.
  • Semantic and hybrid find every file either way. With three short files the test is easy; the same function runs unchanged on a real set, where the numbers spread apart.

Guessing vs measuring

GuessingMeasuring
BasisIt feels betterA hit rate on labelled questions
Comparing optionsOpinionTwo numbers side by side
Catching a regressionBy luckThe score drops

When to measure retrieval

  • You are choosing between chunk sizes, top-k values or retrievers and want evidence.
  • A change to the index needs a check that retrieval did not get worse.
  • You are setting a cutoff and need to see what it refuses on real questions.
Watch out
Six questions demonstrate the method but are far too few to trust. A real test set has dozens to hundreds of questions from real users; LlamaIndex's RetrieverEvaluator computes hit rate and MRR over such a set.
Try it yourself
  • Add three questions of your own, including one no document answers, and decide how to score it.
  • Rebuild nodes with chunk_size=40 and rerun the comparison.
  • Write mrr(retriever): add 1 divided by the position of the right file, or 0 if it is missing.

This is what real progress feels like.