Hybrid search: keywords and meaning together
Hybrid search is a retriever that runs keyword and semantic search together and merges their rankings, so a question is served whichever way it is phrased.
Last updated: 28 Sep, 2026 · LlamaIndex 0.14
BM25: when exact words matter showed each method failing where the other works. Hybrid search runs both and fuses the results, so a part number and a reworded question both land on the right file.
Fusing the two retrievers
QueryFusionRetriever runs several retrievers and merges their results. num_queries=1 uses the question as given; above 1 it asks a model to write extra versions, which is a model call. It looks a model up when created even at 1, so MockLLM fills the slot without a key. reciprocal_rerank merges by rank, not raw score.
from llama_index.core.llms import MockLLM
from llama_index.core.retrievers import QueryFusionRetriever
hybrid = QueryFusionRetriever(
[semantic, keyword], llm=MockLLM(),
similarity_top_k=2, num_queries=1,
mode="reciprocal_rerank", use_async=False,
)View the code here
# Lamps
The LMP-204 desk lamp has a known cable fault. Stop using a lamp with a damaged cable and we will replace it free of charge.
All lamps come with a two year guarantee against electrical faults.
Bulbs are not covered by the refund policy once they have been used.
The LMP-310 floor lamp needs a bulb with an E27 fitting, which is sold separately.
# Refunds
You can get a full refund within 30 days of delivery. The money goes back to the card you paid with within 5 working days of us receiving the item.
Items bought in a sale can be refunded too, but the delivery charge is not returned.
To start a refund, open the order in your account and choose Return an item. Print the label and drop the parcel at any post office.
Personalised items cannot be refunded unless they arrive damaged.
# Delivery
Standard delivery takes 3 to 5 working days and is free on orders over 40.
Express delivery arrives the next working day if you order before 2pm. It costs 6.
We deliver to the mainland only. Parcels to islands take 2 extra working days.
If a parcel has not arrived after 10 working days, contact us and we will send a replacement.
Fusing results for three questions
from llama_index.core.llms import MockLLM
from llama_index.core.retrievers import QueryFusionRetriever
hybrid = QueryFusionRetriever(
[semantic, keyword],
llm=MockLLM(),
similarity_top_k=2,
num_queries=1,
mode="reciprocal_rerank",
use_async=False,
)
for question in ["LMP-204", "Is reimbursement possible?", "Is the LMP-204 refundable?"]:
print(question, [(n.metadata["file_name"], round(n.score, 4)) for n in hybrid.retrieve(question)])LMP-204 [('lamps.md', 0.0333), ('lamps.md', 0.0328)]
Is reimbursement possible? [('refunds.md', 0.0331), ('refunds.md', 0.0167)]
Is the LMP-204 refundable? [('lamps.md', 0.0333), ('lamps.md', 0.0328)]Reading the fused results
- Both kinds of question land right. The part number and the reworded refund question each put the correct file first.
- Fused scores are small rank numbers. They order results but are not similarities, so a cutoff like the one in Refusing when nothing fits: similarity cutoffs does not apply to them.
- A chunk found by only one retriever still scores. Reciprocal rank fusion adds each retriever's contribution, so a single strong finder is enough to rank a chunk.
Single retriever vs hybrid
| One retriever | Hybrid fusion | |
|---|---|---|
| Coverage | Strong one way only | Both codes and meaning |
| Scores | Similarity or BM25 | Small rank-fusion numbers |
| Cost | One search | Two searches, then a merge |
When to combine both retrievers
- Questions mix exact terms and natural phrasing, like a SKU in one sentence and a description in the next.
- One retriever alone keeps missing a class of question you can name.
- You want a single retriever to hand to a query engine that covers both.
num_queries above 1 asks the model to rewrite the question, which is a real model call and needs a working LLM. Keep it at 1 for a keyless run, and do not treat the tiny fused scores as similarities.Related
- Swap the order of
[semantic, keyword]. Does the result change? - Set
similarity_top_k=3on the fusion retriever. - Use
hybridinside a query engine withRetrieverQueryEngine.from_args(hybrid, llm=ExtractiveLLM()).
You understood something today that you didn't yesterday.