LlamaIndexllama-index-core 0.14 · Python 3.10+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
30 small wins to finish your pathNext lesson →

Refusing when nothing fits: similarity cutoffs

A similarity cutoff is a threshold that drops retrieved chunks scoring below it, so an assistant can refuse a question no document answers.

Last updated: 28 Sep, 2026 · LlamaIndex 0.14

Answers with sources: a stand-in model and citations answered from whatever was retrieved. But top-k retrieval always returns k chunks, even for a question your documents do not cover. First see the problem, then fix it.

Retrieving for a question with no answer

Example
retriever = index.as_retriever(similarity_top_k=2)
for question in ["How long does standard delivery take?", "What are your opening hours?"]:
    print(question, [(n.metadata["file_name"], round(n.score, 2)) for n in retriever.retrieve(question)])

No document mentions opening hours, yet two chunks came back for that question, with low scores. A model given those chunks might refuse, or might invent an answer from the nearest thing it was handed.

Adding a similarity cutoff

A node postprocessor runs between retrieval and the model. SimilarityPostprocessor drops nodes scoring below the cutoff before they ever reach the prompt.

python
from llama_index.core.postprocessor import SimilarityPostprocessor

cutoff = SimilarityPostprocessor(similarity_cutoff=0.3)
Project files used on this pageThis lesson builds on a project from earlier lessons. The code below imports these files. Click a file to see its code, or follow the link to the lesson that wrote it. To run the code yourself, keep them in the same folder.
View the code here
help/lamps.md
# Lamps

The LMP-204 desk lamp has a known cable fault. Stop using a lamp with a damaged cable and we will replace it free of charge.

All lamps come with a two year guarantee against electrical faults.

Bulbs are not covered by the refund policy once they have been used.

The LMP-310 floor lamp needs a bulb with an E27 fitting, which is sold separately.
help/refunds.md
# Refunds

You can get a full refund within 30 days of delivery. The money goes back to the card you paid with within 5 working days of us receiving the item.

Items bought in a sale can be refunded too, but the delivery charge is not returned.

To start a refund, open the order in your account and choose Return an item. Print the label and drop the parcel at any post office.

Personalised items cannot be refunded unless they arrive damaged.
help/delivery.md
# Delivery

Standard delivery takes 3 to 5 working days and is free on orders over 40.

Express delivery arrives the next working day if you order before 2pm. It costs 6.

We deliver to the mainland only. Parcels to islands take 2 extra working days.

If a parcel has not arrived after 10 working days, contact us and we will send a replacement.
extractive_llm.py
import re

from llama_index.core.llms import CompletionResponse, CustomLLM, LLMMetadata
from llama_index.core.llms.callbacks import llm_completion_callback


def stems(text):
    """Words longer than three letters, cut to five letters, so refund and refunds match."""
    return {w[:5] for w in re.findall(r"[a-z0-9-]+", text.lower()) if len(w) > 3}


class ExtractiveLLM(CustomLLM):
    """Answers with the context sentence that shares most words with the question."""

    @property
    def metadata(self):
        return LLMMetadata(model_name="extractive")

    @llm_completion_callback()
    def complete(self, prompt, formatted=False, **kwargs):
        context = prompt.split("---------------------")[1]
        question = prompt.split("Query:")[1].split("Answer:")[0]
        asked = stems(question)
        sentences = [s.strip() for s in re.split(r"(?<=[.!?])\s+|\n+", context)]
        sentences = [s for s in sentences if s and not s.startswith("#") and ": " not in s[:20]]
        best = max(sentences, key=lambda s: len(asked & stems(s)), default="")
        if len(asked & stems(best)) < 2:
            return CompletionResponse(text="I could not find that in the documents.")
        return CompletionResponse(text=best)

    @llm_completion_callback()
    def stream_complete(self, prompt, formatted=False, **kwargs):
        yield self.complete(prompt)

Refusing when every chunk is below the cutoff

Example
from llama_index.core.postprocessor import SimilarityPostprocessor
from extractive_llm import ExtractiveLLM

query_engine = index.as_query_engine(
    llm=ExtractiveLLM(),
    similarity_top_k=2,
    node_postprocessors=[SimilarityPostprocessor(similarity_cutoff=0.3)],
)
for question in ["How long does standard delivery take?", "What are your opening hours?"]:
    response = query_engine.query(question)
    print(question, "->", response, f"({len(response.source_nodes)} sources)")

Reading the two responses

  • The delivery question passes. Its top chunk scores above 0.3, so the answer comes through with two sources.
  • The opening-hours question is emptied. Nothing scored above 0.3, so LlamaIndex returns Empty Response without calling the model at all.
  • Empty Response is your refusal signal. When response.source_nodes is empty, your application can show a clear could-not-find message.

No cutoff vs a cutoff

No cutoffWith a cutoff
Weak chunksSent to the modelDropped before the prompt
Unanswerable questionAn invented answer is possibleEmpty Response, a clean refusal
CostA model call every timeNo call when nothing qualifies

When to set a cutoff

  • The assistant must refuse rather than guess when a topic is outside the documents.
  • You want to skip the model call, and its cost, for questions nothing answers.
  • You are tuning how strict retrieval is and want a single number to move.
Watch out
The cutoff is a number to measure, not to guess. Scores depend on the embedding model and the wording of questions; too high refuses good answers, too low refuses nothing. Measuring retrieval: hit rate on labelled questions tests it on labelled questions.
Try it yourself
  • Try cutoffs of 0.1 and 0.5 on both questions.
  • Replace Empty Response with your own refusal message when response.source_nodes is empty.
  • Ask "Can I return a used bulb?" with the cutoff.

You understood something today that you didn't yesterday.