Answers with sources: a stand-in model and citations
A source citation is the list of chunks an answer was built from, kept on the response so a reader can check where each fact came from.
Last updated: 28 Sep, 2026 · LlamaIndex 0.14
Query engines: the prompt the model receives showed the prompt a model receives. This lesson answers it without an API key, using a small stand-in model that picks its answer straight from the retrieved context, and reads back the chunks it used.
Reading the context out of the prompt
A LlamaIndex model is a CustomLLM subclass with metadata and complete. This one starts by naming itself; the answering logic follows in the next piece.
import re
from llama_index.core.llms import CompletionResponse, CustomLLM, LLMMetadata
from llama_index.core.llms.callbacks import llm_completion_callback
def stems(text):
"""Words longer than three letters, cut to five letters, so refund and refunds match."""
return {w[:5] for w in re.findall(r"[a-z0-9-]+", text.lower()) if len(w) > 3}
class ExtractiveLLM(CustomLLM):
"""Answers with the context sentence that shares most words with the question."""
@property
def metadata(self):
return LLMMetadata(model_name="extractive")Answering with the closest sentence, or refusing
It splits the context into sentences, picks the one sharing the most words with the question, and compares words by their first five letters so refund matches refunds. Fewer than two shared words means it says it could not find an answer. A real model writes a better sentence; the flow around it is the same.
@llm_completion_callback()
def complete(self, prompt, formatted=False, **kwargs):
context = prompt.split("---------------------")[1]
question = prompt.split("Query:")[1].split("Answer:")[0]
asked = stems(question)
sentences = [s.strip() for s in re.split(r"(?<=[.!?])\s+|\n+", context)]
sentences = [s for s in sentences if s and not s.startswith("#") and ": " not in s[:20]]
best = max(sentences, key=lambda s: len(asked & stems(s)), default="")
if len(asked & stems(best)) < 2:
return CompletionResponse(text="I could not find that in the documents.")
return CompletionResponse(text=best)
@llm_completion_callback()
def stream_complete(self, prompt, formatted=False, **kwargs):
yield self.complete(prompt)Passing the model to the query engine
from extractive_llm import ExtractiveLLM
query_engine = index.as_query_engine(llm=ExtractiveLLM(), similarity_top_k=2)View the code here
# Lamps
The LMP-204 desk lamp has a known cable fault. Stop using a lamp with a damaged cable and we will replace it free of charge.
All lamps come with a two year guarantee against electrical faults.
Bulbs are not covered by the refund policy once they have been used.
The LMP-310 floor lamp needs a bulb with an E27 fitting, which is sold separately.
# Refunds
You can get a full refund within 30 days of delivery. The money goes back to the card you paid with within 5 working days of us receiving the item.
Items bought in a sale can be refunded too, but the delivery charge is not returned.
To start a refund, open the order in your account and choose Return an item. Print the label and drop the parcel at any post office.
Personalised items cannot be refunded unless they arrive damaged.
# Delivery
Standard delivery takes 3 to 5 working days and is free on orders over 40.
Express delivery arrives the next working day if you order before 2pm. It costs 6.
We deliver to the mainland only. Parcels to islands take 2 extra working days.
If a parcel has not arrived after 10 working days, contact us and we will send a replacement.
Answering two questions with citations
from extractive_llm import ExtractiveLLM
query_engine = index.as_query_engine(llm=ExtractiveLLM(), similarity_top_k=2)
for question in ["How long until my refund money reaches my card?", "Is the LMP-204 lamp safe?"]:
response = query_engine.query(question)
if str(response) == "I could not find that in the documents.":
print(response)
continue
sources = sorted({node.metadata["file_name"] for node in response.source_nodes})
print(f"{response} [{', '.join(sources)}]")The money goes back to the card you paid with within 5 working days of us receiving the item. [delivery.md, refunds.md] The LMP-204 desk lamp has a known cable fault. [lamps.md]
Reading the answer and its sources
- The answer tracks the question. The refund-timing question returns the five-working-days sentence; the lamp question returns the cable-fault sentence.
- source_nodes holds the retrieved chunks. Listing their file names after the answer is a citation a customer or an agent can follow.
- A citation says what was retrieved, not what was used. Two chunks came back per question; one sentence answered it. Showing every source reflects what the model saw.
Sources retrieved vs sources used
| Retrieved sources | Used source | |
|---|---|---|
| Where from | response.source_nodes | The one sentence the model quoted |
| Count here | Two per question | One |
| Honesty | Shows all the model saw | Needs the model to say which it used |
When to show sources
- A customer or support agent needs to verify an answer against the policy it came from.
- You are debugging a wrong answer and want to see which chunks fed it.
- An audit requires every answer to name the document behind it.
Related
- Previous: Query engines: the prompt the model receives
- Next: Refusing when nothing fits: similarity cutoffs
- Ask
"What are your opening hours?". - Print each source node's score next to its file name.
- Set
Settings.llm = ExtractiveLLM()once instead of passingllm=.
Little by little, you're building something great.