LlamaIndex project: a help-centre assistant
The capstone is a runnable project that puts the whole course together: a help-centre assistant that answers from the documents, cites them, respects who is asking, and refuses when nothing fits.
Last updated: 28 Sep, 2026 · LlamaIndex 0.14
Every piece was built in an earlier lesson. Here they come together: local embeddings, the stand-in answering model, per-audience metadata, a permission filter, a similarity cutoff, and citations.
Setting the models and the index
Local embeddings and the extractive stand-in model are set once as defaults. The files are loaded from both folders, each marked with its audience, which is kept out of the embeddings and the prompt.
Settings.embed_model = HuggingFaceEmbedding(model_name="sentence-transformers/all-MiniLM-L6-v2")
Settings.llm = ExtractiveLLM()
documents = SimpleDirectoryReader(".", recursive=True, required_exts=[".md"],
file_metadata=with_audience).load_data()
for document in documents:
document.excluded_embed_metadata_keys = ["audience"]
document.excluded_llm_metadata_keys = ["audience"]
index = VectorStoreIndex.from_documents(documents)Filtering by who is asking
A role maps to the audiences it may read: staff see everything, customers only customer documents. The filter lists each allowed audience and joins them with or.
ALLOWED = {"customer": ["customer"], "staff": ["customer", "staff"]}
filters = MetadataFilters(
filters=[MetadataFilter(key="audience", value=a, operator=FilterOperator.EQ)
for a in ALLOWED[role]],
condition="or",
)Refusing, or answering with sources
The cutoff turns an empty context into a handover to a person. When there are sources, they become a citation on the answer.
response = engine.query(question)
if not response.source_nodes or str(response) == "I could not find that in the documents.":
return "I can't find that in our help centre. A person from the team will reply."
sources = sorted({n.metadata["file_name"] for n in response.source_nodes})
return f"{response} (sources: {', '.join(sources)})"View the code here
# Lamps
The LMP-204 desk lamp has a known cable fault. Stop using a lamp with a damaged cable and we will replace it free of charge.
All lamps come with a two year guarantee against electrical faults.
Bulbs are not covered by the refund policy once they have been used.
The LMP-310 floor lamp needs a bulb with an E27 fitting, which is sold separately.
# Refunds
You can get a full refund within 30 days of delivery. The money goes back to the card you paid with within 5 working days of us receiving the item.
Items bought in a sale can be refunded too, but the delivery charge is not returned.
To start a refund, open the order in your account and choose Return an item. Print the label and drop the parcel at any post office.
Personalised items cannot be refunded unless they arrive damaged.
# Delivery
Standard delivery takes 3 to 5 working days and is free on orders over 40.
Express delivery arrives the next working day if you order before 2pm. It costs 6.
We deliver to the mainland only. Parcels to islands take 2 extra working days.
If a parcel has not arrived after 10 working days, contact us and we will send a replacement.
import re
from llama_index.core.llms import CompletionResponse, CustomLLM, LLMMetadata
from llama_index.core.llms.callbacks import llm_completion_callback
def stems(text):
"""Words longer than three letters, cut to five letters, so refund and refunds match."""
return {w[:5] for w in re.findall(r"[a-z0-9-]+", text.lower()) if len(w) > 3}
class ExtractiveLLM(CustomLLM):
"""Answers with the context sentence that shares most words with the question."""
@property
def metadata(self):
return LLMMetadata(model_name="extractive")
@llm_completion_callback()
def complete(self, prompt, formatted=False, **kwargs):
context = prompt.split("---------------------")[1]
question = prompt.split("Query:")[1].split("Answer:")[0]
asked = stems(question)
sentences = [s.strip() for s in re.split(r"(?<=[.!?])\s+|\n+", context)]
sentences = [s for s in sentences if s and not s.startswith("#") and ": " not in s[:20]]
best = max(sentences, key=lambda s: len(asked & stems(s)), default="")
if len(asked & stems(best)) < 2:
return CompletionResponse(text="I could not find that in the documents.")
return CompletionResponse(text=best)
@llm_completion_callback()
def stream_complete(self, prompt, formatted=False, **kwargs):
yield self.complete(prompt)
# Refund approvals
Refunds over 200 need a team lead's approval before they are paid.
A refund on a personalised item always needs a team lead to check the damage photos first.
Four questions from two roles
The whole program answers the same and different questions for a customer and a staff member. Two of the four are refusals, which is the point: the assistant does not invent.
import os
from llama_index.core import Settings, SimpleDirectoryReader, VectorStoreIndex
from llama_index.core.postprocessor import SimilarityPostprocessor
from llama_index.core.vector_stores import FilterOperator, MetadataFilter, MetadataFilters
from llama_index.embeddings.huggingface import HuggingFaceEmbedding
from extractive_llm import ExtractiveLLM
Settings.embed_model = HuggingFaceEmbedding(model_name="sentence-transformers/all-MiniLM-L6-v2")
Settings.llm = ExtractiveLLM()
def with_audience(path):
audience = "staff" if "/staff/" in path else "customer"
return {"file_name": os.path.basename(path), "audience": audience}
documents = SimpleDirectoryReader(".", recursive=True, required_exts=[".md"], file_metadata=with_audience).load_data()
for document in documents:
document.excluded_embed_metadata_keys = ["audience"]
document.excluded_llm_metadata_keys = ["audience"]
index = VectorStoreIndex.from_documents(documents)
ALLOWED = {"customer": ["customer"], "staff": ["customer", "staff"]}
def ask(question, role):
filters = MetadataFilters(
filters=[MetadataFilter(key="audience", value=a, operator=FilterOperator.EQ) for a in ALLOWED[role]],
condition="or",
)
engine = index.as_query_engine(
similarity_top_k=2,
filters=filters,
node_postprocessors=[SimilarityPostprocessor(similarity_cutoff=0.3)],
)
response = engine.query(question)
if not response.source_nodes or str(response) == "I could not find that in the documents.":
return "I can't find that in our help centre. A person from the team will reply."
sources = sorted({n.metadata["file_name"] for n in response.source_nodes})
return f"{response} (sources: {', '.join(sources)})"
for question, role in [
("How long until my refund money reaches my card?", "customer"),
("Who approves a refund over 200?", "customer"),
("Who approves a refund over 200?", "staff"),
("What are your opening hours?", "customer"),
]:
print(f"[{role}] {question}")
print(" ", ask(question, role))[customer] How long until my refund money reaches my card?
The money goes back to the card you paid with within 5 working days of us receiving the item. (sources: delivery.md, refunds.md)
[customer] Who approves a refund over 200?
I can't find that in our help centre. A person from the team will reply.
[staff] Who approves a refund over 200?
Refunds over 200 need a team lead's approval before they are paid. (sources: refund-approvals.md, refunds.md)
[customer] What are your opening hours?
I can't find that in our help centre. A person from the team will reply.Reading the four answers, including the refusals
- The refund question is answered from the refund file and cited, the happy path.
- The customer's approval question never sees the staff file, so the model finds no matching sentence and says it could not: a real no-answer-found result.
- The staff member gets the approval rule, because the filter offered the staff file.
- The opening-hours question has nothing above the cutoff, so it hands over to a person instead of guessing.
The failure modes this assistant shows
| Situation | What happens | Why |
|---|---|---|
| Answer is in an allowed file | Answered and cited | Retrieval and the stand-in model agree |
| Answer is in a file the role may not see | "I could not find that" | The filter removed the only source |
| No file is close enough | Handover to a person | The similarity cutoff emptied the context |
Pick one to watch it run, step by step.
What we left out
| Topic | What it is for | Where to read |
|---|---|---|
| Response synthesizers | compact, refine and tree_summarize: how many model calls build one answer | Response synthesizers |
| Router query engine | Sending a question to one of several engines | Router query engine |
| Response-quality evals | Scoring answers for faithfulness and relevancy | Evaluating responses |
| Other index types | SummaryIndex and others beyond the vector index | Indexing |
| Agent human-in-the-loop | Pausing an agent for a person to approve an action | Human in the loop |
| Streaming | Sending an answer token by token as it is built | Streaming output |
| Workflow retry | Re-running a failed step with a retry policy | Workflows |
audience on a file leaks it to everyone allowed the fallback. Check the metadata on every file before you trust the filter.Related
- Previous: Integrations: swapping in real backends
- See also: Refusing when nothing fits: similarity cutoffs
- Reference: Putting it all together
- Use the hybrid retriever and the reranker inside
ask, and rerun the hit rate on it. - Persist the index and refresh it when files change.
- Replace
ExtractiveLLMwith a real model through the Groq integration and compare the answers.
This is what real progress feels like.