1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
30 small wins to finish your pathNext lesson →
Query engines: the prompt the model receives
A query engine is the piece that retrieves chunks, writes them into a prompt with the question, and sends that prompt to a model.
Last updated: 28 Sep, 2026 · LlamaIndex 0.14
The retriever from Retrievers: the search step on its own gives back chunks. A query engine takes the next step and asks a model. To see exactly what the model would read, swap in a stand-in that prints the prompt instead of answering.
Setting a mock model to see the prompt
MockLLM is a stand-in included in LlamaIndex. Given no max_tokens, its complete returns the prompt it was handed, so the printed text is exactly what a real model would receive.
from llama_index.core.llms import MockLLM
Settings.llm = MockLLM() # returns the prompt it was given, not an answer
query_engine = index.as_query_engine(similarity_top_k=1)Project files used on this pageThis lesson builds on a project from earlier lessons. The code below imports these files. Click a file to see its code, or follow the link to the lesson that wrote it. To run the code yourself, keep them in the same folder.
View the code here
help/lamps.md
# Lamps
The LMP-204 desk lamp has a known cable fault. Stop using a lamp with a damaged cable and we will replace it free of charge.
All lamps come with a two year guarantee against electrical faults.
Bulbs are not covered by the refund policy once they have been used.
The LMP-310 floor lamp needs a bulb with an E27 fitting, which is sold separately.
help/refunds.md
# Refunds
You can get a full refund within 30 days of delivery. The money goes back to the card you paid with within 5 working days of us receiving the item.
Items bought in a sale can be refunded too, but the delivery charge is not returned.
To start a refund, open the order in your account and choose Return an item. Print the label and drop the parcel at any post office.
Personalised items cannot be refunded unless they arrive damaged.
help/delivery.md
# Delivery
Standard delivery takes 3 to 5 working days and is free on orders over 40.
Express delivery arrives the next working day if you order before 2pm. It costs 6.
We deliver to the mainland only. Parcels to islands take 2 extra working days.
If a parcel has not arrived after 10 working days, contact us and we will send a replacement.
Printing the prompt for a delivery question
from llama_index.core.llms import MockLLM
Settings.llm = MockLLM()
query_engine = index.as_query_engine(similarity_top_k=1)
print(query_engine.query("How long does standard delivery take?"))Output
Context information is below. --------------------- # Delivery Standard delivery takes 3 to 5 working days and is free on orders over 40. Express delivery arrives the next working day if you order before 2pm. It costs 6. We deliver to the mainland only. Parcels to islands take 2 extra working days. --------------------- Given the context information and not prior knowledge, answer the query. Query: How long does standard delivery take? Answer:
Reading the prompt
- The retrieved chunk sits between the dashed lines. That is the only knowledge of your documents the model has.
- No metadata line appears above it. The reader excludes
file_namefrom the model text by default, so citations come from the response, not the prompt. - An instruction follows the context. It tells the model to answer from the context and not from prior knowledge, then gives the question.
MockLLM vs a real model
| MockLLM() | A real model | |
|---|---|---|
| Answer | The prompt, printed back | A written answer from the context |
| Cost / key | None | Tokens, and an API key |
| Use | Seeing what the model is sent | Answering the customer |
When to print the prompt
- A real model gives an odd answer and you want to see the exact context it was given.
- You changed the chunk size or top-k and want to confirm the prompt shrank or grew.
- You are about to pay per token and want to count how much context each question sends.
Watch out
Do not pass
MockLLM(max_tokens=n) for this: with a token count it returns repeated filler text, not the prompt. The prompt echo needs a plain MockLLM().Related
- Previous: Retrievers: the search step on its own
- Next: Answers with sources: a stand-in model and citations
Try it yourself
- Set
similarity_top_k=3and count the chunks in the prompt. - Ask a question with no answer in the documents and read the prompt.
- Print
len(str(response))for top-k 1 and 3: a real model is paid for all of it.
Every expert started right here.