Structured output: a typed answer with Pydantic
Structured output is a query mode that returns a typed object instead of a paragraph, so your code can read named fields rather than parse prose.
Last updated: 28 Sep, 2026 · LlamaIndex 0.14
Every answer so far has been text. When an answer feeds other code, a refund window or a delivery estimate, you want fields you can read. Passing an output_cls to the query engine makes it return an instance of your class.
Describing the answer as a Pydantic model
The schema is a BaseModel. Each field has a type and a description; the description becomes part of the instruction the model sees, so name it well.
from pydantic import BaseModel, Field
class PolicyAnswer(BaseModel):
"""A one-sentence answer with any whole numbers it mentions."""
summary: str = Field(description="the sentence that answers the question")
numbers: list[int] = Field(description="every whole number named in that sentence")Emitting JSON from the context
LlamaIndex appends the schema to the prompt and parses the model's reply as JSON into the class. A real model writes that JSON; here a stand-in reads the best context sentence and returns it as JSON, so the demo runs with no key.
def complete(self, prompt, formatted=False, **kwargs):
context = prompt.split("---------------------")[1]
question = prompt.split("Query:")[1].split("Answer:")[0]
asked = stems(question)
sentences = [s.strip() for s in re.split(r"(?<=[.!?])\s+|\n+", context)]
sentences = [s for s in sentences if s and not s.startswith("#")]
best = max(sentences, key=lambda s: len(asked & stems(s)), default="")
numbers = [int(n) for n in re.findall(r"\d+", best)]
return CompletionResponse(text=json.dumps({"summary": best, "numbers": numbers}))Asking the engine for the typed class
Pass the class as output_cls. The response's .response is then an instance of PolicyAnswer, not a string.
from json_facts_llm import JSONFactsLLM, PolicyAnswer
engine = index.as_query_engine(llm=JSONFactsLLM(), similarity_top_k=2, output_cls=PolicyAnswer)A typed answer for two questions
from json_facts_llm import JSONFactsLLM, PolicyAnswer
engine = index.as_query_engine(llm=JSONFactsLLM(), similarity_top_k=2, output_cls=PolicyAnswer)
for question in ["How long until my refund money reaches my card?", "How long does standard delivery take?"]:
answer = engine.query(question)
print(type(answer.response).__name__, "->", answer.response.numbers, "|", answer.response.summary)Reading the typed objects
- Each answer is a PolicyAnswer. The engine parsed the model's JSON into the class, so
answer.response.numbersis a real list of ints. - The object tracks the question. The refund question yields
[5]; the delivery question yields[3, 5, 40]. Different context, different fields. - No string parsing on your side. Reading
.numbersor.summaryreplaces hunting for figures in a paragraph.
A paragraph vs a typed object
| Plain query | output_cls query | |
|---|---|---|
| Return type | A string | A PolicyAnswer instance |
| Reading a field | Parse the text yourself | answer.response.numbers |
| If a field is missing | Silent | Parsing raises, so you see it |
When to ask for structured output
- The answer feeds other code, a UI field, a database row, or a rule.
- You need specific values pulled out, dates, amounts, yes or no, not a sentence.
- You want the shape of the answer validated, so a missing field is caught early.
Related
- Previous: Metadata filters: who may see which documents
- Next: Chat engines: remembering the conversation
- Add a
refusal: boolfield and set it in the stand-in when no sentence matches. - Ask
"What are your opening hours?"and read what the object holds. - Change a field description and see it appear in the prompt the model is sent.
This is what real progress feels like.