LlamaIndexllama-index-core 0.14 · Python 3.10+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
30 small wins to finish your pathNext lesson →

Structured output: a typed answer with Pydantic

Structured output is a query mode that returns a typed object instead of a paragraph, so your code can read named fields rather than parse prose.

Last updated: 28 Sep, 2026 · LlamaIndex 0.14

Every answer so far has been text. When an answer feeds other code, a refund window or a delivery estimate, you want fields you can read. Passing an output_cls to the query engine makes it return an instance of your class.

Describing the answer as a Pydantic model

The schema is a BaseModel. Each field has a type and a description; the description becomes part of the instruction the model sees, so name it well.

python
from pydantic import BaseModel, Field


class PolicyAnswer(BaseModel):
    """A one-sentence answer with any whole numbers it mentions."""

    summary: str = Field(description="the sentence that answers the question")
    numbers: list[int] = Field(description="every whole number named in that sentence")

Emitting JSON from the context

LlamaIndex appends the schema to the prompt and parses the model's reply as JSON into the class. A real model writes that JSON; here a stand-in reads the best context sentence and returns it as JSON, so the demo runs with no key.

python
def complete(self, prompt, formatted=False, **kwargs):
    context = prompt.split("---------------------")[1]
    question = prompt.split("Query:")[1].split("Answer:")[0]
    asked = stems(question)
    sentences = [s.strip() for s in re.split(r"(?<=[.!?])\s+|\n+", context)]
    sentences = [s for s in sentences if s and not s.startswith("#")]
    best = max(sentences, key=lambda s: len(asked & stems(s)), default="")
    numbers = [int(n) for n in re.findall(r"\d+", best)]
    return CompletionResponse(text=json.dumps({"summary": best, "numbers": numbers}))
Project files used on this pageThis lesson builds on a project from earlier lessons. The code below imports these files. Click a file to see its code, or follow the link to the lesson that wrote it. To run the code yourself, keep them in the same folder.
View the code here
help/lamps.md
# Lamps

The LMP-204 desk lamp has a known cable fault. Stop using a lamp with a damaged cable and we will replace it free of charge.

All lamps come with a two year guarantee against electrical faults.

Bulbs are not covered by the refund policy once they have been used.

The LMP-310 floor lamp needs a bulb with an E27 fitting, which is sold separately.
help/refunds.md
# Refunds

You can get a full refund within 30 days of delivery. The money goes back to the card you paid with within 5 working days of us receiving the item.

Items bought in a sale can be refunded too, but the delivery charge is not returned.

To start a refund, open the order in your account and choose Return an item. Print the label and drop the parcel at any post office.

Personalised items cannot be refunded unless they arrive damaged.
help/delivery.md
# Delivery

Standard delivery takes 3 to 5 working days and is free on orders over 40.

Express delivery arrives the next working day if you order before 2pm. It costs 6.

We deliver to the mainland only. Parcels to islands take 2 extra working days.

If a parcel has not arrived after 10 working days, contact us and we will send a replacement.
json_facts_llm.py
import json
import re

from llama_index.core.llms import CompletionResponse, CustomLLM, LLMMetadata
from llama_index.core.llms.callbacks import llm_completion_callback
from pydantic import BaseModel, Field


class PolicyAnswer(BaseModel):
    """A one-sentence answer with any whole numbers it mentions."""

    summary: str = Field(description="the sentence that answers the question")
    numbers: list[int] = Field(description="every whole number named in that sentence")


def stems(text):
    return {w[:5] for w in re.findall(r"[a-z0-9-]+", text.lower()) if len(w) > 3}


class JSONFactsLLM(CustomLLM):
    """Returns the best context sentence as a JSON object, not a paragraph."""

    @property
    def metadata(self):
        return LLMMetadata(model_name="json-facts")

    @llm_completion_callback()
    def complete(self, prompt, formatted=False, **kwargs):
        context = prompt.split("---------------------")[1]
        question = prompt.split("Query:")[1].split("Answer:")[0]
        asked = stems(question)
        sentences = [s.strip() for s in re.split(r"(?<=[.!?])\s+|\n+", context)]
        sentences = [s for s in sentences if s and not s.startswith("#")]
        best = max(sentences, key=lambda s: len(asked & stems(s)), default="")
        numbers = [int(n) for n in re.findall(r"\d+", best)]
        return CompletionResponse(text=json.dumps({"summary": best, "numbers": numbers}))

    @llm_completion_callback()
    def stream_complete(self, prompt, formatted=False, **kwargs):
        yield self.complete(prompt)

Asking the engine for the typed class

Pass the class as output_cls. The response's .response is then an instance of PolicyAnswer, not a string.

python
from json_facts_llm import JSONFactsLLM, PolicyAnswer

engine = index.as_query_engine(llm=JSONFactsLLM(), similarity_top_k=2, output_cls=PolicyAnswer)

A typed answer for two questions

Example
from json_facts_llm import JSONFactsLLM, PolicyAnswer

engine = index.as_query_engine(llm=JSONFactsLLM(), similarity_top_k=2, output_cls=PolicyAnswer)
for question in ["How long until my refund money reaches my card?", "How long does standard delivery take?"]:
    answer = engine.query(question)
    print(type(answer.response).__name__, "->", answer.response.numbers, "|", answer.response.summary)

Reading the typed objects

  • Each answer is a PolicyAnswer. The engine parsed the model's JSON into the class, so answer.response.numbers is a real list of ints.
  • The object tracks the question. The refund question yields [5]; the delivery question yields [3, 5, 40]. Different context, different fields.
  • No string parsing on your side. Reading .numbers or .summary replaces hunting for figures in a paragraph.

A paragraph vs a typed object

Plain queryoutput_cls query
Return typeA stringA PolicyAnswer instance
Reading a fieldParse the text yourselfanswer.response.numbers
If a field is missingSilentParsing raises, so you see it

When to ask for structured output

  • The answer feeds other code, a UI field, a database row, or a rule.
  • You need specific values pulled out, dates, amounts, yes or no, not a sentence.
  • You want the shape of the answer validated, so a missing field is caught early.
Watch out
Structured output only works when the model returns JSON matching the schema. A model that cannot follow the format makes parsing raise; keep the field descriptions clear and the schema small, and handle the parse error rather than assuming a clean object.
Try it yourself
  • Add a refusal: bool field and set it in the stand-in when no sentence matches.
  • Ask "What are your opening hours?" and read what the object holds.
  • Change a field description and see it appear in the prompt the model is sent.

This is what real progress feels like.