Structured output
Structured output means getting the model's answer in a fixed shape a program can read, such as JSON with named fields, and checking every answer against that shape before using it.
Last updated: 30 Sep, 2026 · groq 1.7 · gpt-oss-120b on Groq
A program that uses a model's answer cannot read it by eye. It has to check the shape every time, because the model can return anything. Pydantic, from Python for AI, does the checking.
The Triage model
from typing import Literal
from pydantic import BaseModel, Field, ValidationError
class Triage(BaseModel):
category: Literal["billing", "shipping", "other"]
priority: int = Field(ge=1, le=5)Literal allows only the three category names. Field(ge=1, le=5) allows priorities from 1 to 5.
A check function
def check(answer):
try:
return Triage.model_validate_json(answer)
except ValidationError:
return Nonecheck returns a Triage for an answer in the right shape and None for anything else.
Checking three answers by hand
print(check('{"category": "billing", "priority": 4}'))
print(check("**Category:** Billing"))
print(check('{"category": "refunds", "priority": 4}'))category='billing' priority=4 None None
- The first is valid and comes back as a
Triage. - The second is the vague prompt's Markdown: not JSON, so
None. - The third is JSON with a category that is not allowed, so
None.
Checking every answer from two prompts
This continues the file built in System prompts, which has ask, tickets, vague and structured, and Few-shot prompting, which adds examples. Add Triage and check from above:
prompts = {
"vague": [{"role": "system", "content": vague}],
"few-shot": [{"role": "system", "content": structured}, *examples],
}
for label, start in prompts.items():
results = [check(ask(start + [{"role": "user", "content": text}])) for text, _ in tickets]
passed = sum(result is not None for result in results)
print(f"{label}: {passed} of {len(results)} answers passed the check")vague: 0 of 7 answers passed the check few-shot: 7 of 7 answers passed the check
- vague: 0 of 7. None of its Markdown answers has the shape.
- few-shot: 7 of 7, including the fee and the cancellation filed as shipping.
shippingis an allowed value, so the check passes. It checks the shape, and it cannot know the right answer.
JSON mode and structured outputs on Groq
Groq can also enforce the shape on its side. response_format={"type": "json_object"} asks for valid JSON, and {"type": "json_schema", ...} with a JSON Schema, the kind Pydantic writes with Triage.model_json_schema(), asks for exactly that shape. It removes shape failures. It does not remove wrong answers, so the check in Evaluating prompts is still needed.
Shape check vs correctness check
| Shape check (Pydantic) | Correctness check (evaluation) | |
|---|---|---|
| Needs | The schema | Tickets with known right answers |
| Catches | Markdown, missing fields, unknown categories | A valid answer in the wrong category |
| Runs | On every answer, in production | On a test set, before a change ships |
When to validate
- Every time a program, not a person, reads the answer.
- Before writing a model's answer to a database or passing it to another system.
- After changing a prompt, to see whether the shape still holds.
None needs a plan: retry the request, send the ticket to a person, or file it as other. Silently dropping failed answers loses tickets.Related
- Previous: Few-shot prompting
- Next: Evaluating prompts
- Reference: Groq structured outputs
- Reference: Pydantic JSON validation
- Print the vague prompt's answers that failed the check.
- Add a third entry to
promptsfor the structured prompt without examples. - Call
check('{"category": "billing", "priority": 9}')and explain the result.
Slow is fine. Stopping is the only problem.