LLM Fundamentalsgpt-oss-120b on Groq · groq 1.7 · Python 3.10+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
19 small wins to finish your pathNext lesson →

Structured output

Structured output means getting the model's answer in a fixed shape a program can read, such as JSON with named fields, and checking every answer against that shape before using it.

Last updated: 30 Sep, 2026 · groq 1.7 · gpt-oss-120b on Groq

A program that uses a model's answer cannot read it by eye. It has to check the shape every time, because the model can return anything. Pydantic, from Python for AI, does the checking.

The Triage model

python
from typing import Literal

from pydantic import BaseModel, Field, ValidationError


class Triage(BaseModel):
    category: Literal["billing", "shipping", "other"]
    priority: int = Field(ge=1, le=5)

Literal allows only the three category names. Field(ge=1, le=5) allows priorities from 1 to 5.

A check function

python
def check(answer):
    try:
        return Triage.model_validate_json(answer)
    except ValidationError:
        return None

check returns a Triage for an answer in the right shape and None for anything else.

Checking three answers by hand

Example
print(check('{"category": "billing", "priority": 4}'))
print(check("**Category:** Billing"))
print(check('{"category": "refunds", "priority": 4}'))
  • The first is valid and comes back as a Triage.
  • The second is the vague prompt's Markdown: not JSON, so None.
  • The third is JSON with a category that is not allowed, so None.

Checking every answer from two prompts

This continues the file built in System prompts, which has ask, tickets, vague and structured, and Few-shot prompting, which adds examples. Add Triage and check from above:

ExampleAPI key
prompts = {
    "vague": [{"role": "system", "content": vague}],
    "few-shot": [{"role": "system", "content": structured}, *examples],
}
for label, start in prompts.items():
    results = [check(ask(start + [{"role": "user", "content": text}])) for text, _ in tickets]
    passed = sum(result is not None for result in results)
    print(f"{label}: {passed} of {len(results)} answers passed the check")
  • vague: 0 of 7. None of its Markdown answers has the shape.
  • few-shot: 7 of 7, including the fee and the cancellation filed as shipping. shipping is an allowed value, so the check passes. It checks the shape, and it cannot know the right answer.

JSON mode and structured outputs on Groq

Groq can also enforce the shape on its side. response_format={"type": "json_object"} asks for valid JSON, and {"type": "json_schema", ...} with a JSON Schema, the kind Pydantic writes with Triage.model_json_schema(), asks for exactly that shape. It removes shape failures. It does not remove wrong answers, so the check in Evaluating prompts is still needed.

Shape check vs correctness check

Shape check (Pydantic)Correctness check (evaluation)
NeedsThe schemaTickets with known right answers
CatchesMarkdown, missing fields, unknown categoriesA valid answer in the wrong category
RunsOn every answer, in productionOn a test set, before a change ships

When to validate

  • Every time a program, not a person, reads the answer.
  • Before writing a model's answer to a database or passing it to another system.
  • After changing a prompt, to see whether the shape still holds.
Watch out. A check that returns None needs a plan: retry the request, send the ticket to a person, or file it as other. Silently dropping failed answers loses tickets.
Try it yourself
  • Print the vague prompt's answers that failed the check.
  • Add a third entry to prompts for the structured prompt without examples.
  • Call check('{"category": "billing", "priority": 9}') and explain the result.

Slow is fine. Stopping is the only problem.