Pydantic AIPydantic AI 2.51 · Python 3.10+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
29 small wins to finish your pathNext lesson →

Evals: scoring the agent on many tickets

An eval is a run of the agent over many cases that scores every answer, so you can measure how good it is and whether a change helped, instead of checking one path at a time.

Last updated: 28 Sep, 2026 · Pydantic AI 2.51

A test from Testing agents with pytest and TestModel proves one path passes. An eval runs many cases and gives a score. Pydantic Evals is the tool for it, and it uses the same Ticket model and scripted stand-in you already have.

A task and an evaluator

An eval needs a task, an async function from input to output, and one or more evaluators that score each output:

python
async def sort(ticket: str) -> str:
    result = await agent.run(ticket)
    return result.output.category   # the task returns what you score
python
class Urgent(Evaluator):
    def evaluate(self, ctx: EvaluatorContext) -> bool:
        # fail a ticket marked urgent that landed in "other"
        return ctx.output != "other" or "urgent" not in ctx.inputs.lower()

A dataset of cases

A Case is one input with the output you expect. EqualsExpected checks the output matches it; your own Urgent checks a rule Pydantic cannot:

python
dataset = Dataset(
    name="ticket sorting",
    cases=[
        Case(name="double charge", inputs="I was charged twice", expected_output="billing"),
        Case(name="urgent broken", inputs="URGENT: the lamp came broken", expected_output="shipping"),
    ],
    evaluators=[EqualsExpected(), Urgent()],
)
Project files used on this pageThis lesson builds on a project from earlier lessons. The code below imports this file. Click a file to see its code, or follow the link to the lesson that wrote it. To run the code yourself, keep it in the same folder.
View the code here
shop_model.py
import re

from pydantic_ai import ModelResponse, TextPart, ToolCallPart
from pydantic_ai.models.function import AgentInfo, FunctionModel


def sort_ticket(text):
    text = text.lower()
    if "charged" in text or "refund" in text:
        return "billing", 4
    if "parcel" in text or "arrived" in text:
        return "shipping", 3
    return "other", 1


def shop_reply(messages, info: AgentInfo) -> ModelResponse:
    prompts = [p.content for m in messages for p in m.parts if p.part_kind == "user-prompt"]
    ticket = prompts[-1]
    last = messages[-1].parts[-1]
    order = re.search(r"A-\d{4}", ticket)

    if order and info.function_tools and last.part_kind == "user-prompt":
        tool = info.function_tools[0].name
        return ModelResponse(parts=[ToolCallPart(tool, {"order_id": order.group()})])

    if last.part_kind == "tool-return" and info.allow_text_output:
        return ModelResponse(parts=[TextPart(f"Order {order.group()}: {last.content}.")])

    category, priority = sort_ticket(ticket)
    if info.output_tools:
        args = {"category": category, "priority": priority}
        return ModelResponse(parts=[ToolCallPart(info.output_tools[0].name, args)])

    return ModelResponse(parts=[TextPart(f"Sorted as {category}.")])


shop_model = FunctionModel(shop_reply, model_name="shop")

Scoring the dataset

evaluate_sync runs the task on every case, then every evaluator on every output:

Example
report = dataset.evaluate_sync(sort, progress=False)
for case in report.cases:
    failed = [name for name, result in case.assertions.items() if not result.value]
    print(f"{case.name:15} {case.output:9} failed: {failed}")
print(f"passed: {report.averages().assertions:.1%}")

Reading the score

  • The first two sorted right, so their failed lists are empty.
  • "I want my money back" has none of the stand-in's keywords, so it came back other and failed EqualsExpected.
  • "URGENT: the lamp came broken" failed both checks, since it was urgent and landed in other.
  • Five of eight checks passed, which is 62.5%. Use the loop to fail a build when the pass rate drops.

Evaluators that need a model

Some qualities have no exact answer: is the reply polite? LLMJudge(rubric="...") from pydantic_evals.evaluators asks a model to decide, and needs a real model to do it. The DeepEval and Ragas courses go deeper into judges and their pitfalls.

Offline vs online evals

OfflineOnline
When it runsBefore release, on a fixed datasetIn production, on real tickets as they arrive
Compared withThe last run of the same datasetNothing; it has no expected output
UsesInput, output and the expected outputOnly the ticket, the answer and the timing

Scoring live tickets online

OnlineEvaluation is a capability: after each finished run it hands the ticket and answer to the evaluators and returns straight away. NotOther fails any ticket the desk could not place:

python
checked = []


def keep(results, failures, context):
    checked.extend(f"{context.inputs}: {r.name} {r.value}" for r in results)


class NotOther(Evaluator):
    def evaluate(self, ctx: EvaluatorContext) -> bool:
        return ctx.output.category != "other"
Example
live = OnlineEvalConfig(default_sink=keep, default_sample_rate=1.0)
desk = Agent(shop_model, output_type=Ticket, name="desk",
             capabilities=[OnlineEvaluation(evaluators=[NotOther()], config=live)])


async def main():
    for ticket in ["I was charged twice", "I want my money back"]:
        print(ticket, "->", (await desk.run(ticket)).output.category)
    await wait_for_evaluations()
    print(*sorted(checked), sep="\n")


asyncio.run(main())

What the live run recorded

  • The money-back ticket failed live, the same weakness the dataset found offline.
  • wait_for_evaluations waits for checks still running, which a script needs and a web app does not.
  • default_sample_rate=1.0 checks every run; on real traffic you run a cheap rule on all and an LLMJudge on a sample.

In production every result is also sent as an OpenTelemetry event, so with Logfire set up as in Observability: tracing runs with Logfire it appears in the Live Evaluations view next to the trace of the run that produced it.

Where you use evals

  • Comparing two prompts or two models on the same dataset.
  • Failing a build when the pass rate drops below a line.
  • Watching quality on live traffic without making the customer wait.
The score is only as good as the cases
An offline score only means as much as its dataset. If a real failure is not a case, the eval will not catch it. Add every ticket the desk gets wrong to the dataset, so a fix is measured and a change that breaks it is caught.
Try it yourself
  • Add "money" and "broken" to sort_ticket and run the eval again; check nothing that passed now fails.
  • Add a case of your own that the stand-in gets wrong.
  • Save the dataset with dataset.to_file("tickets.yaml") and read the file.

Every expert started right here.