Evals: scoring the agent on many tickets
An eval is a run of the agent over many cases that scores every answer, so you can measure how good it is and whether a change helped, instead of checking one path at a time.
Last updated: 28 Sep, 2026 · Pydantic AI 2.51
A test from Testing agents with pytest and TestModel proves one path passes. An eval runs many cases and gives a score. Pydantic Evals is the tool for it, and it uses the same Ticket model and scripted stand-in you already have.
A task and an evaluator
An eval needs a task, an async function from input to output, and one or more evaluators that score each output:
async def sort(ticket: str) -> str:
result = await agent.run(ticket)
return result.output.category # the task returns what you scoreclass Urgent(Evaluator):
def evaluate(self, ctx: EvaluatorContext) -> bool:
# fail a ticket marked urgent that landed in "other"
return ctx.output != "other" or "urgent" not in ctx.inputs.lower()A dataset of cases
A Case is one input with the output you expect. EqualsExpected checks the output matches it; your own Urgent checks a rule Pydantic cannot:
dataset = Dataset(
name="ticket sorting",
cases=[
Case(name="double charge", inputs="I was charged twice", expected_output="billing"),
Case(name="urgent broken", inputs="URGENT: the lamp came broken", expected_output="shipping"),
],
evaluators=[EqualsExpected(), Urgent()],
)View the code here
import re
from pydantic_ai import ModelResponse, TextPart, ToolCallPart
from pydantic_ai.models.function import AgentInfo, FunctionModel
def sort_ticket(text):
text = text.lower()
if "charged" in text or "refund" in text:
return "billing", 4
if "parcel" in text or "arrived" in text:
return "shipping", 3
return "other", 1
def shop_reply(messages, info: AgentInfo) -> ModelResponse:
prompts = [p.content for m in messages for p in m.parts if p.part_kind == "user-prompt"]
ticket = prompts[-1]
last = messages[-1].parts[-1]
order = re.search(r"A-\d{4}", ticket)
if order and info.function_tools and last.part_kind == "user-prompt":
tool = info.function_tools[0].name
return ModelResponse(parts=[ToolCallPart(tool, {"order_id": order.group()})])
if last.part_kind == "tool-return" and info.allow_text_output:
return ModelResponse(parts=[TextPart(f"Order {order.group()}: {last.content}.")])
category, priority = sort_ticket(ticket)
if info.output_tools:
args = {"category": category, "priority": priority}
return ModelResponse(parts=[ToolCallPart(info.output_tools[0].name, args)])
return ModelResponse(parts=[TextPart(f"Sorted as {category}.")])
shop_model = FunctionModel(shop_reply, model_name="shop")
Scoring the dataset
evaluate_sync runs the task on every case, then every evaluator on every output:
report = dataset.evaluate_sync(sort, progress=False)
for case in report.cases:
failed = [name for name, result in case.assertions.items() if not result.value]
print(f"{case.name:15} {case.output:9} failed: {failed}")
print(f"passed: {report.averages().assertions:.1%}")double charge billing failed: [] late parcel shipping failed: [] money back other failed: ['EqualsExpected'] urgent broken other failed: ['EqualsExpected', 'Urgent'] passed: 62.5%
Reading the score
- The first two sorted right, so their failed lists are empty.
- "I want my money back" has none of the stand-in's keywords, so it came back
otherand failedEqualsExpected. - "URGENT: the lamp came broken" failed both checks, since it was urgent and landed in
other. - Five of eight checks passed, which is 62.5%. Use the loop to fail a build when the pass rate drops.
Evaluators that need a model
Some qualities have no exact answer: is the reply polite? LLMJudge(rubric="...") from pydantic_evals.evaluators asks a model to decide, and needs a real model to do it. The DeepEval and Ragas courses go deeper into judges and their pitfalls.
Offline vs online evals
| Offline | Online | |
|---|---|---|
| When it runs | Before release, on a fixed dataset | In production, on real tickets as they arrive |
| Compared with | The last run of the same dataset | Nothing; it has no expected output |
| Uses | Input, output and the expected output | Only the ticket, the answer and the timing |
Scoring live tickets online
OnlineEvaluation is a capability: after each finished run it hands the ticket and answer to the evaluators and returns straight away. NotOther fails any ticket the desk could not place:
checked = []
def keep(results, failures, context):
checked.extend(f"{context.inputs}: {r.name} {r.value}" for r in results)
class NotOther(Evaluator):
def evaluate(self, ctx: EvaluatorContext) -> bool:
return ctx.output.category != "other"live = OnlineEvalConfig(default_sink=keep, default_sample_rate=1.0)
desk = Agent(shop_model, output_type=Ticket, name="desk",
capabilities=[OnlineEvaluation(evaluators=[NotOther()], config=live)])
async def main():
for ticket in ["I was charged twice", "I want my money back"]:
print(ticket, "->", (await desk.run(ticket)).output.category)
await wait_for_evaluations()
print(*sorted(checked), sep="\n")
asyncio.run(main())I was charged twice -> billing I want my money back -> other I want my money back: NotOther False I was charged twice: NotOther True
What the live run recorded
- The money-back ticket failed live, the same weakness the dataset found offline.
- wait_for_evaluations waits for checks still running, which a script needs and a web app does not.
- default_sample_rate=1.0 checks every run; on real traffic you run a cheap rule on all and an
LLMJudgeon a sample.
In production every result is also sent as an OpenTelemetry event, so with Logfire set up as in Observability: tracing runs with Logfire it appears in the Live Evaluations view next to the trace of the run that produced it.
Where you use evals
- Comparing two prompts or two models on the same dataset.
- Failing a build when the pass rate drops below a line.
- Watching quality on live traffic without making the customer wait.
Related
- Previous: Testing agents with pytest and TestModel
- Next: Observability: tracing runs with Logfire
- See also: Observability: tracing runs with Logfire
- Reference: Pydantic Evals
- Add
"money"and"broken"tosort_ticketand run the eval again; check nothing that passed now fails. - Add a case of your own that the stand-in gets wrong.
- Save the dataset with
dataset.to_file("tickets.yaml")and read the file.
Every expert started right here.