Evaluating prompts
Evaluating a prompt means running it on inputs whose right answers you already know and counting how often it gets them right.
Last updated: 30 Sep, 2026 · groq 1.7 · gpt-oss-120b on Groq
Reading answers by eye suggested the structured prompt was better than the vague one. An evaluation turns that into a number. This lesson continues the same file: ask, tickets and both prompts from System prompts, examples from Few-shot prompting, and Triage and check from Structured output.
A score function
def score(start):
right = 0
for text, expected in tickets:
result = check(ask(start + [{"role": "user", "content": text}]))
if result is not None and result.category == expected:
right += 1
return right / len(tickets)An answer counts only if it passes the check and has the right category. The result is the share of tickets right. The labelled tickets are a small test set: cases with known answers, used for measuring and never pasted into a prompt.
Scoring three prompts on seven tickets
candidates = {
"vague": [{"role": "system", "content": vague}],
"structured": [{"role": "system", "content": structured}],
"few-shot": [{"role": "system", "content": structured}, *examples],
}
for label, start in candidates.items():
print(f"{label:10} {score(start):.0%}")vague 0% structured 71% few-shot 71%
- vague 0%: no answer passes the check, so none can be right.
- structured 71%: five of seven.
- few-shot 71%: also five of seven. The examples did not raise the score.
- The number tells you where to look. In the runs in System prompts and Few-shot prompting, both prompts missed the same two tickets, the fee and the cancellation, so the next change should be about those.
Reading answers vs scoring them
| Reading by eye | Scoring | |
|---|---|---|
| Tickets | A handful | As many as you have labels for |
| Result | An impression | A number you can compare |
| Repeatable | No | Yes, after every change |
When to run an evaluation
- After every prompt change, however small.
- Before switching to a cheaper or faster model.
- When a user reports a wrong answer: add that case to the test set first.
ask also uses the model's default temperature, so a rerun can change a score; Project: ticket sorter scores at temperature 0. Tools such as DeepEval, RAGAS and promptfoo are this loop with more metrics and reports.Related
- Previous: Structured output
- Next: Context window
- Inside
score, changeask(...)toask(..., temperature=0)and run the scores twice. - Add three tickets of your own, with labels, to
tickets. - Score priority as well: count an answer right only if its priority is within 1 of a value you choose.
This is what real progress feels like.