LLM Fundamentalsgpt-oss-120b on Groq · groq 1.7 · Python 3.10+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
19 small wins to finish your pathNext lesson →

Evaluating prompts

Evaluating a prompt means running it on inputs whose right answers you already know and counting how often it gets them right.

Last updated: 30 Sep, 2026 · groq 1.7 · gpt-oss-120b on Groq

Reading answers by eye suggested the structured prompt was better than the vague one. An evaluation turns that into a number. This lesson continues the same file: ask, tickets and both prompts from System prompts, examples from Few-shot prompting, and Triage and check from Structured output.

A score function

python
def score(start):
    right = 0
    for text, expected in tickets:
        result = check(ask(start + [{"role": "user", "content": text}]))
        if result is not None and result.category == expected:
            right += 1
    return right / len(tickets)

An answer counts only if it passes the check and has the right category. The result is the share of tickets right. The labelled tickets are a small test set: cases with known answers, used for measuring and never pasted into a prompt.

Scoring three prompts on seven tickets

ExampleAPI key
candidates = {
    "vague": [{"role": "system", "content": vague}],
    "structured": [{"role": "system", "content": structured}],
    "few-shot": [{"role": "system", "content": structured}, *examples],
}
for label, start in candidates.items():
    print(f"{label:10} {score(start):.0%}")
  • vague 0%: no answer passes the check, so none can be right.
  • structured 71%: five of seven.
  • few-shot 71%: also five of seven. The examples did not raise the score.
  • The number tells you where to look. In the runs in System prompts and Few-shot prompting, both prompts missed the same two tickets, the fee and the cancellation, so the next change should be about those.

Reading answers vs scoring them

Reading by eyeScoring
TicketsA handfulAs many as you have labels for
ResultAn impressionA number you can compare
RepeatableNoYes, after every change

When to run an evaluation

  • After every prompt change, however small.
  • Before switching to a cheaper or faster model.
  • When a user reports a wrong answer: add that case to the test set first.
Watch out. Seven tickets are far too few. One ticket is 14 points here, so a single lucky answer moves the score a lot. ask also uses the model's default temperature, so a rerun can change a score; Project: ticket sorter scores at temperature 0. Tools such as DeepEval, RAGAS and promptfoo are this loop with more metrics and reports.
Try it yourself
  • Inside score, change ask(...) to ask(..., temperature=0) and run the scores twice.
  • Add three tickets of your own, with labels, to tickets.
  • Score priority as well: count an answer right only if its priority is within 1 of a value you choose.

This is what real progress feels like.