Project: ticket sorter
The ticket sorter is a two-file program that sorts seven labelled tickets with gpt-oss-120b and reports accuracy, tokens, cost and time for two versions of the prompt.
Last updated: 30 Sep, 2026 · groq 1.7 · gpt-oss-120b on Groq
The overview promised a ticket-sorting prompt, a score for how often it is right, and its cost and speed. One script reports all of them, for the prompt as Few-shot prompting left it and for an improved version.
The prompt in its own file
from typing import Literal
from pydantic import BaseModel, Field, ValidationError
class Triage(BaseModel):
category: Literal["billing", "shipping", "other"]
priority: int = Field(ge=1, le=5)
def check(answer):
try:
return Triage.model_validate_json(answer)
except ValidationError:
return None
FORMAT = """Reply with one line of JSON and nothing else, like this:
{"category": "shipping", "priority": 3}
priority is 1 (can wait) to 5 (urgent)."""
BEFORE = """You sort customer support tickets for an online shop.
Categories:
- billing: payments, charges, refunds
- shipping: parcels, delivery, damaged boxes
- other: anything else
""" + FORMAT
AFTER = """You sort customer support tickets for an online shop.
Categories:
- billing: payments, charges, refunds, and any fee a customer is asked to pay, even at the door
- shipping: where a parcel is, late or lost deliveries, damaged boxes
- other: account changes, cancelling an order, anything else
""" + FORMAT
EXAMPLES = [
{"role": "user", "content": "You took money from my card two times"},
{"role": "assistant", "content": '{"category": "billing", "priority": 4}'},
{"role": "user", "content": "Where is my delivery? It is a week late"},
{"role": "assistant", "content": '{"category": "shipping", "priority": 3}'},
{"role": "user", "content": "Can I change the email on my account?"},
{"role": "assistant", "content": '{"category": "other", "priority": 2}'},
]
TICKETS = [
("I was charged twice for one order", "billing"),
("My parcel has not arrived", "shipping"),
("Can I get a refund for the blue mug?", "billing"),
("How do I change my password?", "other"),
("The parcel arrived but the box was crushed", "shipping"),
("The courier asked me to pay a delivery fee at the door", "billing"),
("I want to cancel my order before it ships", "other"),
]
Triage and check are from Structured output, the examples from Few-shot prompting and the tickets from System prompts. BEFORE is the structured prompt. AFTER changes the category definitions for the two tickets every earlier run got wrong: fees asked for at the door are billing, and cancelling an order is other.
The script
Loading the prompt and the prices
import time
from groq import Groq
from prompt import AFTER, BEFORE, EXAMPLES, TICKETS, check
MODEL = "openai/gpt-oss-120b"
client = Groq()
prices = next(model.pricing for model in client.models.list().data if model.id == MODEL)The prices are read from the model list, as in Cost, so the cost stays current.
Sorting one ticket at temperature 0
def sort_ticket(system, text):
messages = [{"role": "system", "content": system}, *EXAMPLES, {"role": "user", "content": text}]
response = client.chat.completions.create(model=MODEL, messages=messages, temperature=0)
return response.choices[0].message.content, response.usagetemperature=0, from Temperature, makes the scores repeatable, so a change in the score comes from the prompt. The function returns the answer and usage, the token counts.
Scoring one version of the prompt
def run(label, system):
right, tokens_in, tokens_out = 0, 0, 0
started = time.perf_counter()
print(label)
for text, expected in TICKETS:
answer, usage = sort_ticket(system, text)
tokens_in, tokens_out = tokens_in + usage.prompt_tokens, tokens_out + usage.completion_tokens
result = check(answer)
ok = result is not None and result.category == expected
right += ok
print(f" {'ok ' if ok else 'BAD'} {expected:9} {answer}")
seconds = time.perf_counter() - started
cost = tokens_in * float(prices["prompt"]) + tokens_out * float(prices["completion"])
print(f" accuracy {right}/{len(TICKETS)}, {tokens_in} tokens in, {tokens_out} out, "
f"${cost:.6f}, {seconds / len(TICKETS):.2f} seconds a ticket")
run("before", BEFORE)
run("after", AFTER)Each ticket is sorted, checked and compared with its label. The token counts feed the cost, and the clock gives the time per ticket. right += ok adds 1 for True.
Scoring the before and after prompts
Save both files in one folder and run python sort.py:
import time
from groq import Groq
from prompt import AFTER, BEFORE, EXAMPLES, TICKETS, check
MODEL = "openai/gpt-oss-120b"
client = Groq()
prices = next(model.pricing for model in client.models.list().data if model.id == MODEL)
def sort_ticket(system, text):
messages = [{"role": "system", "content": system}, *EXAMPLES, {"role": "user", "content": text}]
response = client.chat.completions.create(model=MODEL, messages=messages, temperature=0)
return response.choices[0].message.content, response.usage
def run(label, system):
right, tokens_in, tokens_out = 0, 0, 0
started = time.perf_counter()
print(label)
for text, expected in TICKETS:
answer, usage = sort_ticket(system, text)
tokens_in, tokens_out = tokens_in + usage.prompt_tokens, tokens_out + usage.completion_tokens
result = check(answer)
ok = result is not None and result.category == expected
right += ok
print(f" {'ok ' if ok else 'BAD'} {expected:9} {answer}")
seconds = time.perf_counter() - started
cost = tokens_in * float(prices["prompt"]) + tokens_out * float(prices["completion"])
print(f" accuracy {right}/{len(TICKETS)}, {tokens_in} tokens in, {tokens_out} out, "
f"${cost:.6f}, {seconds / len(TICKETS):.2f} seconds a ticket")
run("before", BEFORE)
run("after", AFTER)
before
ok billing {"category": "billing", "priority": 4}
ok shipping {"category": "shipping", "priority": 4}
ok billing {"category": "billing", "priority": 4}
ok other {"category": "other", "priority": 2}
ok shipping {"category": "shipping", "priority": 4}
BAD billing {"category": "shipping", "priority": 4}
BAD other {"category": "shipping", "priority": 5}
accuracy 5/7, 1759 tokens in, 524 out, $0.000578, 0.63 seconds a ticket
after
ok billing {"category": "billing", "priority": 4}
ok shipping {"category": "shipping", "priority": 4}
ok billing {"category": "billing", "priority": 4}
ok other {"category": "other", "priority": 2}
ok shipping {"category": "shipping", "priority": 3}
ok billing {"category": "billing", "priority": 4}
ok other {"category": "other", "priority": 4}
accuracy 7/7, 2211 tokens in, 518 out, $0.000642, 0.74 seconds a ticketWhat the two runs show
- before: 5 of 7. The fee and the cancellation are wrong, the same two tickets the structured and few-shot prompts missed.
- after: 7 of 7. Both definitions landed, and no other ticket got worse.
- The fix cost tokens. Input went from 1,759 to 2,211 tokens, because the longer definitions are sent with every ticket, and the cost rose from $0.000578 to $0.000642.
- Both ran in under a second a ticket.
Pick one to watch it run, step by step.
This is the loop AI engineers improve prompts with: change one thing, score it, look at the tokens, the cost and the time, and keep the change or throw it away.
Choosing a model with the same script
Change MODEL to "openai/gpt-oss-20b" and run it again. The same seven tickets give the smaller model's score, cost and speed, and the right choice is the cheapest model whose score clears the bar you agreed on.
Seven tickets vs a real test set
| These seven tickets | A real test set | |
|---|---|---|
| Size | 7 | Dozens to hundreds |
| Source | Written for this course | Real tickets, including awkward ones |
| Seen while writing the prompt | Yes, AFTER was written for two of them | No, kept apart |
LLM topics for later courses
| Topic | What it is for |
|---|---|
| Embeddings | Turning text into vectors to find similar documents, the base of retrieval. |
| Retrieval (RAG) | Fetching the relevant documents and putting them in the prompt, the defence against made-up answers. |
| Tool calling | Letting the model ask your code to run a function, the base of agents. |
| Fine-tuning | Training a model's weights on your own examples when prompting is not enough. |
| Running open weights locally | Downloading a model such as gpt-oss-20b and serving it on your own hardware. |
| Prompt injection | Text in a ticket that tries to override your instructions; covered in the guardrails and red-teaming courses. |
AFTER was written after seeing these tickets, so 7 of 7 here is not proof that it is better on tickets it has never seen. Before trusting a change, score it on new labelled tickets.Related
- Previous: Latency
- See also: LLM Fundamentals overview
- Change
MODELto"openai/gpt-oss-20b"and compare the two reports. - Add three new labelled tickets to
TICKETS, including another fee or cancellation, and run both prompts again. - Remove the examples from
sort_ticketand see what the definitions alone score.
Slow is fine. Stopping is the only problem.