LLM Fundamentalsgpt-oss-120b on Groq · groq 1.7 · Python 3.10+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
19 small wins to finish your path

Project: ticket sorter

The ticket sorter is a two-file program that sorts seven labelled tickets with gpt-oss-120b and reports accuracy, tokens, cost and time for two versions of the prompt.

Last updated: 30 Sep, 2026 · groq 1.7 · gpt-oss-120b on Groq

The overview promised a ticket-sorting prompt, a score for how often it is right, and its cost and speed. One script reports all of them, for the prompt as Few-shot prompting left it and for an improved version.

The prompt in its own file

python
from typing import Literal

from pydantic import BaseModel, Field, ValidationError


class Triage(BaseModel):
    category: Literal["billing", "shipping", "other"]
    priority: int = Field(ge=1, le=5)


def check(answer):
    try:
        return Triage.model_validate_json(answer)
    except ValidationError:
        return None


FORMAT = """Reply with one line of JSON and nothing else, like this:
{"category": "shipping", "priority": 3}
priority is 1 (can wait) to 5 (urgent)."""

BEFORE = """You sort customer support tickets for an online shop.

Categories:
- billing: payments, charges, refunds
- shipping: parcels, delivery, damaged boxes
- other: anything else

""" + FORMAT

AFTER = """You sort customer support tickets for an online shop.

Categories:
- billing: payments, charges, refunds, and any fee a customer is asked to pay, even at the door
- shipping: where a parcel is, late or lost deliveries, damaged boxes
- other: account changes, cancelling an order, anything else

""" + FORMAT

EXAMPLES = [
    {"role": "user", "content": "You took money from my card two times"},
    {"role": "assistant", "content": '{"category": "billing", "priority": 4}'},
    {"role": "user", "content": "Where is my delivery? It is a week late"},
    {"role": "assistant", "content": '{"category": "shipping", "priority": 3}'},
    {"role": "user", "content": "Can I change the email on my account?"},
    {"role": "assistant", "content": '{"category": "other", "priority": 2}'},
]

TICKETS = [
    ("I was charged twice for one order", "billing"),
    ("My parcel has not arrived", "shipping"),
    ("Can I get a refund for the blue mug?", "billing"),
    ("How do I change my password?", "other"),
    ("The parcel arrived but the box was crushed", "shipping"),
    ("The courier asked me to pay a delivery fee at the door", "billing"),
    ("I want to cancel my order before it ships", "other"),
]

Triage and check are from Structured output, the examples from Few-shot prompting and the tickets from System prompts. BEFORE is the structured prompt. AFTER changes the category definitions for the two tickets every earlier run got wrong: fees asked for at the door are billing, and cancelling an order is other.

The script

Loading the prompt and the prices

python
import time

from groq import Groq

from prompt import AFTER, BEFORE, EXAMPLES, TICKETS, check

MODEL = "openai/gpt-oss-120b"
client = Groq()
prices = next(model.pricing for model in client.models.list().data if model.id == MODEL)

The prices are read from the model list, as in Cost, so the cost stays current.

Sorting one ticket at temperature 0

python
def sort_ticket(system, text):
    messages = [{"role": "system", "content": system}, *EXAMPLES, {"role": "user", "content": text}]
    response = client.chat.completions.create(model=MODEL, messages=messages, temperature=0)
    return response.choices[0].message.content, response.usage

temperature=0, from Temperature, makes the scores repeatable, so a change in the score comes from the prompt. The function returns the answer and usage, the token counts.

Scoring one version of the prompt

python
def run(label, system):
    right, tokens_in, tokens_out = 0, 0, 0
    started = time.perf_counter()
    print(label)
    for text, expected in TICKETS:
        answer, usage = sort_ticket(system, text)
        tokens_in, tokens_out = tokens_in + usage.prompt_tokens, tokens_out + usage.completion_tokens
        result = check(answer)
        ok = result is not None and result.category == expected
        right += ok
        print(f"  {'ok ' if ok else 'BAD'} {expected:9} {answer}")
    seconds = time.perf_counter() - started
    cost = tokens_in * float(prices["prompt"]) + tokens_out * float(prices["completion"])
    print(f"  accuracy {right}/{len(TICKETS)}, {tokens_in} tokens in, {tokens_out} out, "
          f"${cost:.6f}, {seconds / len(TICKETS):.2f} seconds a ticket")


run("before", BEFORE)
run("after", AFTER)

Each ticket is sorted, checked and compared with its label. The token counts feed the cost, and the clock gives the time per ticket. right += ok adds 1 for True.

Scoring the before and after prompts

Save both files in one folder and run python sort.py:

ExampleAPI keysort.py
import time

from groq import Groq

from prompt import AFTER, BEFORE, EXAMPLES, TICKETS, check

MODEL = "openai/gpt-oss-120b"
client = Groq()
prices = next(model.pricing for model in client.models.list().data if model.id == MODEL)


def sort_ticket(system, text):
    messages = [{"role": "system", "content": system}, *EXAMPLES, {"role": "user", "content": text}]
    response = client.chat.completions.create(model=MODEL, messages=messages, temperature=0)
    return response.choices[0].message.content, response.usage


def run(label, system):
    right, tokens_in, tokens_out = 0, 0, 0
    started = time.perf_counter()
    print(label)
    for text, expected in TICKETS:
        answer, usage = sort_ticket(system, text)
        tokens_in, tokens_out = tokens_in + usage.prompt_tokens, tokens_out + usage.completion_tokens
        result = check(answer)
        ok = result is not None and result.category == expected
        right += ok
        print(f"  {'ok ' if ok else 'BAD'} {expected:9} {answer}")
    seconds = time.perf_counter() - started
    cost = tokens_in * float(prices["prompt"]) + tokens_out * float(prices["completion"])
    print(f"  accuracy {right}/{len(TICKETS)}, {tokens_in} tokens in, {tokens_out} out, "
          f"${cost:.6f}, {seconds / len(TICKETS):.2f} seconds a ticket")


run("before", BEFORE)
run("after", AFTER)

What the two runs show

  • before: 5 of 7. The fee and the cancellation are wrong, the same two tickets the structured and few-shot prompts missed.
  • after: 7 of 7. Both definitions landed, and no other ticket got worse.
  • The fix cost tokens. Input went from 1,759 to 2,211 tokens, because the longer definitions are sent with every ticket, and the cost rose from $0.000578 to $0.000642.
  • Both ran in under a second a ticket.
One script: accuracy, tokens, cost and time
messagesrequestanswerscoredprompt.pyBEFORE, AFTER, 3 examplesmessages listsystem, examples, ticketgpt-oss-120b on Groqtemperature 0check()Pydantic TriageThe reportaccuracy, tokens, cost, time
Hover or tap a piece to see what it is and which lesson built it.
Trace the run

Pick one to watch it run, step by step.

This is the loop AI engineers improve prompts with: change one thing, score it, look at the tokens, the cost and the time, and keep the change or throw it away.

Choosing a model with the same script

Change MODEL to "openai/gpt-oss-20b" and run it again. The same seven tickets give the smaller model's score, cost and speed, and the right choice is the cheapest model whose score clears the bar you agreed on.

Seven tickets vs a real test set

These seven ticketsA real test set
Size7Dozens to hundreds
SourceWritten for this courseReal tickets, including awkward ones
Seen while writing the promptYes, AFTER was written for two of themNo, kept apart

LLM topics for later courses

TopicWhat it is for
EmbeddingsTurning text into vectors to find similar documents, the base of retrieval.
Retrieval (RAG)Fetching the relevant documents and putting them in the prompt, the defence against made-up answers.
Tool callingLetting the model ask your code to run a function, the base of agents.
Fine-tuningTraining a model's weights on your own examples when prompting is not enough.
Running open weights locallyDownloading a model such as gpt-oss-20b and serving it on your own hardware.
Prompt injectionText in a ticket that tries to override your instructions; covered in the guardrails and red-teaming courses.
Watch out. AFTER was written after seeing these tickets, so 7 of 7 here is not proof that it is better on tickets it has never seen. Before trusting a change, score it on new labelled tickets.
Try it yourself
  • Change MODEL to "openai/gpt-oss-20b" and compare the two reports.
  • Add three new labelled tickets to TICKETS, including another fee or cancellation, and run both prompts again.
  • Remove the examples from sort_ticket and see what the definitions alone score.
PreviousLatency

Slow is fine. Stopping is the only problem.