LLM Fundamentalsgpt-oss-120b on Groq · groq 1.7 · Python 3.10+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
19 small wins to finish your pathNext lesson →

Cost

LLM cost is the number of input tokens times the input price plus the number of output tokens times the output price, with prices quoted per million tokens.

Last updated: 30 Sep, 2026 · groq 1.7 · gpt-oss-120b on Groq

Every response reports how many tokens it used, and the provider lists its prices, so the cost of a request is arithmetic.

Input and output token prices · from the How AI Actually Works + Why Your Prompts Keep Failing session · 22:08 to 26:40

The video asks which costs more, input tokens or output tokens, and the answer is output. Its reason: reading a question takes little effort, but answering it takes work. It then opens price pages: OpenAI's models charge several times more per output token than per input token, and every Claude model charges five times more, prices per million tokens.

The prices in the video are OpenAI's and Anthropic's on the day it was recorded, and prices change often. The code below reads Groq's current prices for gpt-oss from the API.

Syntax: cost = usage.prompt_tokens × input price + usage.completion_tokens × output price.

Reading Groq's prices

ExampleAPI key
from groq import Groq

client = Groq()  # reads GROQ_API_KEY from the environment

for model in client.models.list().data:
    if model.id in ("openai/gpt-oss-20b", "openai/gpt-oss-120b"):
        per_million_in = float(model.pricing["prompt"]) * 1_000_000
        per_million_out = float(model.pricing["completion"]) * 1_000_000
        print(f"{model.id}: ${per_million_in:.3f} in, ${per_million_out:.2f} out, per million tokens")
  • Output costs four times input for both models, the same direction as in the video.
  • The 20b costs half as much as the 120b, for input and for output.

Pricing one ticket

This continues the file with structured from System prompts and examples from Few-shot prompting:

ExampleAPI key
prices = next(model.pricing for model in client.models.list().data if model.id == "openai/gpt-oss-120b")

messages = [{"role": "system", "content": structured}, *examples, {"role": "user", "content": "My parcel has not arrived"}]
response = client.chat.completions.create(model="openai/gpt-oss-120b", messages=messages)
usage = response.usage

cost = usage.prompt_tokens * float(prices["prompt"]) + usage.completion_tokens * float(prices["completion"])
print(usage.prompt_tokens, "tokens in,", usage.completion_tokens, "tokens out")
print(f"${cost:.6f} for this ticket")
print(f"${cost * 50_000:.2f} for 50,000 tickets")
  • 248 tokens in: the system prompt, the three examples and the ticket. The ticket itself is five of them.
  • 58 tokens out, although the JSON answer is about a dozen tokens: the rest is the model's reasoning, which is billed as output.
  • $3.60 for 50,000 tickets at this run's token counts.

Different tokenizers, different counts

Example
import tiktoken

texts = ["My parcel has not arrived", "मेरा पार्सल नहीं आया", "1234567"]
for name in ["o200k_harmony", "cl100k_base"]:
    encoding = tiktoken.get_encoding(name)
    print(name, [len(encoding.encode(text)) for text in texts])

cl100k_base is the tokenizer of OpenAI's GPT-4 and GPT-3.5 models. English and the number come out the same, but the Hindi sentence is 20 tokens there and 7 in o200k_harmony. Count with the tokenizer of the model you pay for, or read usage from the response.

Input vs output tokens

Input tokensOutput tokens
What they areSystem prompt, examples, history, the questionThe answer and any reasoning
Price per tokenLowerHigher
Grow withLonger prompts and longer chatsLonger answers and more reasoning

Ways to spend less

Watch out. Examples are not free. The three examples that Few-shot prompting added are sent with every ticket. Measure what they add to the score against what they add to the bill.
Try it yourself
  • Work out the cost of the vague prompt from System prompts for 50,000 tickets.
  • Run the ticket call with reasoning_effort="low" and compare the output tokens.
  • Count the Hindi sentence from Tokens with o200k_base, the tokenizer o200k_harmony is built on, and compare the counts.

Little by little, you're building something great.