Cost
LLM cost is the number of input tokens times the input price plus the number of output tokens times the output price, with prices quoted per million tokens.
Last updated: 30 Sep, 2026 · groq 1.7 · gpt-oss-120b on Groq
Every response reports how many tokens it used, and the provider lists its prices, so the cost of a request is arithmetic.
The video asks which costs more, input tokens or output tokens, and the answer is output. Its reason: reading a question takes little effort, but answering it takes work. It then opens price pages: OpenAI's models charge several times more per output token than per input token, and every Claude model charges five times more, prices per million tokens.
The prices in the video are OpenAI's and Anthropic's on the day it was recorded, and prices change often. The code below reads Groq's current prices for gpt-oss from the API.
Syntax: cost = usage.prompt_tokens × input price + usage.completion_tokens × output price.
Reading Groq's prices
from groq import Groq
client = Groq() # reads GROQ_API_KEY from the environment
for model in client.models.list().data:
if model.id in ("openai/gpt-oss-20b", "openai/gpt-oss-120b"):
per_million_in = float(model.pricing["prompt"]) * 1_000_000
per_million_out = float(model.pricing["completion"]) * 1_000_000
print(f"{model.id}: ${per_million_in:.3f} in, ${per_million_out:.2f} out, per million tokens")openai/gpt-oss-120b: $0.150 in, $0.60 out, per million tokens openai/gpt-oss-20b: $0.075 in, $0.30 out, per million tokens
- Output costs four times input for both models, the same direction as in the video.
- The 20b costs half as much as the 120b, for input and for output.
Pricing one ticket
This continues the file with structured from System prompts and examples from Few-shot prompting:
prices = next(model.pricing for model in client.models.list().data if model.id == "openai/gpt-oss-120b")
messages = [{"role": "system", "content": structured}, *examples, {"role": "user", "content": "My parcel has not arrived"}]
response = client.chat.completions.create(model="openai/gpt-oss-120b", messages=messages)
usage = response.usage
cost = usage.prompt_tokens * float(prices["prompt"]) + usage.completion_tokens * float(prices["completion"])
print(usage.prompt_tokens, "tokens in,", usage.completion_tokens, "tokens out")
print(f"${cost:.6f} for this ticket")
print(f"${cost * 50_000:.2f} for 50,000 tickets")248 tokens in, 58 tokens out $0.000072 for this ticket $3.60 for 50,000 tickets
- 248 tokens in: the system prompt, the three examples and the ticket. The ticket itself is five of them.
- 58 tokens out, although the JSON answer is about a dozen tokens: the rest is the model's reasoning, which is billed as output.
- $3.60 for 50,000 tickets at this run's token counts.
Different tokenizers, different counts
import tiktoken
texts = ["My parcel has not arrived", "मेरा पार्सल नहीं आया", "1234567"]
for name in ["o200k_harmony", "cl100k_base"]:
encoding = tiktoken.get_encoding(name)
print(name, [len(encoding.encode(text)) for text in texts])o200k_harmony [5, 7, 3] cl100k_base [5, 20, 3]
cl100k_base is the tokenizer of OpenAI's GPT-4 and GPT-3.5 models. English and the number come out the same, but the Hindi sentence is 20 tokens there and 7 in o200k_harmony. Count with the tokenizer of the model you pay for, or read usage from the response.
Input vs output tokens
| Input tokens | Output tokens | |
|---|---|---|
| What they are | System prompt, examples, history, the question | The answer and any reasoning |
| Price per token | Lower | Higher |
| Grow with | Longer prompts and longer chats | Longer answers and more reasoning |
Ways to spend less
- Ask for short answers: a JSON line instead of a paragraph.
- Lower
reasoning_effortwhere the task is simple (Max tokens and finish_reason). - Trim long histories (Context window).
- Use the smaller model where its score is good enough (Choosing a model).
Related
- Previous: Context window
- Next: Latency
- Reference: Groq pricing
- Work out the cost of the vague prompt from System prompts for 50,000 tickets.
- Run the ticket call with
reasoning_effort="low"and compare the output tokens. - Count the Hindi sentence from Tokens with
o200k_base, the tokenizero200k_harmonyis built on, and compare the counts.
Little by little, you're building something great.