LLM Fundamentalsgpt-oss-120b on Groq · groq 1.7 · Python 3.10+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
19 small wins to finish your pathNext lesson →

LLM Fundamentals logoLLM Fundamentals overview

A large language model (LLM) is a neural network that reads text as tokens and writes its answer one token at a time, each token chosen from probabilities it learned from a very large amount of text.

Last updated: 30 Sep, 2026 · groq 1.7 · gpt-oss-120b on Groq

Almost every current LLM is a transformer, the architecture from the 2017 paper Attention Is All You Need. Chat models such as GPT, Claude, Gemini, Llama and Qwen are first trained to predict the next token across a huge amount of text, then tuned to follow instructions and hold a conversation.

You reach a model in one of two ways. A hosted model runs on a provider's computers, and you call it over an API with a key. An open-weight model publishes its weights, the numbers it learned, so anyone can download it, and providers can host it for you.

What happens between a prompt and an answer
Text becomes tokenspieces, each a numberEvery next token scoredone pass of the modelOne token is pickedtemperature decidesRepeatuntil it stopsTokens become textthe answer you read
Hover or tap a piece to see what it is and which lesson built it.
Trace a prompt

Pick one to watch it run, step by step.

gpt-oss-120b on Groq

Every example calls gpt-oss-120b, an open-weight model that OpenAI released in August 2025 under the Apache 2.0 licence. Groq hosts it, and Groq's free plan needs only an account and an API key. The code uses Groq's Python library, groq, to send requests, tiktoken, OpenAI's tokenizer library, to count tokens, and pydantic to check answers.

The ideas do not depend on Groq or on this model. Tokens, sampling settings, chat messages, context windows and prices work the same way with OpenAI, Anthropic and Gemini; the client library and the model name are what change.

Tokens, sampling, prompts and limits

Everything in this course, and where it is going
How a model writesControlling the outputPromptsIn productiontokensnext-token predictiongenerating texttemperaturetop_p and top_kseedschat messagessystem prompts and examplesevaluating promptscontext windowcostproject: ticket sorter
Hover or tap a piece to see what it is and which lesson built it.
PartWhat you runLessons
How a model writesCount tokens, turn scores into probabilities, stream an answerTokens, Next-token prediction, Generating text
Controlling the outputChange temperature, top_p, the length cap and the seed, and compare answersTemperature, top_p and top_k, Seeds
PromptsWrite a system prompt and examples, check and score the answersSystem prompts, Few-shot prompting, Evaluating prompts
In productionMeasure the context window, the cost and the time of a requestContext window, Cost, Latency

The ticket sorter

One task runs through the prompt lessons: sort an online shop's support tickets into billing, shipping or other, with a priority, as one line of JSON. It starts as a vague instruction in System prompts, fails in ways you can see, and ends in Project: ticket sorter as a small program that reports how many tickets it got right, how many tokens that took, what it cost and how long it ran.

Python and a free Groq key

  • Python 3.10 or later, with functions, lists, dictionaries and Pydantic, as in Python for AI.
  • A free Groq API key, created in Installation and setup.
Watch out. Every reply on these pages is a real answer from gpt-oss-120b. Run the same code and your wording will differ, because a chat model picks its words at random unless told not to. Temperature shows how that works and how to switch it off.
Back toAll frameworks

Every expert started right here.