LLM Fundamentals overview
A large language model (LLM) is a neural network that reads text as tokens and writes its answer one token at a time, each token chosen from probabilities it learned from a very large amount of text.
Last updated: 30 Sep, 2026 · groq 1.7 · gpt-oss-120b on Groq
Almost every current LLM is a transformer, the architecture from the 2017 paper Attention Is All You Need. Chat models such as GPT, Claude, Gemini, Llama and Qwen are first trained to predict the next token across a huge amount of text, then tuned to follow instructions and hold a conversation.
You reach a model in one of two ways. A hosted model runs on a provider's computers, and you call it over an API with a key. An open-weight model publishes its weights, the numbers it learned, so anyone can download it, and providers can host it for you.
Pick one to watch it run, step by step.
gpt-oss-120b on Groq
Every example calls gpt-oss-120b, an open-weight model that OpenAI released in August 2025 under the Apache 2.0 licence. Groq hosts it, and Groq's free plan needs only an account and an API key. The code uses Groq's Python library, groq, to send requests, tiktoken, OpenAI's tokenizer library, to count tokens, and pydantic to check answers.
The ideas do not depend on Groq or on this model. Tokens, sampling settings, chat messages, context windows and prices work the same way with OpenAI, Anthropic and Gemini; the client library and the model name are what change.
Tokens, sampling, prompts and limits
| Part | What you run | Lessons |
|---|---|---|
| How a model writes | Count tokens, turn scores into probabilities, stream an answer | Tokens, Next-token prediction, Generating text |
| Controlling the output | Change temperature, top_p, the length cap and the seed, and compare answers | Temperature, top_p and top_k, Seeds |
| Prompts | Write a system prompt and examples, check and score the answers | System prompts, Few-shot prompting, Evaluating prompts |
| In production | Measure the context window, the cost and the time of a request | Context window, Cost, Latency |
The ticket sorter
One task runs through the prompt lessons: sort an online shop's support tickets into billing, shipping or other, with a priority, as one line of JSON. It starts as a vague instruction in System prompts, fails in ways you can see, and ends in Project: ticket sorter as a small program that reports how many tickets it got right, how many tokens that took, what it cost and how long it ran.
Python and a free Groq key
- Python 3.10 or later, with functions, lists, dictionaries and Pydantic, as in Python for AI.
- A free Groq API key, created in Installation and setup.
Related
- Next: Installation and setup
- Reference: Groq documentation
- Reference: Introducing gpt-oss
Every expert started right here.