1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
What are token embeddings and positional encodings? Why does a transformer need position information?
30-second answerSay your answer out loud first, then reveal.
Token embeddings
- An embedding matrix of shape (vocab_size × d_model), e.g. 128K × 4096.
- Looking up row
token_idgives the token's vector, which is learned during training. - Many models tie it with the output projection ("weight tying") to save parameters.
Positional encoding schemes
| Scheme | How | Notes |
|---|---|---|
| Sinusoidal (Vaswani et al. 2017) | Fixed sin/cos of different frequencies added to embeddings | No learned params; limited extrapolation |
| Learned absolute | A trainable vector per position (GPT-2, BERT) | Can't go beyond trained max length |
| RoPE (rotary) | Rotates query/key vectors by an angle proportional to position; attention then depends on relative distance | Used by Llama, Mistral, Qwen and others; can be extended (Q40) |
| ALiBi | Adds a distance-based penalty to attention scores | Good length extrapolation |
Why RoPE is popular: relative position falls out naturally from the rotation maths (the dot product of rotated q and k depends on the position difference), it adds no parameters, and there are good techniques for extending context length.
Follow-ups to expect
- How do models extend context from 8K to 128K? (Q40.)
Related
This is what real progress feels like.