Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q4EasyConcept

What are token embeddings and positional encodings? Why does a transformer need position information?

30-second answerSay your answer out loud first, then reveal.

Token embeddings

  • An embedding matrix of shape (vocab_size × d_model), e.g. 128K × 4096.
  • Looking up row token_id gives the token's vector, which is learned during training.
  • Many models tie it with the output projection ("weight tying") to save parameters.

Positional encoding schemes

SchemeHowNotes
Sinusoidal (Vaswani et al. 2017)Fixed sin/cos of different frequencies added to embeddingsNo learned params; limited extrapolation
Learned absoluteA trainable vector per position (GPT-2, BERT)Can't go beyond trained max length
RoPE (rotary)Rotates query/key vectors by an angle proportional to position; attention then depends on relative distanceUsed by Llama, Mistral, Qwen and others; can be extended (Q40)
ALiBiAdds a distance-based penalty to attention scoresGood length extrapolation

Why RoPE is popular: relative position falls out naturally from the rotation maths (the dot product of rotated q and k depends on the position difference), it adds no parameters, and there are good techniques for extending context length.

Follow-ups to expect

  • How do models extend context from 8K to 128K? (Q40.)

This is what real progress feels like.