1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
What are embeddings, and how is semantic similarity computed?
30-second answerSay your answer out loud first, then reveal.
Intuition: "How do I reset my password?" and "I forgot my login credentials" share almost no words but produce nearby vectors. Keyword search would miss the match; embeddings catch it.
How embedding models are trained: usually with contrastive learning. Pairs of related texts (question/answer, title/body) are pulled together, and unrelated texts are pushed apart. Examples: OpenAI text-embedding-3, Cohere Embed, Voyage, Google Gemini embeddings, and open models such as BGE, E5, GTE, Nomic and Jina.
Similarity metrics
- Cosine similarity = (A · B) / (‖A‖ ‖B‖). Ranges from −1 to 1 and ignores vector length.
- Dot product = A · B. Fast. Equals cosine if vectors are L2-normalised (most embedding APIs return normalised vectors).
- Euclidean (L2) distance. For normalised vectors, ranking by L2 equals ranking by cosine.
Use whatever metric the embedding model was trained with, which is usually documented.
Important properties
- Asymmetric retrieval: queries are short, documents long. Many models use different prefixes or modes for "query" and "document" (e.g. E5's
query:/passage:). Forgetting this measurably hurts retrieval. - Max input length: text beyond the model's token limit is truncated silently. Another reason to chunk.
- Vectors from different models are not comparable. Changing the model means re-embedding everything (Q45).
Common mistakes
- Saying "embeddings understand meaning perfectly." They struggle with negation ("not refundable"), exact identifiers (SKU-4471), numbers and rare jargon. This is why hybrid search exists (Q7).
Related
Slow is fine. Stopping is the only problem.