Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q3EasyConcept

What are embeddings, and how is semantic similarity computed?

30-second answerSay your answer out loud first, then reveal.

Intuition: "How do I reset my password?" and "I forgot my login credentials" share almost no words but produce nearby vectors. Keyword search would miss the match; embeddings catch it.

How embedding models are trained: usually with contrastive learning. Pairs of related texts (question/answer, title/body) are pulled together, and unrelated texts are pushed apart. Examples: OpenAI text-embedding-3, Cohere Embed, Voyage, Google Gemini embeddings, and open models such as BGE, E5, GTE, Nomic and Jina.

Similarity metrics

  • Cosine similarity = (A · B) / (‖A‖ ‖B‖). Ranges from −1 to 1 and ignores vector length.
  • Dot product = A · B. Fast. Equals cosine if vectors are L2-normalised (most embedding APIs return normalised vectors).
  • Euclidean (L2) distance. For normalised vectors, ranking by L2 equals ranking by cosine.

Use whatever metric the embedding model was trained with, which is usually documented.

Important properties

  • Asymmetric retrieval: queries are short, documents long. Many models use different prefixes or modes for "query" and "document" (e.g. E5's query: / passage:). Forgetting this measurably hurts retrieval.
  • Max input length: text beyond the model's token limit is truncated silently. Another reason to chunk.
  • Vectors from different models are not comparable. Changing the model means re-embedding everything (Q45).

Common mistakes

  • Saying "embeddings understand meaning perfectly." They struggle with negation ("not refundable"), exact identifiers (SKU-4471), numbers and rare jargon. This is why hybrid search exists (Q7).

Slow is fine. Stopping is the only problem.