Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q31IntermediateConcept

Explain vector quantization and Matryoshka embeddings. What are the trade-offs?

30-second answerSay your answer out loud first, then reveal.

Quantization types

TypeSize vs float32Accuracy impactNotes
float162x smallerNegligibleEasy win
int8 (scalar)4x smallerSmallWidely supported
Binary32x smallerNoticeable aloneUse Hamming distance for candidates, then rescore with full vectors
Product quantization (PQ)8–64xModerateClassic in FAISS / IVF-PQ

Rescoring pattern

  1. Search the quantized index for the top 200 candidates (fast, in RAM).
  2. Load the full-precision vectors for those 200 (from disk) and recompute exact similarity.
  3. Return the top 10. This recovers most of the lost accuracy.

Matryoshka Representation Learning (MRL)

  • Training puts the most important information in the leading dimensions.
  • Truncate to 256 or 512 dims for a cheap first-stage search; use full dims for rescoring (an adaptive retrieval funnel).
  • Supported by several modern embedding models (e.g. OpenAI text-embedding-3 via its dimensions parameter, plus several open models).

When to use: large corpora (tens of millions of vectors and up), memory-constrained deployments, cost pressure. Always measure the recall drop on your eval set.

Little by little, you're building something great.