1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
Explain vector quantization and Matryoshka embeddings. What are the trade-offs?
30-second answerSay your answer out loud first, then reveal.
Quantization types
| Type | Size vs float32 | Accuracy impact | Notes |
|---|---|---|---|
| float16 | 2x smaller | Negligible | Easy win |
| int8 (scalar) | 4x smaller | Small | Widely supported |
| Binary | 32x smaller | Noticeable alone | Use Hamming distance for candidates, then rescore with full vectors |
| Product quantization (PQ) | 8–64x | Moderate | Classic in FAISS / IVF-PQ |
Rescoring pattern
- Search the quantized index for the top 200 candidates (fast, in RAM).
- Load the full-precision vectors for those 200 (from disk) and recompute exact similarity.
- Return the top 10. This recovers most of the lost accuracy.
Matryoshka Representation Learning (MRL)
- Training puts the most important information in the leading dimensions.
- Truncate to 256 or 512 dims for a cheap first-stage search; use full dims for rescoring (an adaptive retrieval funnel).
- Supported by several modern embedding models (e.g. OpenAI text-embedding-3 via its
dimensionsparameter, plus several open models).
When to use: large corpora (tens of millions of vectors and up), memory-constrained deployments, cost pressure. Always measure the recall drop on your eval set.
Related
Little by little, you're building something great.