1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
How do you choose an embedding model for a RAG system?
30-second answerSay your answer out loud first, then reveal.
Criteria
- Quality on your data: leaderboards are averages over generic datasets. A model that ranks first overall may lose on legal Hindi text. Build 100–300 query→relevant-doc pairs and compare recall@10 / nDCG.
- Languages: multilingual and cross-lingual needs (query in Hindi, doc in English)? Pick models trained for that (Q28).
- Domain: code, legal, biomedical and finance have specialised models, or you can fine-tune (Q42).
- Max tokens: 512 vs 8K affects chunking freedom.
- Dimensions: 384 vs 1536 vs 3072. Higher can be better but costs more memory and is slower. Matryoshka models let you truncate dimensions (Q31).
- Latency / throughput: affects query time and bulk indexing time.
- Deployment: API (easy, but data leaves your network) vs self-hosted open models (control, privacy, GPU cost).
- Cost: per-token API cost for indexing millions of documents, plus re-indexing.
- Stability / versioning: API providers deprecate models. Plan for migrations.
Example decision. "For an Indian bank's internal docs in English and Hindi with strict data residency, I'd shortlist two or three open multilingual models (BGE-M3-class, multilingual E5-class), self-host them, and pick based on recall@10 on 200 real employee queries, with dimension size as the tie-breaker."
Common mistakes
- Choosing by leaderboard rank alone.
- Mixing vectors from two models in one index.
Related
You understood something today that you didn't yesterday.