1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
Your users ask questions in Hindi, English and Hinglish, but most documents are in English. How do you design retrieval?
30-second answerSay your answer out loud first, then reveal.
Options
- Multilingual embeddings (cross-lingual): one index, and queries in any language match English docs. Simple, but quality varies by language pair and drops on code-mixed or romanised text.
- Query translation: translate or normalise the query to English with an LLM ("PF nikalne ka process kya hai?" → "What is the process to withdraw PF?") and search the English index. Often the most robust option, but it adds latency.
- Hybrid: search with both the original and the translated query, then fuse with RRF.
- Index translation: translate documents into other languages. Expensive, and doubles the index, but useful for key content.
Hinglish specifics
- Romanised Hindi ("kitni chuttiyan milti hain") isn't well covered by many models' training data. An LLM normalisation step helps a lot.
- BM25 won't match transliterated terms to English docs, so rely on translation for the keyword leg.
Generation: detect the user's language and respond in it (or as the user prefers), while keeping citations to the English sources.
Evaluation: separate eval slices for English, Hindi (Devanagari), and Hinglish (romanised). Track recall@k per slice. Averages hide failures in minority languages.
Related
Slow is fine. Stopping is the only problem.