Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q28IntermediateScenario

Your users ask questions in Hindi, English and Hinglish, but most documents are in English. How do you design retrieval?

30-second answerSay your answer out loud first, then reveal.

Options

  1. Multilingual embeddings (cross-lingual): one index, and queries in any language match English docs. Simple, but quality varies by language pair and drops on code-mixed or romanised text.
  2. Query translation: translate or normalise the query to English with an LLM ("PF nikalne ka process kya hai?" → "What is the process to withdraw PF?") and search the English index. Often the most robust option, but it adds latency.
  3. Hybrid: search with both the original and the translated query, then fuse with RRF.
  4. Index translation: translate documents into other languages. Expensive, and doubles the index, but useful for key content.

Hinglish specifics

  • Romanised Hindi ("kitni chuttiyan milti hain") isn't well covered by many models' training data. An LLM normalisation step helps a lot.
  • BM25 won't match transliterated terms to English docs, so rely on translation for the keyword leg.

Generation: detect the user's language and respond in it (or as the user prefers), while keeping citations to the English sources.

Evaluation: separate eval slices for English, Hindi (Devanagari), and Hinglish (romanised). Track recall@k per slice. Averages hide failures in minority languages.

Slow is fine. Stopping is the only problem.