1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
How do you evaluate quality for multilingual and Indic-language users?
30-second answerSay your answer out loud first, then reveal.
What to measure per language slice
| Dimension | Example check |
|---|---|
| Task accuracy | Same rubric as English, graded by native speakers or calibrated judges |
| Language match | Reply in the user's language and script (or as configured) |
| Fluency and naturalness | Native-speaker rating; avoid literal translations |
| Terminology | Domain terms (banking, legal) correct in the local language, or kept in English where users prefer |
| Cross-lingual retrieval | Hindi query → English doc retrieval recall |
| Guardrails | Toxicity / PII detection recall in that language; injection in other scripts |
| Cost / latency | Token inflation for non-Latin scripts |
Practical tips
- Code-mixed inputs ("mera refund kab aayega?") are extremely common in India; make them a separate slice.
- LLM judges can be lenient or wrong in low-resource languages, so calibrate per language against native speakers.
- Keep slice sizes large enough for meaningful comparisons (Q19).
Related
Little by little, you're building something great.