You're asked to build a model for Hindi legal Q&A. Walk through your end-to-end plan, from base model choice to deployment.
1. Scope and risk: who are the users (lawyers vs citizens)? What tasks (statute lookup, case summarisation, drafting)? This is high stakes: wrong legal advice causes harm, so design for citations, abstention and disclaimers.
2. Evaluation first
- 300–500 questions written and answered by legal experts, in Devanagari and romanised/Hinglish, with source citations.
- Metrics: correctness (expert-graded), citation accuracy, faithfulness, abstention on unanswerable questions, language quality.
3. Baselines
- Frontier API models + RAG over curated sources (Indian statutes, relevant judgments, with version/amendment metadata).
- Open multilingual models with good Indic tokenization (measure tokens per word, Q3).
- This often gets you most of the way. Measure the gap.
4. Close the gaps by type
| Gap | Fix |
|---|---|
| Missing or outdated legal knowledge | Better RAG corpus, chunking by section, hybrid search (section numbers) |
| Weak Hindi legal language / terminology | Continued pretraining on Hindi legal text (if a large corpus exists), or SFT |
| Format, citation style, abstention | SFT with LoRA on expert-curated examples; DPO on preference pairs |
| Retrieval misses Hindi queries | Multilingual / fine-tuned embeddings; query translation |
5. Training details: 5–20K high-quality SFT examples (expert-written + synthetic then expert-verified); keep general data in the mix to avoid forgetting (Q30); evaluate every checkpoint.
6. Deployment: RAG + fine-tuned generator, a citation verifier, "consult a lawyer" escalation, logging and an expert feedback loop, periodic re-indexing as laws change.
7. Cost decision: self-host (data residency, volume) vs API (quality, speed to market), decided on the eval and on cost per query.
Related
You understood something today that you didn't yesterday.