Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q24IntermediateConcept

How do you handle follow-up questions in a conversational RAG chatbot?

30-second answerSay your answer out loud first, then reveal.
The chat history and the follow-up go to a condense LLM that writes a standalone question, which is used for retrieval, and generation uses the retrieved context plus the follow-up and recent history.

Design points

  1. Query condensation: cheap, fast model; few-shot examples; keep the user's language.
  2. Retrieval decision: classify whether new retrieval is needed (new information need) or the turn is chit-chat or a formatting request on the previous answer.
  3. History management: keep the last N turns verbatim plus a summary of older turns; don't stuff full past contexts back in.
  4. Context carry-over: sometimes reuse previously retrieved chunks plus new ones (e.g. "explain point 3 more").
  5. Entity tracking: "he", "that product" resolved via condensation.

Failure modes

  • Condensation drops important constraints ("for India" from three turns ago).
  • Topic switches incorrectly blended with old context.
  • Evaluate on multi-turn test conversations, not only single questions.

This is what real progress feels like.