Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q47HardScenario

Users ask "How many of our vendor contracts have an auto-renewal clause?" and RAG gives wrong answers. Why, and what would you build?

30-second answerSay your answer out loud first, then reveal.

Why RAG fails

  • Retrieval returns the top 10 of maybe 400 relevant contracts, so the LLM counts what it sees.
  • Even 1M-token contexts can't hold thousands of contracts reliably, and counting over long context is error-prone.

Solution: extract, then query

Extract, then query: contracts go through batch LLM extraction to a schema, validation with human spot-checks, and into a structured table with source span links; a router sends aggregate, filter and list-all questions to text-to-SQL over the table and specific-clause questions to RAG over the contract text, and both return answers with evidence links.

Design details

  1. Schema design with the business: which fields matter (renewal, termination, liability cap, governing law, payment terms).
  2. Extraction quality: structured outputs; store a supporting quote and page for every field; confidence scores; human review of low-confidence items; evaluate on a labelled sample (field-level precision and recall).
  3. Incremental updates: extract fields when a new contract arrives.
  4. Query routing: detect aggregation intent ("how many", "list all", "average", "which ones").
  5. Answers with evidence: "142 contracts have auto-renewal" plus a downloadable list with links to the exact clauses.

General lesson: RAG is for finding and reading; databases are for counting and filtering. Convert unstructured data to structured data when the questions are analytical.

You understood something today that you didn't yesterday.