Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q23IntermediateConcept

How do you handle tables and structured data in RAG?

30-second answerSay your answer out loud first, then reveal.
A router sends unstructured questions to vector or hybrid RAG and aggregation or filter questions to text-to-SQL on a read-only database with validation, and both paths end in an LLM answer that shows sources or SQL.

For tables inside documents

  1. Extract with structure (layout-aware parsers); store as Markdown or HTML so rows and columns survive.
  2. Keep headers with data. If you split a large table, repeat the header row in each piece.
  3. Embed a summary, return the table: an LLM describes the table ("Quarterly revenue by region for FY2026, columns: ...") and that description gets embedded. The full table is passed to the LLM on retrieval (a multi-vector pattern).
  4. Row-level chunks for lookup tables: each row becomes "Product: X, Price: Y, Warranty: Z".

For databases / large structured data: text-to-SQL

  • Provide the schema with column descriptions and example values; few-shot examples.
  • Safety: read-only credentials, query validation, row limits, timeouts, no DDL/DML.
  • Show the generated SQL or the source rows for transparency.

Why vector search fails here: "What's the average order value in Maharashtra last month?" needs filtering and arithmetic over thousands of rows. No top-k retrieval of chunks can compute that correctly.

Slow is fine. Stopping is the only problem.