Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q6EasyConcept

Why is building an eval set with the customer so important, and how do you do it?

30-second answerSay your answer out loud first, then reveal.

Process

  1. Sample real data: 100–300 cases stratified by type, difficulty, language and source. Include known-hard cases and examples the AI should refuse or escalate.
  2. Define the rubric with SMEs: e.g. for summaries: correct facts, no critical omissions, correct recommendation, appropriate tone. Make the criteria binary or 1–3 scales with examples.
  3. Label: SMEs write or approve reference outputs; for extraction, field-level ground truth.
  4. Calibrate: two SMEs label an overlapping subset. If they disagree often, the task definition is unclear, so fix the definition first.
  5. Automate scoring where possible: exact matches for fields, an LLM judge calibrated against SME grades for free text.
  6. Version and grow: add production failures; keep a held-out portion to avoid overfitting prompts.

Why customers value it: it gives them confidence, it's an asset they keep (useful for vendor comparisons and future upgrades), and it gives SMEs ownership of quality.

Practical tips: SME time is scarce, so make labelling easy (a simple UI or spreadsheet), batch sessions, and show them how their labels improved the system.

Little by little, you're building something great.