Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q13EasyConcept

How do you evaluate LLMs used for classification tasks (intent detection, routing, tagging)?

30-second answerSay your answer out loud first, then reveal.

Example confusion insight

Actual \ PredictedBillingRefundTechOther
Billing90802
Refund128314
Tech10954

Billing ↔ Refund confusion suggests overlapping definitions. Fix the label descriptions and examples, or merge the classes if downstream handling is similar.

Metrics to know

  • Precision: of items predicted X, how many are X (cost of false alarms).
  • Recall: of actual X, how many were found (cost of misses).
  • Macro vs micro averages: macro treats classes equally (important for rare but critical classes such as "fraud" or "complaint").
  • Calibration: does "0.9 confidence" mean right 90% of the time? LLM self-reported confidence is often poorly calibrated, so use logprobs or empirical calibration.

LLM-specific checks: output always in the allowed label set (constrained decoding or validation); stability across prompt paraphrases; the cost and latency vs a fine-tuned small classifier.

Slow is fine. Stopping is the only problem.