1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
How do you evaluate LLMs used for classification tasks (intent detection, routing, tagging)?
30-second answerSay your answer out loud first, then reveal.
Example confusion insight
| Actual \ Predicted | Billing | Refund | Tech | Other |
|---|---|---|---|---|
| Billing | 90 | 8 | 0 | 2 |
| Refund | 12 | 83 | 1 | 4 |
| Tech | 1 | 0 | 95 | 4 |
Billing ↔ Refund confusion suggests overlapping definitions. Fix the label descriptions and examples, or merge the classes if downstream handling is similar.
Metrics to know
- Precision: of items predicted X, how many are X (cost of false alarms).
- Recall: of actual X, how many were found (cost of misses).
- Macro vs micro averages: macro treats classes equally (important for rare but critical classes such as "fraud" or "complaint").
- Calibration: does "0.9 confidence" mean right 90% of the time? LLM self-reported confidence is often poorly calibrated, so use logprobs or empirical calibration.
LLM-specific checks: output always in the allowed label set (constrained decoding or validation); stability across prompt paraphrases; the cost and latency vs a fine-tuned small classifier.
Related
Slow is fine. Stopping is the only problem.