Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q2EasyConcept

What are the main types of evals, and when do you use each?

30-second answerSay your answer out loud first, then reveal.
Each grader type (code-based checks, LLM-as-judge, human review) can be used both offline on golden datasets and online on live traffic.
GraderStrengthsWeaknessesUse for
Code-basedFast, cheap, deterministicOnly for checkable propertiesFormat, JSON validity, required fields, numeric correctness, SQL execution results, tests passing
LLM-as-judgeScales, handles nuanceBiases, needs calibration, costs tokensFaithfulness, helpfulness, tone, policy adherence
HumanMost reliable for subjective / expert tasksSlow, expensive, inconsistent without guidelinesCalibrating judges, high-stakes domains, launch reviews

Rule of thumb: use code checks wherever possible, LLM judges where necessary, and humans to validate the judges and audit samples.

You understood something today that you didn't yesterday.