DeepEvaldeepeval 4.2 · Python 3.9+
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
23 small wins to finish your pathNext lesson

LLM as a judge quiz

Question 1 of 5

Your G-Eval criteria say "gives the same shipping news as the expected answer", evaluation_params has only ACTUAL_OUTPUT, and a wrong answer scores 1.0. Why?

Getting one wrong here is cheaper than getting it wrong in your own code.