DeepEvaldeepeval 4.2.8 · Python 3.10+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
34 small wins to finish your pathNext lesson →

LLM as a judge quiz

Question 1 of 5

A G-Eval metric built with only <code>criteria</code> scored a correct paraphrase of the return policy far below 1. What fixed it?

Getting one wrong here is cheaper than getting it wrong in your own code.