Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →

Q21IntermediateConcept

How do you evaluate multi-turn conversations?

30-second answerSay your answer out loud first, then reveal.

Approaches

ApproachHowGood for
Next-turn eval on real historiesFeed the first N turns from logs, evaluate the bot's next replyContext handling, regression on real data
User simulatorLLM plays a user with persona, goal, hidden info, patience levelEnd-to-end task success, many variations
Scripted scenariosFixed user messages with assertions per turnDeterministic policy tests (e.g. must ask for verification)
Conversation-level judgeJudge reads the full transcript against a rubricOverall helpfulness, tone, compliance

Conversation-level criteria examples

  • Goal achieved (verify via backend state where possible: was the ticket created correctly?).
  • Asked necessary clarifying questions; didn't ask redundant ones.
  • Consistent facts across turns; no contradictions.
  • Followed escalation rules; respected refusals or boundaries.
  • Number of turns to resolution (efficiency).

User simulator tips: give the simulator a hidden goal and constraints ("you want a refund but the order is 40 days old"); vary personas; cap turns; and validate that simulator behaviour is realistic by comparing with real conversations.

Little by little, you're building something great.