1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
51 small wins to finish your pathNext question →
How do you evaluate multi-turn conversations?
30-second answerSay your answer out loud first, then reveal.
Approaches
| Approach | How | Good for |
|---|---|---|
| Next-turn eval on real histories | Feed the first N turns from logs, evaluate the bot's next reply | Context handling, regression on real data |
| User simulator | LLM plays a user with persona, goal, hidden info, patience level | End-to-end task success, many variations |
| Scripted scenarios | Fixed user messages with assertions per turn | Deterministic policy tests (e.g. must ask for verification) |
| Conversation-level judge | Judge reads the full transcript against a rubric | Overall helpfulness, tone, compliance |
Conversation-level criteria examples
- Goal achieved (verify via backend state where possible: was the ticket created correctly?).
- Asked necessary clarifying questions; didn't ask redundant ones.
- Consistent facts across turns; no contradictions.
- Followed escalation rules; respected refusals or boundaries.
- Number of turns to resolution (efficiency).
User simulator tips: give the simulator a hidden goal and constraints ("you want a refund but the order is 40 days old"); vary personas; cap turns; and validate that simulator behaviour is realistic by comparing with real conversations.
Related
Little by little, you're building something great.