Multi-turn evaluation
Multi-turn evaluation is the scoring of a whole conversation instead of one answer: a ConversationalTestCase holds the turns, and conversational metrics judge each assistant reply against everything said before it.
Last updated: 05 Oct, 2026 · DeepEval 4.2.8
Every test case so far had one question and one answer, even the agent's traces. A support chat runs over several messages, and a reply can be right on its own and still wrong for the conversation, such as asking for an order number the customer gave two messages earlier.
Keeping the conversation in a list
The video starts from one fact: LLMs are stateless. Each API request knows nothing about the one before, not even what the model said a second ago. The simplest answer is conversation buffer memory: a list of messages that grows every turn. On turn one the request carries the system prompt and the first user message. On turn two it carries the system prompt, the first user message, the assistant's reply and the second user message. Each message has a role: system, user or assistant.
The video builds this memory by hand in plain Python for an agent memory module. The TechNest chat below does the same with its own messages, then lets DeepEval check whether the replies use what the list holds.
The ConversationalTestCase API
from deepeval.test_case import ConversationalTestCase, Turn
test_case = ConversationalTestCase(turns=[
Turn(role="user", content="Hi, I'm Priya. I ordered a ProBook X1 last week, order TN-1002."),
Turn(role="assistant", content="...", retrieval_context=["..."]), # the chunks this reply used
Turn(role="user", content="What was my order number again?"),
Turn(role="assistant", content="..."),
])
KnowledgeRetentionMetric(model=judge).measure(test_case)A Turn is one message with a role, "user" or "assistant", and its content. The docs say retrieval_context and tools_called belong only on assistant turns. Conversational metrics take a ConversationalTestCase, not an LLMTestCase.
Answering with the earlier turns
def reply(history, message):
"""Answer one message, given the earlier turns of the conversation."""
contexts = retrieve(" ".join([turn.content for turn in history if turn.role == "user"] + [message]))
context_block = "\n\n".join(f"[{i+1}] {c}" for i, c in enumerate(contexts))
messages = [{"role": "system", "content": CHAT_PROMPT}]
messages += [{"role": turn.role, "content": turn.content} for turn in history]
messages.append({"role": "user", "content": f"Context:\n{context_block}\n\nCustomer: {message}"})
response = groq.chat.completions.create(model=CHAT_MODEL, messages=messages, temperature=0)
return response.choices[0].message.content.strip(), contextsThis is the buffer from the video. The earlier turns go into the request between the system prompt and the new message, and the search also uses the earlier user messages, so a follow-up such as "How much RAM does it have?" still finds the ProBook X1.
Running a conversation with or without memory
def run_conversation(user_messages, memory=True):
"""Send the messages one by one; without memory, every reply sees only the newest message."""
turns = []
for message in user_messages:
response, contexts = reply(turns if memory else [], message)
turns.append(Turn(role="user", content=message))
turns.append(Turn(role="assistant", content=response, retrieval_context=contexts))
return turnsThe list of turns is the bot's memory and the test case at once. With memory=False every reply gets an empty history, the stateless model the video starts from.
- written in Custom judge model
- written in TechNest RAG app
- written in TechNest RAG app
View the code here
import os
from deepeval.models import DeepEvalBaseLLM
from openai import AsyncOpenAI, OpenAI
GROQ_URL = "https://api.groq.com/openai/v1"
class GroqJudge(DeepEvalBaseLLM):
"""A DeepEval judge model that runs on Groq."""
def __init__(self, model="openai/gpt-oss-120b"):
self.model_name = model
key = os.environ["GROQ_API_KEY"]
# on a 429 (rate limit) the client waits and tries again, up to 8 times
self.client = OpenAI(api_key=key, base_url=GROQ_URL, max_retries=8)
self.async_client = AsyncOpenAI(api_key=key, base_url=GROQ_URL, max_retries=8)
def load_model(self):
return self.client
def get_model_name(self):
return self.model_name
def request(self, prompt, schema):
request = {"model": self.model_name, "messages": [{"role": "user", "content": prompt}], "temperature": 0}
if schema is not None:
# ask Groq for JSON in the shape of the metric's Pydantic schema
json_schema = {"name": schema.__name__, "schema": schema.model_json_schema()}
request["response_format"] = {"type": "json_schema", "json_schema": json_schema}
return request
def generate(self, prompt, schema=None):
reply = self.client.chat.completions.create(**self.request(prompt, schema))
text = reply.choices[0].message.content
return schema.model_validate_json(text) if schema else text
async def a_generate(self, prompt, schema=None):
reply = await self.async_client.chat.completions.create(**self.request(prompt, schema))
text = reply.choices[0].message.content
return schema.model_validate_json(text) if schema else text
judge = GroqJudge(os.environ.get("JUDGE_MODEL", "openai/gpt-oss-120b"))
import json
import os
import re
from openai import OpenAI
groq = OpenAI(api_key=os.environ["GROQ_API_KEY"], base_url="https://api.groq.com/openai/v1")
CHAT_MODEL = "qwen/qwen3.8-27b"
SYSTEM_PROMPT = """You are a helpful customer support assistant for TechNest, an online electronics store.
Answer the customer's question using ONLY the information provided in the context below.
If the context does not contain enough information to answer fully, say so honestly.
Keep your answer concise, factual, and friendly. Do not invent any details not present in the context.
Reply in two or three plain sentences, with no lists or tables."""
with open("catalog.json", encoding="utf-8") as f:
CATALOG = json.load(f)
SKIP = {"a", "an", "and", "are", "can", "do", "does", "for", "how", "i", "in", "is", "it",
"long", "much", "my", "of", "on", "s", "technest", "the", "to", "what", "with", "you", "your"}
def words(text):
"""The words in a text that carry meaning, in lower case."""
return {w for w in re.findall(r"[a-z0-9]+", text.lower()) if w not in SKIP}
def retrieve(question, top_k=3):
asked = words(question)
ranked = sorted(CATALOG, key=lambda item: len(asked & words(item["title"] + " " + item["content"])), reverse=True)
return [item["content"] for item in ranked[:top_k]]
def generate(question, contexts):
context_block = "\n\n".join(f"[{i+1}] {c}" for i, c in enumerate(contexts))
messages = [
{"role": "system", "content": SYSTEM_PROMPT},
{"role": "user", "content": f"Context:\n{context_block}\n\nCustomer question: {question}"},
]
response = groq.chat.completions.create(model=CHAT_MODEL, messages=messages, temperature=0)
return response.choices[0].message.content.strip()
def answer(question, top_k=3):
contexts = retrieve(question, top_k)
return generate(question, contexts), contexts
[
{
"id": "prod_001",
"category": "product",
"title": "ProBook X1 Laptop",
"content": "The TechNest ProBook X1 is a 14-inch laptop featuring an Intel Core i7-13th Gen processor, 16GB DDR5 RAM, and a 512GB NVMe SSD. It has a battery life of 12 hours, weighs 1.4kg, and comes with a backlit keyboard. Price: $1,299. Includes a 2-year manufacturer warranty."
},
{
"id": "prod_002",
"category": "product",
"title": "PixelPhone 15",
"content": "The TechNest PixelPhone 15 is a 6.7-inch AMOLED smartphone with a 50MP triple camera system, 8GB RAM, 256GB storage, and a 5,000mAh battery supporting 65W fast charging. Available in Midnight Black and Arctic White. Price: $899. Includes a 1-year warranty."
},
{
"id": "prod_003",
"category": "product",
"title": "SoundPods Pro",
"content": "The TechNest SoundPods Pro are true wireless earbuds with active noise cancellation (ANC), 8 hours of playback per charge plus 24 hours with the case, and IPX4 water resistance. They connect via Bluetooth 5.3 and support multipoint pairing with two devices simultaneously. Price: $149."
},
{
"id": "prod_004",
"category": "product",
"title": "UltraTab S2 Tablet",
"content": "The TechNest UltraTab S2 is a 11-inch tablet powered by a Snapdragon 870 processor with 8GB RAM and 128GB storage expandable via microSD. It features a 120Hz display, a 7,500mAh battery, and supports the TechNest Stylus Pen sold separately. Price: $549. Includes a 1-year warranty."
},
{
"id": "prod_005",
"category": "product",
"title": "SmartWatch X",
"content": "The TechNest SmartWatch X features continuous heart rate monitoring, SpO2 tracking, GPS, and 7-day battery life. It is water-resistant up to 50 metres. Compatible with both Android and iOS. Price: $299. Includes a 1-year warranty and a free extra silicone band."
},
{
"id": "prod_006",
"category": "product",
"title": "ProCam 4K Action Camera",
"content": "The TechNest ProCam 4K shoots 4K video at 60fps and 20MP photos. It is waterproof up to 10 metres without a case, has built-in image stabilisation (EIS), and includes a touch screen. Battery life is 90 minutes of 4K recording. Price: $229. Includes a 1-year warranty."
},
{
"id": "prod_007",
"category": "product",
"title": "BassBuds Max Headphones",
"content": "The TechNest BassBuds Max are over-ear wireless headphones with 40-hour battery life, hybrid active noise cancellation, and a premium 40mm driver for deep bass. They fold flat for travel and include a carrying case. Price: $199. Compatible with all Bluetooth devices."
},
{
"id": "prod_008",
"category": "product",
"title": "SoundBar 360",
"content": "The TechNest SoundBar 360 is a 2.1 soundbar with a 120W output, built-in subwoofer, Dolby Atmos support, and HDMI ARC connectivity. It also supports Bluetooth streaming and has an optical audio input. Dimensions: 90cm wide. Price: $349. Includes a 2-year warranty."
},
{
"id": "policy_001",
"category": "policy",
"title": "Return Policy",
"content": "TechNest accepts returns within 30 days of the original purchase date. Items must be in their original packaging with all accessories included. Customers are responsible for return shipping costs unless the item arrives defective or damaged. Refunds are processed within 5 to 7 business days of receiving the returned item. Digital downloads and opened software are non-refundable."
},
{
"id": "policy_002",
"category": "policy",
"title": "Shipping Policy",
"content": "TechNest offers free standard shipping on all orders over $50 within the continental US. Standard shipping takes 3 to 5 business days. Expedited shipping (1 to 2 business days) is available for $9.99. Same-day delivery is available in select cities for $19.99. Orders placed before 2pm local time are dispatched the same day."
},
{
"id": "policy_003",
"category": "policy",
"title": "Warranty Policy",
"content": "All TechNest products include a minimum 1-year manufacturer warranty covering defects in materials and workmanship. The ProBook X1 and SoundBar 360 include a 2-year warranty. Warranty does not cover physical damage, water damage (unless the product is rated waterproof), or damage from unauthorised modifications. To make a warranty claim, contact support@technest.com with your order number and a description of the issue."
},
{
"id": "policy_004",
"category": "policy",
"title": "Payment Policy",
"content": "TechNest accepts Visa, Mastercard, American Express, PayPal, and Apple Pay. All transactions are encrypted using 256-bit SSL. Buy Now Pay Later is available via Klarna for orders over $100, with 0% interest for 3 monthly instalments. TechNest does not store full card details — payments are processed securely by Stripe."
},
{
"id": "faq_001",
"category": "faq",
"title": "Order Tracking",
"content": "To track your order, visit technest.com/orders and enter your order number and email address. A shipping confirmation email with a tracking link is sent within 24 hours of dispatch. If you have not received your tracking email after 48 hours, check your spam folder or contact support@technest.com."
},
{
"id": "faq_002",
"category": "faq",
"title": "International Shipping",
"content": "TechNest ships to over 40 countries. International shipping rates start at $14.99 and delivery takes 7 to 14 business days. Import duties and taxes are the responsibility of the customer and are not included in the product price. Free shipping promotions apply to US orders only."
},
{
"id": "faq_003",
"category": "faq",
"title": "Bulk and Business Orders",
"content": "TechNest offers volume discounts for businesses purchasing 10 or more units of any single product. Discounts range from 10% for 10 to 49 units up to 25% for 100 or more units. Contact business@technest.com with your requirements for a custom quote. A dedicated account manager is assigned for orders over $10,000."
}
]
The chat.py file
Save the chat as chat.py, next to technest.py and catalog.json. Its prompt says where each kind of fact comes from: the context for products and policies, the conversation for what the customer said.
from deepeval.test_case import Turn
from technest import CHAT_MODEL, groq, retrieve
CHAT_PROMPT = """You are a helpful customer support assistant for TechNest, an online electronics store.
Use the context for facts about products and policies, and the conversation for what the customer told you.
If neither has the answer, ask the customer for what you need. Reply in one or two plain sentences."""
def reply(history, message):
"""Answer one message, given the earlier turns of the conversation."""
contexts = retrieve(" ".join([turn.content for turn in history if turn.role == "user"] + [message]))
context_block = "\n\n".join(f"[{i+1}] {c}" for i, c in enumerate(contexts))
messages = [{"role": "system", "content": CHAT_PROMPT}]
messages += [{"role": turn.role, "content": turn.content} for turn in history]
messages.append({"role": "user", "content": f"Context:\n{context_block}\n\nCustomer: {message}"})
response = groq.chat.completions.create(model=CHAT_MODEL, messages=messages, temperature=0)
return response.choices[0].message.content.strip(), contexts
def run_conversation(user_messages, memory=True):
"""Send the messages one by one; without memory, every reply sees only the newest message."""
turns = []
for message in user_messages:
response, contexts = reply(turns if memory else [], message)
turns.append(Turn(role="user", content=message))
turns.append(Turn(role="assistant", content=response, retrieval_context=contexts))
return turnsScoring a conversation with and without history
from deepeval.metrics import KnowledgeRetentionMetric, TurnRelevancyMetric
from deepeval.test_case import ConversationalTestCase
from chat import run_conversation
from judge import judge
messages = [
"Hi, I'm Priya. I ordered a ProBook X1 last week, order TN-1002.",
"How much RAM and storage does it have?",
"What was my order number again?",
]
metrics = [KnowledgeRetentionMetric(model=judge), TurnRelevancyMetric(model=judge)]
for memory in (True, False):
turns = run_conversation(messages, memory=memory)
print("with history" if memory else "without history")
for turn in turns[1::2]: # the assistant's replies
print(" bot:", turn.content)
test_case = ConversationalTestCase(turns=turns)
for metric in metrics:
metric.measure(test_case)
print(f" {metric.__name__}: {metric.score:.2f} {metric.reason}")with history bot: Hi Priya, thanks for reaching out. How can I help you with your ProBook X1 order? bot: The ProBook X1 comes with 16GB of DDR5 RAM and a 512GB NVMe SSD. bot: Your order number is TN-1002. Knowledge Retention: 1.00 The score is 1.00 because there are no attritions indicating any forgetfulness or loss of previously established knowledge. Turn Relevancy: 1.00 The score is 1.0 because there are no irrelevant messages in the conversation. without history bot: Hi Priya, thanks for reaching out. How can I help you with your ProBook X1 order? bot: I need to know which product you are asking about, as the RAM and storage vary between the PixelPhone 15, UltraTab S2, and ProBook X1. bot: I don't have access to your personal order details, so I cannot retrieve your order number. Please check your confirmation email or log in to your account at technest.com/orders to find it. Knowledge Retention: 0.33 The score is 0.33 because the LLM forgot previously established details, asking again about the known product (ProBook X1) and claiming it cannot retrieve the known order number (TN-1002). Turn Relevancy: 1.00 The score is 1.0 because there are no irrelevant messages; no messages were identified as irrelevant.
What the bot forgot without history
- With history, the second reply gives the ProBook X1's 16GB of RAM and 512GB SSD although the message only says "it", and the third answers
TN-1002. Knowledge retention and turn relevancy both score 1.00. - Without history, the first reply is the same, since the first message carries everything it needs. The second asks which product Priya means, and the third says it has no access to her order details, although she gave the number two messages earlier.
- Knowledge retention drops to 0.33. The docs define it as assistant turns without a forgotten fact over all assistant turns: only the first of the three replies kept what Priya said. The judge's reason names both slips, the product and TN-1002.
- Turn relevancy stays at 1.00 in both runs. Every forgetful reply was still on the topic of the conversation, so this metric alone would have passed the bot without memory.
Simulating a customer
Writing every customer message by hand limits a test to the conversations you thought of. DeepEval's ConversationSimulator lets a model play the customer: it reads a scenario, writes each next customer message, sends it to your bot through a callback, and keeps going for a fixed number of turns. The result is a ConversationalTestCase that the same metrics score.
The scenario as a ConversationalGolden
golden = ConversationalGolden(
scenario="A TechNest customer's order TN-1042 has not arrived. ...", # what the simulated customer wants
expected_outcome="The assistant repeats order number TN-1042.",
user_description="A polite customer who writes short messages.",
)The bot as a callback
The simulator calls an async function with the customer's new message and the turns so far, which already end with that message. Passing turns[:-1] as the history gives the bot memory; passing an empty list takes it away.
async def model_callback(input, turns):
history = turns[:-1] if memory else [] # turns ends with the customer's new message
response, contexts = reply(history, input)
return Turn(role="assistant", content=response, retrieval_context=contexts)A fixed number of turns
By default the simulator stops once the judge decides expected_outcome was met. A stopping_controller that always returns proceed() leaves only max_user_simulations to end the conversation, so both runs get three customer messages. The simulator's own model plays the customer; here it is the Groq judge.
Simulating the TN-1042 customer with and without history
from deepeval.dataset import ConversationalGolden
from deepeval.metrics import KnowledgeRetentionMetric
from deepeval.simulator import ConversationSimulator
from deepeval.simulator.controller import proceed
from deepeval.test_case import Turn
from chat import reply
from judge import judge
golden = ConversationalGolden(
scenario=("A TechNest customer's order TN-1042 has not arrived. They give the order number in their first "
"message, ask about shipping times, and in their last message ask what their order number was, without saying it again."),
expected_outcome="The assistant repeats order number TN-1042.",
user_description="A polite customer who writes short messages.",
)
def callback_for(memory):
async def model_callback(input, turns):
history = turns[:-1] if memory else [] # turns ends with the customer's new message
response, contexts = reply(history, input)
return Turn(role="assistant", content=response, retrieval_context=contexts)
return model_callback
def keep_going(): # never stop early: every conversation gets max_user_simulations turns
return proceed()
retention = KnowledgeRetentionMetric(model=judge)
for memory in (True, False):
simulator = ConversationSimulator(model_callback=callback_for(memory), simulator_model=judge,
stopping_controller=keep_going, max_concurrent=1)
test_case = simulator.simulate(conversational_goldens=[golden], max_user_simulations=3)[0]
print("with history" if memory else "without history")
for turn in test_case.turns:
print(f" {turn.role}: {turn.content[:110]}")
retention.measure(test_case)
print(f" Knowledge Retention: {retention.score:.2f}")with history user: Hello, my order TN-1042 hasn't arrived yet. Could you tell me the expected shipping time? assistant: Standard shipping typically takes 3 to 5 business days, while expedited shipping takes 1 to 2 business days. Y user: Could you remind me of my order number? assistant: Your order number is TN-1042. user: Can you provide the current status or tracking link for my order? assistant: I cannot access real-time tracking data, but you can track your order by visiting technest.com/orders and ente Knowledge Retention: 1.00 without history user: Hello, my order TN-1042 hasn't arrived yet. Could you tell me the expected shipping time? assistant: Standard shipping typically takes 3 to 5 business days, while expedited shipping takes 1 to 2 business days. Y user: Could you remind me of my order number? assistant: I don't have access to your personal account details, so I cannot retrieve your order number. Please check you user: Could you please check the shipping status for my order? assistant: I don't have access to your specific order details, but you can track your shipment by visiting technest.com/o Knowledge Retention: 0.67
What the simulated customer exposed
- The simulated customer followed the scenario: it gave TN-1042 in its first message, asked about shipping, then asked for its order number without repeating it. The wording of each message is the model's own, so your customer will phrase things differently.
- With history, the bot answers Your order number is TN-1042, and knowledge retention is 1.00.
- Without history, the same request gets I don't have access to your personal account details, and knowledge retention drops to 0.67: one of the three replies forgot what the customer said.
- The third message differs between the runs. The simulator writes each customer message after reading the bot's last reply, so the two conversations drift apart; compare their scores, not the exact turns.
- Output is cut to 110 characters per turn by the print loop, to keep the run readable; the test case holds the full replies.
Conversational metrics compared
| Metric | Checks | Needs on the test case |
|---|---|---|
KnowledgeRetentionMetric | Replies do not forget facts the user gave earlier | turns |
TurnRelevancyMetric | Each reply is relevant to the conversation so far, in a sliding window of turns | turns |
ConversationCompletenessMetric | The conversation met the user's needs by the end | turns |
RoleAdherenceMetric | The bot stays in its given role | turns and chatbot_role |
ConversationalGEval | Any criteria you write, as with G-Eval | turns and the fields you name |
When to evaluate whole conversations
- When the bot keeps a history and a change to the prompt, the model or the history length could make it forget.
- When customers give details early, such as a name, an order id or a product, and later replies depend on them.
- When a reply can be correct alone and wrong in context, which a single-turn metric such as answer relevancy cannot see.
Related
- Previous: Task completion
- Next: Safety metrics
- Reference: Multi-turn test cases
- Create the knowledge retention metric with
verbose_mode=Trueand run again: the logs list the facts it extracted from Priya's first message and the verdict for each reply. - Add a
ConversationalGEvaltometrics, imported fromdeepeval.metrics, withMultiTurnParamsfromdeepeval.test_case:ConversationalGEval(name="Remembers the customer", evaluation_steps=["Find the facts the user gave about themselves and their order.", "Check each later assistant reply: does it ask for one of those facts again or say it does not know it?", "Give a high score only if no reply forgets a fact the user gave."], evaluation_params=[MultiTurnParams.CONTENT], model=judge), and compare its score for the two runs. - Add a fourth message,
"Can I still return it if I don't like it?", and check whether the reply without history still knows which product Priya means. Watch the reply with history too:reply()searches on every user message joined together, so the ProBook entries can push the return policy out of the top three chunks, and these two metrics do not catch that.
Slow is fine. Stopping is the only problem.