DAG metric
DAGMetric is a DeepEval metric that scores a test case by walking a decision tree of small judge questions, where every path ends in a score you chose in advance.
Last updated: 05 Oct, 2026 · DeepEval 4.2.8
G-Eval lets the judge pick the number. When your rule is a list of known checks, such as a return reply must give the 30-day window, then who pays shipping and how long the refund takes, you can fix the score for each outcome yourself and let the judge answer only yes-or-no and pick-one questions. DAG usually means directed acyclic graph; DeepEval names its class DeepAcyclicGraph. Either way: nodes connected one way, with no loops.
The DAGMetric API
from deepeval.metrics import DAGMetric
from deepeval.metrics.dag import DeepAcyclicGraph
dag = DeepAcyclicGraph(root_nodes=[first_node]) # checks the whole graph when it is built
metric = DAGMetric(name="Return reply", dag=dag, model=judge)
metric.measure(test_case) # score = the verdict's score / 10A graph has three kinds of node. A TaskNode asks the judge to pull text out of the test case. A BinaryJudgementNode asks a yes-or-no question. A NonBinaryJudgementNode asks the judge to pick one of the answers you list. Each judgement ends in a verdict, which either gives a score from 0 to 10 or leads to the next node.
Extracting the return facts with a TaskNode
extract = TaskNode(
instructions="List every fact the reply states about returns: the time limit, who pays return shipping, and how long the refund takes.",
output_label="Return facts in the reply", # the name the next nodes see it under
evaluation_params=[SingleTurnParams.ACTUAL_OUTPUT],
)The task reads only the reply. Its output goes to every node connected below it, under its output_label.
Checking the window with a BinaryJudgementNode
window = BinaryJudgementNode(criteria="Do the facts give a time limit of 30 days for returns?")
window.add_verdict(verdict=False, score=0) # no window: the path ends at 0
window.add_verdict(verdict=True, then=details) # window present: go on to the detailsThis is the gate. A reply without the 30-day window scores 0 whatever else it says.
Grading the details with a NonBinaryJudgementNode
details = NonBinaryJudgementNode(criteria="Which of these two facts are present: who pays return shipping, and the 5 to 7 business day refund time?")
details.add_verdict(verdict="Both", score=10)
details.add_verdict(verdict="Only one", score=6)
details.add_verdict(verdict="Neither", score=3)The judge's answer is limited to the three verdict strings. Each one maps to a fixed score, so the same answer always gets the same number.
Connecting the task to both judgements
extract.add_node(window) # the window check reads the extracted facts
extract.add_node(details) # so does the details checkadd_verdict creates a VerdictNode for you. DeepEval 4.2.8 still accepts the older way of passing children=[VerdictNode(...)] to a node, with a DeprecationWarning that it will be removed.
A binary node with one verdict
A yes-or-no node needs both answers covered. Build the graph with only the False verdict:
from deepeval.metrics.dag import BinaryJudgementNode, DeepAcyclicGraph
window = BinaryJudgementNode(criteria="Do the facts give a time limit of 30 days for returns?")
window.add_verdict(verdict=False, score=0)
dag = DeepAcyclicGraph(root_nodes=[window])Traceback (most recent call last):
File "main.py", line 5, in <module>
dag = DeepAcyclicGraph(root_nodes=[window])
ValueError: BinaryJudgementNode must have exactly 2 children.DeepAcyclicGraph checks every node when it is built, before any judge call. A binary node must have exactly one True and one False verdict, and each verdict must have a score or a then, not both. Adding the True verdict fixes it.
- written in Custom judge model
View the code here
import os
from deepeval.models import DeepEvalBaseLLM
from openai import AsyncOpenAI, OpenAI
GROQ_URL = "https://api.groq.com/openai/v1"
class GroqJudge(DeepEvalBaseLLM):
"""A DeepEval judge model that runs on Groq."""
def __init__(self, model="openai/gpt-oss-120b"):
self.model_name = model
key = os.environ["GROQ_API_KEY"]
# on a 429 (rate limit) the client waits and tries again, up to 8 times
self.client = OpenAI(api_key=key, base_url=GROQ_URL, max_retries=8)
self.async_client = AsyncOpenAI(api_key=key, base_url=GROQ_URL, max_retries=8)
def load_model(self):
return self.client
def get_model_name(self):
return self.model_name
def request(self, prompt, schema):
request = {"model": self.model_name, "messages": [{"role": "user", "content": prompt}], "temperature": 0}
if schema is not None:
# ask Groq for JSON in the shape of the metric's Pydantic schema
json_schema = {"name": schema.__name__, "schema": schema.model_json_schema()}
request["response_format"] = {"type": "json_schema", "json_schema": json_schema}
return request
def generate(self, prompt, schema=None):
reply = self.client.chat.completions.create(**self.request(prompt, schema))
text = reply.choices[0].message.content
return schema.model_validate_json(text) if schema else text
async def a_generate(self, prompt, schema=None):
reply = await self.async_client.chat.completions.create(**self.request(prompt, schema))
text = reply.choices[0].message.content
return schema.model_validate_json(text) if schema else text
judge = GroqJudge(os.environ.get("JUDGE_MODEL", "openai/gpt-oss-120b"))
The return_dag.py file
All the pieces in one file, return_dag.py, next to judge.py:
from deepeval.metrics import DAGMetric
from deepeval.metrics.dag import BinaryJudgementNode, DeepAcyclicGraph, NonBinaryJudgementNode, TaskNode
from deepeval.test_case import SingleTurnParams
from judge import judge
extract = TaskNode(
instructions="List every fact the reply states about returns: the time limit, who pays return shipping, and how long the refund takes.",
output_label="Return facts in the reply",
evaluation_params=[SingleTurnParams.ACTUAL_OUTPUT],
)
window = BinaryJudgementNode(criteria="Do the facts give a time limit of 30 days for returns?")
details = NonBinaryJudgementNode(criteria="Which of these two facts are present: who pays return shipping, and the 5 to 7 business day refund time?")
extract.add_node(window)
extract.add_node(details)
window.add_verdict(verdict=False, score=0)
window.add_verdict(verdict=True, then=details)
details.add_verdict(verdict="Both", score=10)
details.add_verdict(verdict="Only one", score=6)
details.add_verdict(verdict="Neither", score=3)
metric = DAGMetric(name="Return reply", dag=DeepAcyclicGraph(root_nodes=[extract]), model=judge)Scoring three support replies with the DAG
Three replies to a return question: one with all three facts, one with only the window, and one with no window at all.
from deepeval.test_case import LLMTestCase
from return_dag import metric
replies = [
"You have 30 days to return it. Return shipping is on you unless it arrived defective, and refunds take 5 to 7 business days.",
"You have 30 days to send it back in the original packaging.",
"Sure, you can return it any time. Just contact us.",
]
for reply in replies:
metric.measure(LLMTestCase(input="Can I return my SoundPods Pro?", actual_output=reply))
print(metric.score, "|", reply)
print(" ", metric.reason)1.0 | You have 30 days to return it. Return shipping is on you unless it arrived defective, and refunds take 5 to 7 business days.
The score is 1.0 because the DAG traversal shows the TaskNode listed the required return facts, the BinaryJudgementNode verified the 30‑day time limit, the NonBinaryJudgementNode confirmed both the shipping‑payer and refund‑time facts, and the final VerdictNode returned "Both", meeting the metric criteria.
0.3 | You have 30 days to send it back in the original packaging.
The score is 0.3 because the DAG traversal shows the BinaryJudgementNode confirmed the 30‑day time limit, but the NonBinaryJudgementNode found neither a shipping‑responsibility fact nor a refund‑time fact, leading the final VerdictNode to output "Neither", indicating incomplete return information and thus a low score.
0.0 | Sure, you can return it any time. Just contact us.
The score is 0.0 because the DAG traversal shows the TaskNode generated facts with no 30‑day limit, the BinaryJudgementNode evaluated the 30‑day criterion as False, and the final VerdictNode deterministically returned False, leading to a zero score for the Return reply metric.Which path each reply took
- The full reply scores 1.0. The window check answered True, the details check picked Both, and that verdict's score of 10 becomes 1.0.
- The window-only reply scores 0.3. It passed the gate, then the details check picked Neither, whose score is 3. Its mention of the original packaging earns nothing, because no node asks about it.
- The "any time" reply scores 0.0. The window check answered False and that verdict ends the path at 0; the details check never ran.
- Each reason retells the path: the node verdicts and the final verdict. DAGMetric writes it with one more judge call, which
include_reason=Falseskips.
DAGMetric vs G-Eval
| DAGMetric | GEval | |
|---|---|---|
| Who sets the score | You, on each verdict | The judge |
| Judge questions | Several small ones, one per node on the path, plus one for the reason | One overall judgement (plus the steps, if you gave only criteria) |
| Setup | Nodes and verdicts to wire | Criteria or steps |
| Good for | Rules with gates and known outcomes | Overall quality that is hard to list |
When to use a DAG metric
- When one missing fact should fail the answer outright, like the 30-day return window, and the rest earns partial credit.
- When two people should agree on the score for a given outcome, because the number comes from your verdicts, not the judge.
- When you want to see which check failed:
verbose_mode=Trueprints each node's verdict and reason.
verbose_mode when a score looks off and read the root's verdict first.Related
- Previous: Threshold and strict mode
- Next: Custom metrics
- Reference: DAG (Deep Acyclic Graph)
- Add
window.add_verdict(verdict=True, score=10)before the last line of the broken example and run it: the graph builds without an error. - Pass
verbose_mode=TruetoDAGMetricinreturn_dag.pyand run only the second reply to read each node's verdict. - Change the score of
"Neither"to 0 inreturn_dag.pyand run again: the window-only reply now scores 0.0, the same as a reply with no window.
Little by little, you're building something great.