DeepEvaldeepeval 4.2.8 · Python 3.10+
Dashboard
0%
1
Curious builder0 XP earned · 300 to level 2
0 daysFinish a lesson to begin
Badge collection0 of 6 unlocked
34 small wins to finish your pathNext lesson →

DAG metric

DAGMetric is a DeepEval metric that scores a test case by walking a decision tree of small judge questions, where every path ends in a score you chose in advance.

Last updated: 05 Oct, 2026 · DeepEval 4.2.8

G-Eval lets the judge pick the number. When your rule is a list of known checks, such as a return reply must give the 30-day window, then who pays shipping and how long the refund takes, you can fix the score for each outcome yourself and let the judge answer only yes-or-no and pick-one questions. DAG usually means directed acyclic graph; DeepEval names its class DeepAcyclicGraph. Either way: nodes connected one way, with no loops.

The DAGMetric API

python
from deepeval.metrics import DAGMetric
from deepeval.metrics.dag import DeepAcyclicGraph

dag = DeepAcyclicGraph(root_nodes=[first_node])  # checks the whole graph when it is built
metric = DAGMetric(name="Return reply", dag=dag, model=judge)
metric.measure(test_case)                        # score = the verdict's score / 10

A graph has three kinds of node. A TaskNode asks the judge to pull text out of the test case. A BinaryJudgementNode asks a yes-or-no question. A NonBinaryJudgementNode asks the judge to pick one of the answers you list. Each judgement ends in a verdict, which either gives a score from 0 to 10 or leads to the next node.

Extracting the return facts with a TaskNode

python
extract = TaskNode(
    instructions="List every fact the reply states about returns: the time limit, who pays return shipping, and how long the refund takes.",
    output_label="Return facts in the reply",       # the name the next nodes see it under
    evaluation_params=[SingleTurnParams.ACTUAL_OUTPUT],
)

The task reads only the reply. Its output goes to every node connected below it, under its output_label.

Checking the window with a BinaryJudgementNode

python
window = BinaryJudgementNode(criteria="Do the facts give a time limit of 30 days for returns?")
window.add_verdict(verdict=False, score=0)      # no window: the path ends at 0
window.add_verdict(verdict=True, then=details)  # window present: go on to the details

This is the gate. A reply without the 30-day window scores 0 whatever else it says.

Grading the details with a NonBinaryJudgementNode

python
details = NonBinaryJudgementNode(criteria="Which of these two facts are present: who pays return shipping, and the 5 to 7 business day refund time?")
details.add_verdict(verdict="Both", score=10)
details.add_verdict(verdict="Only one", score=6)
details.add_verdict(verdict="Neither", score=3)

The judge's answer is limited to the three verdict strings. Each one maps to a fixed score, so the same answer always gets the same number.

Connecting the task to both judgements

python
extract.add_node(window)   # the window check reads the extracted facts
extract.add_node(details)  # so does the details check

add_verdict creates a VerdictNode for you. DeepEval 4.2.8 still accepts the older way of passing children=[VerdictNode(...)] to a node, with a DeprecationWarning that it will be removed.

A binary node with one verdict

A yes-or-no node needs both answers covered. Build the graph with only the False verdict:

Example
from deepeval.metrics.dag import BinaryJudgementNode, DeepAcyclicGraph

window = BinaryJudgementNode(criteria="Do the facts give a time limit of 30 days for returns?")
window.add_verdict(verdict=False, score=0)
dag = DeepAcyclicGraph(root_nodes=[window])

DeepAcyclicGraph checks every node when it is built, before any judge call. A binary node must have exactly one True and one False verdict, and each verdict must have a score or a then, not both. Adding the True verdict fixes it.

Project files used on this pageThis lesson builds on a project from earlier lessons. The code below imports this file. Click a file to see its code, or follow the link to the lesson that wrote it. To run the code yourself, keep it in the same folder.
View the code here
judge.py
import os

from deepeval.models import DeepEvalBaseLLM
from openai import AsyncOpenAI, OpenAI

GROQ_URL = "https://api.groq.com/openai/v1"


class GroqJudge(DeepEvalBaseLLM):
    """A DeepEval judge model that runs on Groq."""

    def __init__(self, model="openai/gpt-oss-120b"):
        self.model_name = model
        key = os.environ["GROQ_API_KEY"]
        # on a 429 (rate limit) the client waits and tries again, up to 8 times
        self.client = OpenAI(api_key=key, base_url=GROQ_URL, max_retries=8)
        self.async_client = AsyncOpenAI(api_key=key, base_url=GROQ_URL, max_retries=8)

    def load_model(self):
        return self.client

    def get_model_name(self):
        return self.model_name

    def request(self, prompt, schema):
        request = {"model": self.model_name, "messages": [{"role": "user", "content": prompt}], "temperature": 0}
        if schema is not None:
            # ask Groq for JSON in the shape of the metric's Pydantic schema
            json_schema = {"name": schema.__name__, "schema": schema.model_json_schema()}
            request["response_format"] = {"type": "json_schema", "json_schema": json_schema}
        return request

    def generate(self, prompt, schema=None):
        reply = self.client.chat.completions.create(**self.request(prompt, schema))
        text = reply.choices[0].message.content
        return schema.model_validate_json(text) if schema else text

    async def a_generate(self, prompt, schema=None):
        reply = await self.async_client.chat.completions.create(**self.request(prompt, schema))
        text = reply.choices[0].message.content
        return schema.model_validate_json(text) if schema else text


judge = GroqJudge(os.environ.get("JUDGE_MODEL", "openai/gpt-oss-120b"))

The return_dag.py file

All the pieces in one file, return_dag.py, next to judge.py:

python
from deepeval.metrics import DAGMetric
from deepeval.metrics.dag import BinaryJudgementNode, DeepAcyclicGraph, NonBinaryJudgementNode, TaskNode
from deepeval.test_case import SingleTurnParams

from judge import judge

extract = TaskNode(
    instructions="List every fact the reply states about returns: the time limit, who pays return shipping, and how long the refund takes.",
    output_label="Return facts in the reply",
    evaluation_params=[SingleTurnParams.ACTUAL_OUTPUT],
)
window = BinaryJudgementNode(criteria="Do the facts give a time limit of 30 days for returns?")
details = NonBinaryJudgementNode(criteria="Which of these two facts are present: who pays return shipping, and the 5 to 7 business day refund time?")

extract.add_node(window)
extract.add_node(details)
window.add_verdict(verdict=False, score=0)
window.add_verdict(verdict=True, then=details)
details.add_verdict(verdict="Both", score=10)
details.add_verdict(verdict="Only one", score=6)
details.add_verdict(verdict="Neither", score=3)

metric = DAGMetric(name="Return reply", dag=DeepAcyclicGraph(root_nodes=[extract]), model=judge)

Scoring three support replies with the DAG

Three replies to a return question: one with all three facts, one with only the window, and one with no window at all.

ExampleAPI key
from deepeval.test_case import LLMTestCase

from return_dag import metric

replies = [
    "You have 30 days to return it. Return shipping is on you unless it arrived defective, and refunds take 5 to 7 business days.",
    "You have 30 days to send it back in the original packaging.",
    "Sure, you can return it any time. Just contact us.",
]
for reply in replies:
    metric.measure(LLMTestCase(input="Can I return my SoundPods Pro?", actual_output=reply))
    print(metric.score, "|", reply)
    print("   ", metric.reason)

Which path each reply took

  • The full reply scores 1.0. The window check answered True, the details check picked Both, and that verdict's score of 10 becomes 1.0.
  • The window-only reply scores 0.3. It passed the gate, then the details check picked Neither, whose score is 3. Its mention of the original packaging earns nothing, because no node asks about it.
  • The "any time" reply scores 0.0. The window check answered False and that verdict ends the path at 0; the details check never ran.
  • Each reason retells the path: the node verdicts and the final verdict. DAGMetric writes it with one more judge call, which include_reason=False skips.

DAGMetric vs G-Eval

DAGMetricGEval
Who sets the scoreYou, on each verdictThe judge
Judge questionsSeveral small ones, one per node on the path, plus one for the reasonOne overall judgement (plus the steps, if you gave only criteria)
SetupNodes and verdicts to wireCriteria or steps
Good forRules with gates and known outcomesOverall quality that is hard to list

When to use a DAG metric

  • When one missing fact should fail the answer outright, like the 30-day return window, and the rest earns partial credit.
  • When two people should agree on the score for a given outcome, because the number comes from your verdicts, not the judge.
  • When you want to see which check failed: verbose_mode=True prints each node's verdict and reason.
Watch out. The judge still answers every node, so a DAG is not fully deterministic. A wrong yes or no at the root sends the whole reply down the wrong path. Turn on verbose_mode when a score looks off and read the root's verdict first.
Try it yourself
  • Add window.add_verdict(verdict=True, score=10) before the last line of the broken example and run it: the graph builds without an error.
  • Pass verbose_mode=True to DAGMetric in return_dag.py and run only the second reply to read each node's verdict.
  • Change the score of "Neither" to 0 in return_dag.py and run again: the window-only reply now scores 0.0, the same as a reply with no window.

Little by little, you're building something great.