Write and Test an Evaluation MetricBeta
A metric is a question you ask about every conversation your pipeline has, such as "Was the answer grounded in the documents?" or "Did the agent use the search tool?" Write the metric, test it on real sessions until its answers match your own judgment, and then use it in experiments to measure your pipeline.
About This Task
A metric is only useful if you trust it. A metric that marks good answers as bad, or the other way round, leads you to the wrong conclusions about your pipeline. That's why you test every metric on real sessions before you save it, and read the rationale it gives for each result.
For background on metrics, scores, labels, and LLM judges, see Pipeline Evaluation.
Prerequisites
- Your pipeline has sessions to evaluate. Every query sent to the pipeline creates one. If you have none yet, run a few queries in Playground. Pick queries that represent what your users ask, including ones you expect the pipeline to struggle with.
- To use an LLM judge, your workspace must have a model connected. You can add one in Settings. For details, see Using Hosted Models and External Services.
Decide What to Measure
Before you open the editor, write down the question your metric answers, in one sentence. A good metric question is:
- Specific: "Does the answer cite at least one retrieved document?" is easier to measure than "Is the answer good?"
- About one thing: If you want to measure groundedness and tone, create two criteria or two metrics. When one metric mixes several qualities, a low result doesn't tell you what to fix.
- Answerable from the conversation: The metric reads what the pipeline received, retrieved, called, and replied. It can't see anything outside the session.
Then decide how to measure it:
- Choose LLM-as-a-judge for qualities that need judgment, such as groundedness, relevance, helpfulness, or tone.
- Choose Code for things you can check exactly, such as whether a tool was called, whether a turn failed, or how many tokens a conversation used. Code is faster, cheaper, and returns the same result every time.
Create the Metric
- Open your pipeline and click Evaluation.
- Go to Metrics>New Metric.
- Give the metric a name that says what it measures, for example grounded_answer. You'll see this name as a column in your experiment results.
- Under Method, select LLM-as-a-judge or Code, and then follow the matching section below. ::: Tip LLM Judges vs Code Metrics Use LLM-as-a-judge when you need to judge a quality that's hard to measure exactly, such as groundedness, helpfulness, or tone. Use code for qualities you can measure exactly, such as whether a tool was called or how many tokens a conversation used. :::
Write an LLM Judge
- In What should the model judge?, describe the criterion as instructions to the model.
The model can read the whole session: the turns, the retrieved documents, and the tool calls. Tell it what to look at instead of pasting content in. For example: "Check whether every claim in the final answer is supported by the retrieved documents. Ignore greetings and follow-up questions." - In What should it answer?, choose how the model answers:
- A label, from a list: Add at least two labels, for example grounded, partly grounded, and not grounded. Use this when you can.
A model picks between clearly described options more reliably than it picks a number. - A score, within a range: Set Lowest and Highest under Range for every number. The range defaults to 0 to 1.
The model is told the range, so scores stay comparable between sessions.
- A label, from a list: Add at least two labels, for example grounded, partly grounded, and not grounded. Use this when you can.
- In Answer it, select Once per session to judge the whole conversation, or Per turn to judge each answer separately.
Per turn shows you exactly which answer in a conversation went wrong. - Optional: Click Add criterion to judge more qualities in the same metric.
Criteria judged together tend to influence each other. If you need independent results, create a separate metric for each criterion. - In Which model should judge?, choose the model. Select Connected only to show only models you can use right now.
The model's configuration is saved in the metric, so later changes to the model don't change this metric's results.
The form writes the metric as Python code. Select Code to read it. If you edit that code by hand, the metric stays code from then on.
Write a Code Metric
- In the Evaluation Function editor, write exactly one public function that takes a session. Prefix any helper functions with an underscore so they're not treated as the metric.
A new metric opens with starter code that checks whether an agent called a tool. Change it to fit your question. - Return a result with a rationale:
Score(score=..., rationale=...)for a number.Label(label=..., rationale=...)for a category.NotApplicable(reason=...)when the question doesn't apply to a session, so it doesn't count toward your results.
- To read the session, import helpers from
dc_haystack_utilities.evaluation. Click Helpers to see the full list. For example,traces(session)returns the session's turns,replies(trace)returns the pipeline's answers,component_output(trace, component)returns what a component produced, andfailed(trace)tells you whether a turn failed.
Click the info icon next to the editor to see the full contract, including how to return several metrics from one function with@metrics.
This example scores a session by the share of turns that ran without failing:
from dc_haystack_utilities.evaluation import Score, failed, traces
def turns_without_failure(session):
runs = traces(session)
failures = sum(1 for run in runs if failed(run))
return Score(
score=1 - failures / len(runs),
rationale=f"{failures} of {len(runs)} turns failed",
)
The Returns field above the editor shows what your metric produces, read from your code or from your last try. Check that it matches what you intended.
Test the Metric on Real Sessions
Test before you save. A try runs whatever is in the editor, saved or not. It doesn't create a version or add results to any experiment, so you can experiment freely.
- In Try it, choose a mode:
- One session: Use this while you write, to get quick feedback on a single conversation.
- Several sessions: Use this to check the metric on a range of conversations at once. Each session gets its own column.
- Pick sessions from the list. It shows your pipeline's most recent sessions.
Choose sessions where you already know the right answer: a clearly good one, a clearly bad one, and a borderline case. If the metric gets those right, you can trust it on the rest. - Click Try, or Try on all in several-sessions mode. You can also press Ctrl+Enter (Command+Enter on Mac).
- Read the result and the rationale for each session. Ask yourself whether you'd give the same answer. If not, refine the criterion or the code, and try again.
If your code raises an error, the editor highlights the line that failed. Open the try to see the full traceback. An error in your function says nothing about your pipeline, so fix the function and try again. If a try couldn't run at all, for example because the session is too large, click Retry.
Compare Your Drafts
Every try appears in Try history, grouped into Current draft and Earlier draft. When you try a session again after an edit, the history tells you whether the result changed, so you know which edit made the difference.
- Click Compare to see two tries side by side.
- Click Restore to bring an earlier draft back into the editor. If your current edits aren't saved anywhere else, they're kept in the history first, so you don't lose them.
Tries last only until you leave the page. Save the metric when you're happy with it.
Save the Metric
- Click Save metric. For an existing metric, the button shows the next version, for example Save as v3.
If the code hasn't changed, there's nothing new to save. - If experiments already use this metric, a message tells you how many. Their next run uses the version you're saving. Past runs keep the results of the version they used, so you can still read them.
To look at an earlier version, choose it from Version at the top of the page. Saving changes to an older version creates a new version and doesn't overwrite anything. Click Back to latest to return. To drop unsaved edits, click Discard.
What To Do Next
Add your metric to an experiment and run it on your pipeline's sessions to see how your pipeline performs.
Was this page helpful?