Evaluating Step-by-Step
Why Should I Evaluate?
Before bringing your pipeline to production, you want to make sure it's useful, provides relevant results, and serves the intended purpose. By evaluating your app, you can determine how effective it is and how satisfied the users are. You can share your pipeline prototype with external users and collect their feedback in Haystack Platform interface. Prototypes are customizable, you can add your brand logo and colors to them.
User feedback tells you how people feel about the answers, but only for the answers they rate. To measure every conversation consistently, and to check whether a change made your pipeline better or worse, use metrics and experiments. A metric answers one question about each session, such as "Was the answer grounded in the documents?" An experiment runs your metrics over the sessions you choose and shows you where the pipeline does well and where it struggles. For background, see Pipeline Evaluation.
The two approaches work best together: user feedback shows you what matters to your users, and metrics let you measure it at scale.
Detailed Steps
Here’s a breakdown of the steps we recommend you take to evaluate your pipeline:
- Upload files to your workspace. The pipeline you’re evaluating will run on these files. The more files you upload, the more difficult the task of finding the right answer is.
- Create a pipeline. You can start with one of the ready-made templates.
- Share your pipeline with users. At this stage, you want to get an idea of how people use your pipeline and if it works OK. You’re not aiming for high user satisfaction at this point. You just want to determine if it’s ready for an evaluation through experiments. If most of your users are happy with the results, it means it’s good to go. Otherwise, try tweaking your app until its performance is satisfactory.
- Collect user feedback. It's best to have domain experts and qualified users try your pipeline and then use their feedback to improve it.
- Optimize and improve your pipeline.
Measure Your Pipeline With Metrics and Experiments
Once your pipeline works and people use it, measure it. These steps help you find out how well it performs, where it fails, and whether your changes help.
- Collect sessions to measure. Every query sent to your pipeline creates a session: from Playground, a shared prototype, a deployed app, or the API. If your pipeline isn't in use yet, run a set of queries that represent what your users ask, including difficult ones.
- Decide what good looks like. Write down two or three questions that matter most for your use case, for example "Is the answer grounded in the documents?" and "Did the agent search before answering?" Use what you learned from user feedback to choose them.
- Turn each question into a metric and test it. Use an LLM judge for qualities that need judgment, and code for things you can check exactly. Test each metric on sessions where you already know the right answer until you trust its results. See Write and Test an Evaluation Metric.
- Create an experiment and run it. Group your metrics into an experiment and run it on the sessions you want to measure. See Run an Experiment.
- Find the weak spots. Look at the distribution for each metric, then open the lowest-scoring sessions and read the rationales. They tell you whether the problem is in your pipeline or in your metric.
- Improve and measure again. Change your pipeline, send the same or similar queries to it, and run the experiment on the new sessions. Compare the results with the previous run to see whether the change helped.
Repeat steps five and six as your pipeline evolves. You can also run your experiment on recent production sessions to check that your pipeline still behaves as expected.
Evaluate a Single Turn from a Trace
If a trace shows a turn that looks wrong, you can run one of the pipeline's Experiments on just that turn without leaving the trace. For steps, see Evaluate a Turn from a Trace.
Run the Evaluation
- Open a pipeline-scoped Trace in the Trace inspector.
- Click Evaluate this turn in the Trace header. This opens a new Evaluation tab for that Trace.
- In the Evaluation tab, select one of the pipeline's Experiments.
- Click Run to start the Experiment for this turn.
Before you run it, the tab tells you that the run:
- Is hand-picked, meaning it only evaluates the turn you chose, not a full dataset.
- Is kept in the Experiment's run history so you and others can find it again.
- Does not feed into the Experiment's trends, since it only covers one turn.
The run scope is the turn's session: if the turn is part of a chat, the run covers the whole session the turn belongs to; outside of a chat, it covers just that turn's own query.
Read the Results
Results appear grouped for readability:
- Rows are grouped first by Evaluator, in the order the Experiment defines, and then by Metric key.
- Rows scored at a finer grain than the turn keep their original addresses and appear in the order the pipeline run produced them.
- Metrics scored at the session level are shown separately, under a "Session-level, for context" group, since they describe the whole session rather than just this turn.
- Rationales and any "not applicable" reasons a Metric reports appear inline with its row.
- Rows that errored or didn't finish show the same details panel used elsewhere in Evaluation.
Share a Diagnosis
Each result includes an Open in Experiment link. Clicking it takes you to the Experiment detail page opened on that specific run, so you can share the link with a teammate and they'll land on the same run you're looking at, not just the Experiment's latest run.
Results shown in the Evaluation tab only last as long as the page stays open. To come back to them later, use the Open in Experiment link rather than reopening the Trace.
If the pipeline has no Experiments yet, the Evaluation tab links you to the Experiments section so you can create one first.
Related Information
Was this page helpful?