Run an ExperimentBeta
An experiment measures your pipeline with a set of metrics on the sessions you choose. Use it to see how your pipeline performs overall, find the conversations where it struggles, and check whether a change made it better.
About This Task
An experiment is a question you come back to. You create it once, with the metrics that matter for your use case, and run it whenever you want an up-to-date answer: after you change a prompt, switch models, or when new production traffic comes in.
Experiments don't run your pipeline again. They measure sessions your pipeline already had. To measure a change, send queries to the changed pipeline first, and then run the experiment on those sessions.
For background on experiments and metrics, see Pipeline Evaluation.
Prerequisites
- You have at least one metric you've tested and trust. To create one, see Write and Test an Evaluation Metric.
- Your pipeline has sessions to measure. Every query sent to the pipeline, from Playground, a deployed app, or the API, creates one.
Create an Experiment
- Go to your pipeline and click Evaluation.
- Click Experiments, and then click New experiment.
- In Experiment name, type a name that says what the experiment answers, for example Groundedness and cost.
- In Metrics, choose the metrics to run.
Pick a small set that covers what matters most, for example one metric for answer quality and one for cost. The order you pick them in is the order of the columns in the results. - Click Save.
Experiments always use the latest version of each metric. When you improve a metric, the next run picks up the change, and past runs keep their original results.
Run the Experiment
- Open the experiment and click Run experiment.
- In Sessions to judge, choose the sessions to measure.
Choose sessions that represent what you want to learn about. To see how the pipeline does for real users, pick recent production sessions. To check a change, pick the sessions you created after the change. - If the experiment has already measured some of these sessions, choose what to do under Sessions this experiment has already judged:
- Fill in the gaps only: Measure only sessions without a result. Use this to add new sessions to the picture without repeating work.
- Judge everything again: Measure all selected sessions. Use this after you change a metric and want every session measured with the new version.
- Check the number of results the run produces. It's the number of sessions times the number of metric results. LLM judges call a model for each result, so a large run takes longer and costs more.
- Click Run.
A run can measure up to 100 sessions and produce up to 500 results. Your workspace can have up to three runs in progress at once.
Read the Results
When the run finishes, the results show:
- Distribution: One chart for each metric. For a Score, you see a histogram with the median and how many sessions scored below 0.5. For a Label, you see how many sessions got each label.
Start here to get the overall picture. A wide spread or a cluster of low scores tells you where to look. - Sessions: A grid with one row for each session and one column for each metric.
Click a cell to see the value for each turn, with the rationale. This is how you find out why a session scored low, and whether the fix belongs in your pipeline or in your metric.
For metrics that answer per turn, select how to roll up each session: Mean per session for the typical answer, or Worst turn per session to find conversations where a single answer went wrong.
Some cells may show a problem instead of a value:
- Code broke: The metric's code raised an error. This says nothing about your pipeline. Fix the metric and run the experiment again with Fill in the gaps only.
- Run failed: The metric couldn't run on this session, for example because the session's traces are no longer available.
- No metric: The metric ran but returned no value for this session. Check the metric's code to see which sessions it skips.
Sessions where a metric doesn't apply are counted as N/A and left out of the distribution.
The results don't tell you whether your pipeline passes. You decide what's good enough for your use case. Write down your threshold before you look at the results, so the numbers don't change your mind about what you expected.
Compare Runs
Each run is kept in the experiment's history. Choose a run from the run picker to see its results, along with who started it, when, and which metric versions it used.
To compare two runs, make sure they measured similar sessions. If one run covers easy questions and the other covers hard ones, the difference comes from the questions, not your pipeline. To compare two versions of your pipeline fairly:
- Run the same set of queries on each version.
- Run the experiment once on the sessions from each version.
- Compare the distributions and look at sessions that changed.
Evaluate a Single Turn From a Trace
When you spot a problem while inspecting a trace, you can measure that turn on the spot:
- Go to your pipeline and click Analytics > Traces, then open a trace.
- Click Evaluate this turn.
- In Choose an Experiment, choose the experiment to run, and click Run on this turn.
If the pipeline has no experiments yet, click Go to Experiments to create one. - Click Open in Experiment to see the result with its rationale.
The run is kept in the experiment's history and marked as hand-picked, so you can tell it apart from runs over a broader set of sessions.
What To Do Next
- Inspect the lowest-scoring sessions in Trace with Built-In Traces to find what went wrong.
- Improve your pipeline, send new queries, and run the experiment again. For ideas, see Improving Your Question Answering Pipeline.
- If a metric judged a session differently from how you would, refine it in Write and Test an Evaluation Metric.
- Pipeline Evaluation
- Evaluating Step-by-Step
Was this page helpful?