Pipeline EvaluationBeta
Evaluation turns "the answers seem fine" into numbers you can track. You decide what good looks like for your use case and write it down as a metric. Then you measure your pipeline against it on conversations it already had.
What You Evaluate
You evaluate the conversations and answers your pipeline already produced. This lets you see how your pipeline behaves on real conversations, without having to prepare a test dataset first.
A turn is one pipeline run: one query and everything the pipeline did to answer it, including retrieved documents, tool calls, and the final reply. The platform records each turn as a trace.
A session is a conversation: all the turns that share one chat. A single query sent outside a chat counts as a session with one turn.
Every query you send to your pipeline, from Playground, a deployed app, or the API, becomes a session you can evaluate. Your evaluation reflects how people use the pipeline, not how you expect them to use it.
You evaluate whole sessions, so a metric can judge things a single answer can't show, such as whether the pipeline remembered earlier questions or whether an agent called the right tool.
Main Benefits
Evaluation lets you:
- Check whether a pipeline is ready for production by measuring it on a realistic set of conversations before you share it widely.
- See whether a change helped by measuring before and after you change a prompt, a model, or a retrieval setting.
- Find where the pipeline fails by inspecting the sessions with the lowest scores instead of reading every conversation.
- Catch problems that nobody reports by measuring production conversations.
User feedback answers some of these questions too, but it depends on people taking the time to rate answers. Metrics run on every session you choose, so you get consistent coverage.
LLM-as-a-Judge and Code
You can evaluate in two ways: with an LLM-as-a-judge or with code.
LLM-as-a-judge asks a language model to rate the session against criteria you describe in plain language. Use it for qualities that need judgment: groundedness, helpfulness, tone, or whether the answer addresses the question. You write the criteria, choose the answers the model can give, and pick the model, without writing code.
A code evaluation relies on a Python function that reads the session and computes the value. Use it for things you can check exactly: whether a tool was called, whether a turn failed, how many tokens the conversation used, or whether the answer contains a required phrase. Code is fast, cheap, and gives the same result every time.
Many teams combine both: code for the facts and an LLM judge for the qualities.
Both methods produce one metric that the platform stores as code. The LLM-as-a-judge form writes that code for you. You can switch to the code view to see it, and edit it by hand if you need more control.
An LLM judge is a model and can be wrong, so its rationales matter. They show you whether the judge reasoned correctly.
Metrics
A metric is a question you ask about every session, answered the same way each time. For example:
- Did the answer stay grounded in the retrieved documents?
- Did the agent call the search tool before answering?
- How polite was the reply, on a scale from 0 to 1?
- How many tokens did the conversation use?
A metric returns a value and a rationale: the measurement and the reason for it. The rationale lets you check whether the metric judged correctly, which is how you learn to trust it.
Scores and Labels
Each metric returns one of two kinds of value. A score is a number, such as 0.8. Use it for things that come in degrees, such as relevance or cost. Scores let you see a distribution and spot sessions at the low end.
A label is one value from a list you define, such as called or skipped, or grounded, partly grounded, and not grounded. Use it for things that fall into categories. A yes-or-no question is a label with two values.
When you ask a model to judge, prefer labels. A model picks between clearly described categories more reliably than it places a value on a numeric scale.
A metric can also report that it doesn't apply to a session, for example a tool-call metric on a session where no tool was needed. The platform counts those sessions separately, so they don't distort your results.
Session, Turn, or Component
A metric can answer once for the whole session, once for each turn, or for a component within a turn. A component is one building block of your pipeline, such as the retriever, which finds documents. Choose the level that matches your question:
- Per session: "Did the conversation reach a resolution?"
- Per turn: "Was this answer grounded?"
- Per component: "Did the retriever return relevant documents?"
When a metric answers per turn, you can roll the results up per session using the average or the lowest-scoring turn. The lowest-scoring turn is useful when a single bad answer ruins the conversation.
Metric Versions
Each time you save a metric with new code, you get a new version. Experiments always use the latest version and record which version produced each result. You can improve a metric over time and still read past results exactly as they were measured.
Metrics belong to your workspace, so you can reuse one metric across several pipelines.
Experiments
An experiment is a named set of metrics you run over your pipeline's sessions. For example, an experiment called "Groundedness and cost" could combine an LLM judge for groundedness with a code metric for token usage.
Each time you run an experiment, you pick the sessions to measure. The run gives you a distribution for each metric, which shows how scores spread across sessions, so you see the overall picture at a glance. It also gives you a grid of sessions and metrics, so you can find the sessions behind a low score and inspect each turn and its rationale.
An experiment is how you answer the same question repeatedly: run it now, change your pipeline, send new queries, and run it again.
Comparing Runs
Two runs are comparable only if they measure similar sessions. If one run covers simple questions and the other covers hard ones, the difference in scores reflects the questions, not your pipeline. To compare two pipeline versions fairly, both need to answer the same or similar queries.
A Measurement, Not a Verdict
Results tell you what happened, not whether it's good enough. There is no built-in pass or fail. You decide what threshold matters for your use case: 0.7 groundedness may be fine for an internal assistant and not good enough for a customer-facing one.
Related Information
Was this page helpful?