Write and Test a Metric
Write a metric and try it against real sessions, side by side, before using it in an experiment.
About This Task
A metric judges a session and returns a measurement: a flag, a score, or feedback. You write a metric either by filling in a form that asks a model to judge a session, or by writing a Python function yourself. Writing and trying a metric both happen on the same page, so you can check what a metric measures before you add it to an experiment.
Every metric lives in your workspace and keeps a history of versions. Saving adds a version only when the code actually changed — saving the same code again doesn't add a new version, and going back to older code returns to the version that already holds it.
Trying a metric is separate from saving it. A try runs whatever is currently in the editor, saved or not, against one or more sessions you pick. Nothing about a try is recorded anywhere: no version is created, and no result reaches a chart. Tries exist only for as long as you stay on this page — leaving it clears them.
Open the Metric Editor
You can open the metric editor from two places:
- On the pipeline's Evaluation > Metrics page, click New metric, or click an existing metric to edit it.
- When creating or editing an experiment, click New metric next to the metrics list. This takes you to the metric editor page and closes the experiment dialog — once you save or leave the metric editor, you return to the metrics list rather than to the experiment you were creating, so add your new metric to the experiment afterward.
Both entry points open the same full-page editor, with the metric definition on the left and the test bench on the right.
Write the Metric
-
In the top bar, type a name for your metric. The name is what every chart and report labels this measurement with — renaming it later relabels past results too, rather than starting a new measurement.
-
Under Method, choose how to define the metric:
- Form: Describe what the model should judge, choose whether it answers with one of two labels or a number, and pick the model that judges sessions. The form writes the underlying Python code for you.
- Code: Write a Python function yourself. The function takes a session and returns a flag, a score, or feedback. Use the Helpers button in the editor to see the functions you can call to read a session's turns, tool calls, and token usage. Click the info icon next to the editor to see the function's contract.
Editing the code by hand after using the form makes the metric code from then on — switching back to the form isn't possible once you've made manual edits.
-
The editor shows what your function currently returns (Flag, Score, or Feedback) next to the code, read from the code itself or from your last try.
Try the Metric
Use the test bench on the right to run your metric before saving it:
- Choose One session to try the metric against a single session, or Several sessions to try it against a batch at once.
- Pick the session or sessions to try it on.
- Click Try, or press Ctrl+Enter (Cmd+Enter on Mac) from inside the editor.
Reading Tries in One-Session Mode
Every try you run appears in Try history, grouped under Current draft or Earlier draft depending on whether the code has changed since that try ran:
- Select a try to see its result, rationale, and the session it ran on.
- If the same session was tried on an earlier draft of the code and the answer changed, the selected try shows that contrast, with options to Compare the two tries side by side or Restore the earlier draft.
- If a try's evaluation function raised an error, the line that broke is marked in the editor and in the try detail, along with the error message and a traceback you can copy.
- A try that fails to run (for example, because of a temporary platform issue) shows Retry in place, rather than a toast notification that disappears before you see it.
Reading Tries in Several-Sessions Mode
Trying several sessions at once builds a matrix with sessions as rows and each batch of tries as a column. Cells show what each session returned, flag when an answer changed from a previous column, and let you restore an earlier draft from any cell.
Save a Version
Click Save metric (for a new metric) or Save as vN (for an existing one) once you're done writing and trying your metric. If the draft matches a version that already exists, Save stays disabled until you change the code or name.
If the metric is already used by one or more experiments, saving it changes the version those experiments use the next time they run. The editor warns you about this before you save.
View an Older Version
Use the version list in the top bar to load an earlier version as your current draft. The editor shows a banner confirming you're viewing an older version. Editing or saving from there creates a new version on top of it. Click Back to latest to return to the most recent version without losing your changes elsewhere.
Leave the Page
Because tries aren't saved anywhere, the editor warns you before you navigate away if you have unsaved changes. Save your metric first if you want to keep your edits; the try history itself can't be kept across page visits.
Related Information
Was this page helpful?