Reading the results
The scorecard, the case-by-case grid, and the comparison that tells you whether a change helped.
Run dataset, pick one, and it scores your agent's current revision against every case. You can watch the progress and cancel: cases not yet started are skipped, and cases in flight finish first.
The scorecard
The headline is cases passing, for example "14/20 cases passing", and a case passes when all of its gating checks pass. Alongside it you may see:
- a count of incomplete cases, which did not finish
- a count of cases that had no applicable checks, worth investigating: those scored nothing at all
- info, marking checks that are tracked but never fail a case
Do not chase 20/20. A dataset that always scores full marks has stopped telling you anything, and usually means the awkward cases are missing. Somewhere in the high teens with a stable set of known failures is a healthier place to be.
Cases by evaluators
A grid: one row per case, one column per check, so you can see at a glance whether a failure is one bad case or one bad check.
Each cell is Pass, Fail, Skipped, Errored, or No pass bar (scored, but with no threshold set). Open a cell for the score, the threshold it was judged against, and the judge's reasoning.
Read the reasoning before you act. Often the agent was right and the check was wrong, which is a signal to fix your criterion rather than your agent.
The two patterns to look for:
A column mostly failing is a problem with the check, or with something systemic. Twenty cases do not all fail faithfulness by coincidence.
A row mostly failing is one hard case. Read it: sometimes it is genuinely unfair, more often it is the interesting one.
The rest of a run
Summary is a written commentary on how the run went, generated when it completes. You can regenerate it.
Steps and Tool calls show what the agent actually did per case, which is where you find out it never searched the collection.
Artifacts holds anything it produced.
Trace ID is for your administrators. If your organisation collects traces, that id finds the full detail.
Comparing two runs
The reason to do any of this.
Compare with, then pick another completed run of the same dataset. You get:
Score deltas per evaluator, run A against run B, with the difference.
Spec changes between the two, since a run records which revision of the agent it used. So you see the score moved and what changed: instructions edited, tools added or removed, temperature changed, collections pinned or unpinned.
Per-case score changes, which is where the useful detail is. An unchanged average can hide four cases improving and four getting worse, and that is usually more interesting than the average.
Compare runs of the same dataset. Comparing across datasets tells you nothing, which is why the dialog only offers same-dataset runs.
Revisions
Every save of an agent creates a revision. The Revisions view lists them with a summary (instructions, how many tools, sub-agents, pinned collections) and lets you select two and Compare to see exactly what changed.
Pair that with run comparison and you can answer the question this feature exists for: did the thing I changed on Tuesday make it better?
The dashboard
Evaluations in the left rail shows eval health across every agent you can access, with only show what needs attention.
Trends draw a pass-rate line over recent runs.
A dashed break in a trend line marks an evaluator version change. Scores either side are not comparable, so the line is cut rather than joined.
That is a deliberate refusal to draw a reassuring line through a discontinuity. If you see one, do not read across it.
You may also see runs marked as predating a scoring rework, whose numbers are not comparable with today's. Older runs are excluded from trends rather than silently mixed in.
What to do with a failure
Read the reasoning first. Was the agent wrong, or the check?
Check whether it used its tools. A confident answer with no collection search is the classic failure, and it is usually fixed in the instructions rather than the model.
Change one thing, then re-run. Changing instructions and cases together means you learn nothing from the result.
Keep the failing case. Once it passes, it is a case that can never quietly regress.