Choosing checks
The kinds of evaluator, what each is good for, and how a case ends up passing or failing.
An evaluator is one check applied to every case. You pick them on the dataset, and can override them per case.
They come in five groups, and the useful distinction is cheap and certain versus expensive and judgemental.
Behavioural: cheap, deterministic
Mechanical checks with no model involved. They cost nothing and always give the same answer.
- Output type matches: did it come back in the shape expected?
- Latency under threshold: was it fast enough?
- Used tools: did it call the tools it should have?
Used tools is more useful than it looks. Set it to check the agent searched your collection rather than answering from memory, and you catch the single commonest quiet failure. It also inverts: pass when the tools are not used, which is how you check a guardrail ("this agent must never send email").
Start here. These checks are free and catch real problems.
Structured: is it faithful to the documents?
For agents that answer from a collection. They pull the answer apart and check it against the retrieved material.
- Faithfulness: is every claim supported by the documents, or did it embellish?
- Answer relevancy: does it actually answer the question asked?
- Context precision and context recall: did retrieval find the right material, and enough of it?
Faithfulness is the one to run on any agent whose job is answering from your documents. It is the check for confident invention, which is the failure people most fear and least often test.
The two context checks are diagnostic: when faithfulness is fine but answers are thin, they tell you whether the problem is the agent or the collection underneath it.
Heuristic: a judge with a rubric
A model scores each answer against a general rubric.
- Completeness: did it cover what it should?
- Safety: is there anything problematic?
Fast and qualitative. Treat the scores as a signal rather than a measurement: they are opinions, and they wobble a little between runs.
Data task: did the artifact come out right?
For agents producing tables and charts, these compare what came out against an expected shape: table shape, columns, schema, values, chart type, plus an overall data analysis quality judge.
Only relevant if your agent produces data artifacts. Ignore otherwise.
Custom: your own criterion
Add custom evaluator lets you write the check yourself. A label (which becomes the column heading in results) and a criterion in plain language.
The criterion is where the work is. Vague criteria produce meaningless scores:
Weak: The response is good quality.
Strong: Uses our house terminology: "member" not "user", "plan" not "package". Never promises a specific timeline. Ends with a next step the reader can take.
The dialog suggests anchoring the scale, and it is worth doing. Taking the same criterion and saying what the ends look like:
Anchored: 0 = uses "user" or "package", or promises a timeline. 1 = uses "member" and "plan" throughout, makes no timeline promise, and ends with a next step.
A judge told what 0 and 1 look like is far more consistent than one left to interpret "good".
Add as many as you need, each with a distinct label.
How a case passes or fails
Worth understanding, because the headline number depends on it.
Each check produces a score from 0 to 1, and some have a threshold: score at or above it, the check passes. A check with no threshold is scored and shown but has no pass bar, and results show it as such rather than pretending.
A case passes when all of its gating checks pass. Skipped and errored checks do not count against it.
Gating is the important word. A check can be marked informational: it is scored, displayed and tracked over time, but it never fails a case. Use that for things you want to watch without blocking on, like a stylistic judge whose opinion you do not want deciding whether a run is green.
Repeats, if you run a case more than once, take the mean score, and the case passes only if every repeat passes. Deliberate: an answer that passes two times in three is not a passing answer, it is an intermittent one, and intermittent is signal rather than noise.
Picking a sensible set
For most agents, start with:
- Used tools, to confirm it searches rather than guesses
- Faithfulness, if it answers from documents
- One custom judge for whatever "right" means in your job
Three checks you understand beat nine you do not. You can tell what a failure means, which is the whole point.
Add more when a real failure slips past the ones you have, and use informational rather than gating for anything you are not yet sure you trust.
Every model-based check (structured, heuristic, custom, and the data analysis judge) costs a model call per case, on top of running the agent itself. Twenty cases with four judges is eighty judge calls per run.
The Cost view breaks eval spend down by agent, by kind (the agent under test, the scoring, the summaries) and by day. Worth a look before putting a large dataset on a frequent schedule.