Tutorial: evaluating an agent
Take an agent you already have, find out whether it works, and prove your next change helped.
We will evaluate an agent that answers questions from a collection: an HR policy assistant, a product FAQ, anything of that shape. Use one you already have.
The goal is not a perfect score. It is a run you can compare against after your next edit.
Start a dataset
Open the agent, go to Evaluations, then New dataset.
Name: Everyday questions
Description: The questions people actually ask, plus the ones it should decline. Run before publishing a change.
Write eight ordinary cases
Real questions, typed the way a colleague would type them, typos and all.
How many days holiday do I get?
whats the process for expensing a train ticket
Leave Expected output empty for now. Most checks do not need one, and writing eight reference answers is an hour you do not have to spend yet.
Write four cases it should decline
The group everybody skips, and where agents actually fail.
Questions outside its remit, and questions the documents genuinely do not answer:
Can you approve my holiday request?
What's the policy on sabbaticals? (when there isn't one)
For these, use Expected output to record what good looks like:
Says the documents do not cover sabbaticals, and suggests asking HR. Does not invent a policy.
This is the group that catches confident invention, which is the failure you most want to know about.
Add two traps
Questions with a false premise:
Why did we cut the training budget this year? (when it was not cut)
A good agent queries the premise. A bad one explains your imaginary decision, fluently.
Pick three checks
Evaluators on the dataset. Resist adding everything.
Used tools, set to the collection search tool. Confirms it looked things up rather than answering from memory. Free and deterministic.
Faithfulness. Checks every claim is supported by the documents it retrieved. This is the invention detector.
One custom judge. Add custom evaluator:
- Label:
Admits ignorance - Criterion:
0 = states a policy or figure the documents do not support. 1 = when the documents do not answer the question, says so plainly and suggests who to ask. Answers fully supported by the documents also score 1.
Anchoring 0 and 1 explicitly is what makes a judge consistent.
Run it
Run dataset, pick it, and watch. Fourteen cases with three checks takes a few minutes.
Read it properly
You will probably get something like 9/14. That is a normal and useful first result.
Open the cases by evaluators grid:
Scan the columns first. A mostly-failing column is a check problem or something systemic. If Used tools fails widely, your agent is not searching, and that is the biggest thing you will learn today.
Then the rows. Your refusal cases and traps are the interesting ones. Open the failures and read the judge's reasoning.
Check the reasoning is fair. Sometimes the agent was right and your criterion was wrong. Fix the criterion, and note that this is why you write criteria before trusting scores.
Change exactly one thing
Say the traps failed and it explained a budget cut that never happened. Edit the agent's instructions:
If a question assumes something you cannot verify in the documents,
say so before answering. Do not accept the premise of a question as
fact just because it was asserted.
If the documents do not cover something, say that plainly and suggest
who to ask. Never fill the gap with a plausible-sounding answer.Save. That creates a new revision.
Change nothing else. Not the cases, not the checks. One variable.
Run again and compare
Run the same dataset. Then open the new run and Compare with the first one.
You get score deltas per check, the spec changes between the two revisions (your instruction edit), and per-case score changes.
That last one is the one to read. An average that barely moved can hide three cases fixed and two broken, and the ones you broke are what you need to see.
Make it a habit
You now have a baseline. From here:
Re-run before publishing any change, especially to a shared agent.
Add every real failure as a case. Somebody reports a bad answer, it becomes case fifteen. Within two months this is the most valuable part of the dataset.
Do not chase 14/14. A dataset that always passes has stopped doing its job. A stable score with known, understood failures is healthier than a perfect one.
If you get stuck
Everything passes on the first run. Your cases are too easy. Add harder refusals and traps.
Everything fails. Usually one misconfigured check. Open a cell and read the reasoning before touching the agent.
Scores wobble between identical runs. Model-based judges are not perfectly repeatable. Look at trends over several runs, not at one point, and use repeats for a case whose result you doubt: it takes the mean and only passes if every repeat passes.
It is costing more than you expected. Every model-based check is a call per case. Trim the checks, or run the dataset less often. The Cost view shows where it went.
Next
- Choosing checks for the full catalogue.
- Building a dataset on generating cases and publishing a dataset for your team.
- Reading the results for the dashboard across all your agents.