Evaluations
How you find out whether your agent is any good, and whether your last change helped.
You built an agent. You tried it a few times, it seemed fine, you shared it.
Then you improved the instructions. Is it better now? You try it twice, it seems fine. But you did not re-check the awkward cases from last week, and you cannot remember exactly what it used to say.
An evaluation is that checking, written down so it can be repeated.
What it actually is
Three things:
Cases. Questions you would ask the agent, with your notes on what a good answer looks like. Twenty of them, say.
Checks. What "good" means, made explicit. Did it stay on topic? Did it stick to the documents rather than inventing? Did it use the right tool? Did it come back fast enough?
A run. Asteria puts every case to your agent and applies every check, then shows you a scorecard.
Change the agent, run it again, and compare. That is the whole idea, and everything else is detail.
Why it is worth the afternoon
You cannot remember what it used to do. Nobody can. A run from three weeks ago can.
A fix in one place breaks something elsewhere. You add "always cite the source" and it starts refusing questions it used to answer. Twenty cases catch that in five minutes.
"It seems better" is not a reason to ship. Especially when you are about to publish the agent to your whole organisation.
You will inherit an agent one day. Its evaluation is the only honest description of what it is supposed to do.
When to bother
Not for every agent. An agent you built for yourself, use twice a week and can eyeball is fine without one. The effort only pays back when being wrong has a cost or an audience.
Worth it when:
- The agent is shared, so somebody else relies on it being right
- It touches something consequential: numbers, policy, anything customer-facing
- You are changing it regularly and want to know whether you are improving it
- It runs unattended on a schedule, where nobody sees a bad answer until later
The vocabulary
| Word | Means |
|---|---|
| Dataset | A set of cases plus the checks to apply |
| Case | One question, optionally with a reference answer |
| Evaluator | One check, applied to every case |
| Run | One execution of a dataset against the agent |
| Revision | A saved version of the agent's setup, created every time you save |
| Template | A dataset published by a colleague or shipped with the product, ready to copy |
Where it lives
Two places, for two different jobs.
On the agent, an Evaluations section: build datasets, run them, read results for that agent.
Evaluations in the left rail of the Agents area, a dashboard across every agent you can see, with a "needs attention" filter. Useful once you have more than a couple.
Next
- Building a dataset: writing cases, generating them, starting from a template.
- Choosing checks: what each kind of evaluator is for.
- Reading the results: the scorecard, the matrix, and comparing two runs.
- Tutorial: evaluating an agent.