Asteria Docs
Evaluations

Evaluations

How you find out whether your agent is any good, and whether your last change helped.

You built an agent. You tried it a few times, it seemed fine, you shared it.

Then you improved the instructions. Is it better now? You try it twice, it seems fine. But you did not re-check the awkward cases from last week, and you cannot remember exactly what it used to say.

An evaluation is that checking, written down so it can be repeated.

What it actually is

Three things:

Cases. Questions you would ask the agent, with your notes on what a good answer looks like. Twenty of them, say.

Checks. What "good" means, made explicit. Did it stay on topic? Did it stick to the documents rather than inventing? Did it use the right tool? Did it come back fast enough?

A run. Asteria puts every case to your agent and applies every check, then shows you a scorecard.

Change the agent, run it again, and compare. That is the whole idea, and everything else is detail.

Why it is worth the afternoon

You cannot remember what it used to do. Nobody can. A run from three weeks ago can.

A fix in one place breaks something elsewhere. You add "always cite the source" and it starts refusing questions it used to answer. Twenty cases catch that in five minutes.

"It seems better" is not a reason to ship. Especially when you are about to publish the agent to your whole organisation.

You will inherit an agent one day. Its evaluation is the only honest description of what it is supposed to do.

When to bother

Not for every agent. An agent you built for yourself, use twice a week and can eyeball is fine without one. The effort only pays back when being wrong has a cost or an audience.

Worth it when:

  • The agent is shared, so somebody else relies on it being right
  • It touches something consequential: numbers, policy, anything customer-facing
  • You are changing it regularly and want to know whether you are improving it
  • It runs unattended on a schedule, where nobody sees a bad answer until later

The vocabulary

WordMeans
DatasetA set of cases plus the checks to apply
CaseOne question, optionally with a reference answer
EvaluatorOne check, applied to every case
RunOne execution of a dataset against the agent
RevisionA saved version of the agent's setup, created every time you save
TemplateA dataset published by a colleague or shipped with the product, ready to copy

Where it lives

Two places, for two different jobs.

On the agent, an Evaluations section: build datasets, run them, read results for that agent.

Evaluations in the left rail of the Agents area, a dashboard across every agent you can see, with a "needs attention" filter. Useful once you have more than a couple.

Next

On this page