Building a dataset
Twenty good questions beat two hundred mediocre ones. Here is how to pick them.
A dataset is a named set of cases and the checks applied to them. An agent can have several: one for everyday questions, one for the awkward ones, one for the things it should refuse.
New dataset, give it a name and a line about when to run it, and start adding cases.
What makes a case
Input, the question, exactly as a real person would type it. Type the messy version, not the tidy one, because the messy version is what your colleagues will send.
Expected output (optional), a reference answer. Some checks need one, most do not. Write it when there is a definitively right answer, and leave it empty when there is not.
Metadata (optional), your own labels such as {"difficulty": "hard"}, useful for making sense of results later.
Evaluator overrides (optional), checks for this case alone, otherwise it inherits the dataset's.
Choosing your twenty
This is the part that determines whether the whole exercise is useful, and the instinct (write twenty ordinary questions) produces a dataset that passes forever and teaches you nothing.
A good spread:
A few genuinely ordinary questions. The bread and butter. If these ever fail, something is badly wrong.
The awkward ones. Ambiguous, under-specified, or with an assumption buried in them. "How many days do I get?" when the answer depends on the country.
The ones it should refuse or hedge. Outside its remit, or where the honest answer is "the documents do not say". This is the group people leave out, and it is where agents most often fail: confidently answering something they should have declined.
The ones that bit you. Every time the agent gets something wrong in real use, add it as a case. Over a couple of months this becomes the most valuable part of the dataset, because it is a list of mistakes it can no longer make quietly.
One or two with a trap. A question with a false premise. "Why did we discontinue the Pro plan?" when you did not. Does it correct you, or invent a reason?
Twenty cases chosen like that tell you more than two hundred variations on "what is our refund policy".
Generating cases
If staring at a blank list is the blocker, Generate drafts some for you: describe the domain, say how many, optionally give a couple of examples for flavour.
You then review every draft, editing, deleting or regenerating individually before accepting. Nothing is saved until you accept.
Generated cases are a starting point, not a dataset. They tend to cluster around the obvious, which is exactly the group that teaches you least.
Use them for the ordinary questions, then add the awkward ones, the refusals and the traps by hand. Those are the ones that come from knowing the job, and no generator knows your job.
Starting from a template
Start from a template copies a dataset that somebody in your organisation published, or a built-in starter from Asteria Labs.
You get your own copy. Later changes to the original never reach it, and your edits never touch theirs.
Publish your own with Publish to catalog, and everyone in your organisation can read and copy every test case in it. Removing it later does not affect copies people already made.
Shared question banks
When several agents fork the same template, the Shared question banks view groups them and ranks them by their latest pass rate. Useful for a standard everybody is meant to meet: an agent well below the others is visible.
Read that ranking with one caveat in mind, which the view states: forking copies the questions at that moment. A copy edited since may no longer be asking the same things, so a lower score can mean a harder dataset rather than a worse agent.
Keeping a dataset honest
Add the real failures. The single highest-value habit. Real mistakes beat imagined ones.
Do not edit a case to make it pass. Tempting, and it destroys the point. If a case is genuinely unfair, fix it. If it is fair and failing, that is the dataset doing its job.
Prune the ones that never fail. A case that has passed forty runs running costs money and tells you nothing. Keep a few as canaries and retire the rest.
Leave it alone when comparing. Changing cases and the agent at the same time means you cannot tell which moved the score.