A demo shows that an agent can complete a task once, on an example chosen by the person giving the demo. It says nothing about the next hundred cases. The gap between a convincing demo and a dependable system is closed by testing, and testing an agent is different from testing ordinary software.
Why normal tests are not enough
Ordinary code gives the same output for the same input, so a test either passes or fails. A language model can answer the same request in slightly different ways each time. Two different answers can both be correct, and an answer that looks right can be wrong. Testing has to measure how often the agent is right, not whether one run passed.
Build the evaluation set from real cases
Collect real examples of the task: actual tickets, actual invoices, actual requests. For each one, record what the correct outcome is. Fifty is a workable start, a few hundred is better.
Include the difficult ones on purpose. Requests with missing information, contradictory details, a customer writing in two languages, and cases where the right action is to do nothing and ask a person. An evaluation set made only of clean examples will report a success rate that production never matches.
Score the steps, not only the result
An agent can reach the right answer for the wrong reason, for example by guessing a value it should have looked up. Check the path as well as the destination.
- Did it call the right tools, with correct arguments?
- Did it avoid actions it should not take?
- Did it stop and ask for help when it lacked information?
- How many steps and how much model usage did it need?
The last point matters for cost. An agent that solves the task in four steps and one that needs thirty are not equally good, even if both finish.
Decide what good enough means before you measure
Agree on the threshold in advance. For a task where every output is reviewed by a person, a lower success rate may be acceptable because errors are caught. For anything that acts on its own, the bar is higher. Setting the number after seeing the results invites wishful thinking.
Run the set on every change
Changing a prompt, switching to a newer model or adding a tool can improve one kind of case and break another. Run the full evaluation set after each change and compare with the previous result. This is the agent equivalent of a regression test, and it is what makes it safe to keep improving the system after launch.
Test in production, carefully
Before the agent acts for real, run it in shadow mode: it processes live cases and proposes actions, a person does the work as usual, and the two are compared. This reveals situations the evaluation set did not contain. After launch, review a sample of real runs each week and add the failures to the set.
Summary
Test with real cases, include the hard ones, score the steps as well as the outcome, fix the pass mark in advance and rerun on every change. We build this evaluation set as part of every AI agent project. If you have an agent that works in demos and want to know whether it is ready, talk to us.