Eval cases and evaluators
An eval case is one turn with the agent and the evaluators its reply and tool calls must pass.
The parts of an eval case
EvalCase is a record with four parts.
| Part | Meaning |
|---|---|
|
Names the case in the report. Unique within the suite. |
|
What the agent’s command handler receives: the user message as a |
|
The tool calls a recording carried, as data. The runner loads their results into the stubs through the |
|
The checks the turn must pass, as a list or as trailing arguments. Empty when the case only collects evidence. |
A case that throws while its recorded calls are loaded fails with a result labeled setup.
The agent is not called.
Each case runs in a fresh session, so no case can read the session memory of an earlier case.
A model-based judge is asked about the user message the agent sent the model, which Interaction.userMessage() carries from the runtime trace.
When the trace has no user message, the judge is asked about the case’s command instead, as text: a String command as is, any other command as its JSON.
Interaction.asked() returns whichever of the two applies.
Preparing the stubs
The stubs the agent’s tools call are prepared before the experiment runs, in the test, for example in a JUnit @BeforeEach.
Every case of a run reads the same stubs, so a stub holds the answers of all cases at once and keys them by the tool’s arguments: an order by its id, the routes by their destination.
State that lives in the runtime, such as an entity a tool reads, is seeded the same way, through the component client before the run.
A replayed case is the exception: its recorded results are loaded per case, see Binding tools to stubs.
Evaluators
An Evaluator is one check over the reply and the traced tool calls of a turn. It reports a result: pass, fail, or inconclusive when the evidence it needs is absent. A case carries any number of evaluators, and they are all of the same type:
-
the built-in evaluators, created through the factories in Evaluators;
-
a criterion scored by a model, created through a Judge;
-
a custom evaluator of your own.
An evaluator given to experimentRunner.cases(…).evaluator(…) runs on every case of the experiment instead of one.
A case with five built-in evaluators:
private EvalCase<String> fullRefund() {
return EvalCase.of(
"full-refund",
"Order o_9 arrived broken. I want my money back.",
Evaluators.shouldCallToolsInOrder("getOrder", "issueRefund"), (1)
Evaluators.shouldCallToolWith("issueRefund", "amountCents", 4999), (2)
Evaluators.replyShouldMatch("refund(ed)?.*49\\.99"), (3)
Evaluators.shouldMakeAtMostToolCalls(2), (4)
Evaluators.shouldMakeAtMostModelCalls(3)
);
}
| 1 | getOrder must be called before issueRefund. Other calls may come between. |
| 2 | issueRefund must receive amountCents as 4999. Numbers compare by value, so an int matches a recorded long. |
| 3 | The reply must match the regular expression somewhere. Anchor it for a full match. |
| 4 | Budgets. At most two tool calls and three model calls for this turn. |
Evaluator labels
Every evaluator class carries an @EvalLabel: the label used in the report to group its results.
A built-in evaluator uses its label as is, for example tools. A custom evaluator gets the custom-eval: prefix,
for example custom-eval:refund-within-total. The prefix keeps custom results apart from built-in results.
A case refuses an evaluator without a @EvalLabel. Labels must be unique across all evaluators in an experiment. The rule is per class, not per instance.
Gate.passRateShouldBeAtLeast(evaluatorClass, minimumRate) takes the evaluator’s class and reads the label from it.
The built-in classes are nested in Evaluators, and the tables below list each one next to its label.
Built-in evaluators on tool calls
| Label and class | Created with | Passes when |
|---|---|---|
|
|
Every named tool was called, in any order. Other calls are allowed. |
|
|
The named tools were called in this relative order. Calls to other tools may come between. Inconclusive when one of the named tools was never called. |
|
|
A call to the tool carried the argument with this value. Inconclusive when the tool was never called. |
|
|
The result of the tool contains the text, case-insensitively. Inconclusive when the tool was never called or the trace carries no result for it. |
|
|
None of the named tools was called. |
Built-in evaluators on the reply
| Label and class | Created with | Passes when |
|---|---|---|
|
|
The reply contains every given text, case-insensitively. |
|
|
The reply matches the regular expression anywhere. Anchor the expression for a full match. |
|
|
The reply contains none of the given texts, case-insensitively. |
|
|
The reply matches the regular expression nowhere. |
|
|
The reply contains no payment card number. A card number is a sequence of 13 to 19 digits, with spaces and dashes allowed between them, that passes the Luhn checksum. |
|
|
The reply contains no sequence of digits within the length range that passes the Luhn checksum. The same checksum validates identifiers other than payment cards, each with its own length, so give the range of the one you are looking for. |
The String forms of replyShouldMatch and replyShouldNotMatch compile the expression with Pattern.DOTALL, so . also matches line breaks.
The Pattern forms use the pattern as given, with its own flags.
A regular expression that does not compile throws PatternSyntaxException when the case is created.
replyShouldNotContain, replyShouldNotMatch, replyShouldNotContainPaymentCard and replyShouldNotContainLuhnNumber state what a reply must not carry.
Adversarial tests builds cases from them to check an agent under attack.
Built-in budgets
| Label and class | Created with | Passes when |
|---|---|---|
|
|
The agent made at most this many tool calls. |
|
|
The agent made at most this many model calls. A reply takes at least one. |
|
|
Input and output tokens together stay within the limit. Inconclusive when the model reported no tokens. |
|
|
The agent command completed within the duration. Inconclusive when the trace carries no timing. |
Custom evaluator
A custom evaluator is a named class annotated with @EvalLabel that implements Evaluator.
evaluate(…) receives the case as EvalCase<?> and the Interaction, the reply and everything the runtime traced for the turn, and returns an EvalResult: a pass, a fail with a reason, or an inconclusive result.
In the report, its label gets the custom-eval: prefix, see Evaluator labels.
An anonymous class or a lambda cannot carry the annotation, so a case refuses it.
RefundWithinTotal is an evaluator that checks that the amount the agent passed to issueRefund does not exceed the total of the order.
/** The refund the agent issued must not exceed the order's total. */
@EvalLabel("refund-within-total") (1)
public final class RefundWithinTotal implements Evaluator {
private final int totalCents;
public RefundWithinTotal(int totalCents) {
this.totalCents = totalCents;
}
@Override
public EvalResult evaluate(EvalCase<?> evalCase, Interaction interaction) {
var refund = interaction
.toolCalls()
.stream()
.filter(call -> call.name().equals("issueRefund"))
.findFirst();
if (refund.isEmpty()) {
return EvalResult.inconclusive("no refund was issued"); (2)
}
var amount = ((Number) refund.get().arguments().get("amountCents")).intValue(); (3)
return amount <= totalCents
? EvalResult.pass()
: EvalResult.fail("refunded " + amount + " cents of a " + totalCents + " cent order");
}
}
| 1 | The label. The report shows it as custom-eval:refund-within-total. Gate.passRateShouldBeAtLeast(RefundWithinTotal.class, minimumRate) rates the results of this evaluator. |
| 2 | No refund call means nothing to check. Report an inconclusive result rather than a fail; the tools evaluator on the case reports the missing call. |
| 3 | Tool arguments arrive as the runtime decoded them from the model’s JSON: a number is a Number, an object is a Map. |
It goes into a case next to the built-in evaluators.
@Test
public void aRefundNeverExceedsTheOrderTotal() {
var refund = EvalCase.of(
"refund-within-total",
"Order o_9 arrived broken. I want my money back.",
Evaluators.shouldCallTool("issueRefund"),
new RefundWithinTotal(orders.getOrder("o_9").totalCents()) (1)
);
var report = new ExperimentRunner(testKit).cases(refund).agent(OrderAgent::ask).run();
assertThat(report.passed()).withFailMessage(report::render).isTrue();
}
| 1 | The evaluator carries what it needs to know about the case. |
Inconclusive results
An evaluator reads only what it names and is inconclusive when that evidence is absent or does not support a verdict.
A tool that was never called fails tools and makes tool-order and tool-arguments inconclusive, so one missing call is reported once.
An inconclusive result is neither a pass nor a fail.
A case with inconclusive results and no failed result passes.
The report counts inconclusive results per evaluator label, and Gate.passRateShouldBeAtLeast(evaluatorClass, minimumRate) rates an evaluator over its conclusive results.
See also
-
Replaying recorded interactions derives cases with the built-in evaluators from recordings.
-
Adversarial tests states what a breached reply looks like, with the same evaluators.