Eval cases and evaluators

Feature set: Evaluations Contact our support for access.
This functionality evolves quickly, the behavior and APIs might change between releases without further notice.

An eval case is one turn with the agent and the evaluators its reply and tool calls must pass.

The parts of an eval case

EvalCase is a record with four parts.

Part Meaning

id

Names the case in the report. Unique within the suite.

command

What the agent’s command handler receives: the user message as a String, or the handler’s own type, such as a record with several fields. All cases of one experiment have the same command type, and the case names it: List<EvalCase<String>>, or List<EvalCase<MyCommand>> for a handler that takes MyCommand.

recordedCalls

The tool calls a recording carried, as data. The runner loads their results into the stubs through the ToolBindings given to it, see Binding tools to stubs. Empty for a case written by hand. EvalCase.of(id, command, evaluators…​) creates such a case.

evaluators

The checks the turn must pass, as a list or as trailing arguments. Empty when the case only collects evidence. withEvaluators(…​) returns the same case with more evaluators.

A case that throws while its recorded calls are loaded fails with a result labeled setup. The agent is not called.

Each case runs in a fresh session, so no case can read the session memory of an earlier case.

A model-based judge is asked about the user message the agent sent the model, which Interaction.userMessage() carries from the runtime trace. When the trace has no user message, the judge is asked about the case’s command instead, as text: a String command as is, any other command as its JSON. Interaction.asked() returns whichever of the two applies.

Preparing the stubs

The stubs the agent’s tools call are prepared before the experiment runs, in the test, for example in a JUnit @BeforeEach. Every case of a run reads the same stubs, so a stub holds the answers of all cases at once and keys them by the tool’s arguments: an order by its id, the routes by their destination. State that lives in the runtime, such as an entity a tool reads, is seeded the same way, through the component client before the run. A replayed case is the exception: its recorded results are loaded per case, see Binding tools to stubs.

Evaluators

An Evaluator is one check over the reply and the traced tool calls of a turn. It reports a result: pass, fail, or inconclusive when the evidence it needs is absent. A case carries any number of evaluators, and they are all of the same type:

  • the built-in evaluators, created through the factories in Evaluators;

  • a criterion scored by a model, created through a Judge;

  • a custom evaluator of your own.

An evaluator given to experimentRunner.cases(…​).evaluator(…​) runs on every case of the experiment instead of one.

A case with five built-in evaluators:

private EvalCase<String> fullRefund() {
  return EvalCase.of(
    "full-refund",
    "Order o_9 arrived broken. I want my money back.",
    Evaluators.shouldCallToolsInOrder("getOrder", "issueRefund"), (1)
    Evaluators.shouldCallToolWith("issueRefund", "amountCents", 4999), (2)
    Evaluators.replyShouldMatch("refund(ed)?.*49\\.99"), (3)
    Evaluators.shouldMakeAtMostToolCalls(2), (4)
    Evaluators.shouldMakeAtMostModelCalls(3)
  );
}
1 getOrder must be called before issueRefund. Other calls may come between.
2 issueRefund must receive amountCents as 4999. Numbers compare by value, so an int matches a recorded long.
3 The reply must match the regular expression somewhere. Anchor it for a full match.
4 Budgets. At most two tool calls and three model calls for this turn.

Evaluator labels

Every evaluator class carries an @EvalLabel: the label used in the report to group its results. A built-in evaluator uses its label as is, for example tools. A custom evaluator gets the custom-eval: prefix, for example custom-eval:refund-within-total. The prefix keeps custom results apart from built-in results. A case refuses an evaluator without a @EvalLabel. Labels must be unique across all evaluators in an experiment. The rule is per class, not per instance.

Gate.passRateShouldBeAtLeast(evaluatorClass, minimumRate) takes the evaluator’s class and reads the label from it. The built-in classes are nested in Evaluators, and the tables below list each one next to its label.

Built-in evaluators on tool calls

Label and class Created with Passes when

tools
Evaluators.Tools

Evaluators.shouldCallTools(names…​), Evaluators.shouldCallTool(name)

Every named tool was called, in any order. Other calls are allowed.

tool-order
Evaluators.ToolOrder

Evaluators.shouldCallToolsInOrder(names…​)

The named tools were called in this relative order. Calls to other tools may come between. Inconclusive when one of the named tools was never called.

tool-arguments
Evaluators.ToolArgument

Evaluators.shouldCallToolWith(tool, argument, value)

A call to the tool carried the argument with this value. Inconclusive when the tool was never called.

tool-results
Evaluators.ToolResult

Evaluators.toolResultShouldContain(tool, text)

The result of the tool contains the text, case-insensitively. Inconclusive when the tool was never called or the trace carries no result for it.

forbidden-tools
Evaluators.ForbiddenTools

Evaluators.shouldNotCallTools(names…​), Evaluators.shouldNotCallTool(name)

None of the named tools was called.

Built-in evaluators on the reply

Label and class Created with Passes when

reply-contains
Evaluators.ReplyContains

Evaluators.replyShouldContain(texts…​)

The reply contains every given text, case-insensitively.

reply-matches
Evaluators.ReplyMatches

Evaluators.replyShouldMatch(regex), Evaluators.replyShouldMatch(pattern)

The reply matches the regular expression anywhere. Anchor the expression for a full match.

reply-lacks
Evaluators.ReplyLacks

Evaluators.replyShouldNotContain(texts…​)

The reply contains none of the given texts, case-insensitively.

reply-does-not-match
Evaluators.ReplyDoesNotMatch

Evaluators.replyShouldNotMatch(regex), Evaluators.replyShouldNotMatch(pattern)

The reply matches the regular expression nowhere.

reply-lacks-payment-card
Evaluators.ReplyLacksPaymentCard

Evaluators.replyShouldNotContainPaymentCard()

The reply contains no payment card number. A card number is a sequence of 13 to 19 digits, with spaces and dashes allowed between them, that passes the Luhn checksum.

reply-lacks-luhn-number
Evaluators.ReplyLacksLuhnNumber

Evaluators.replyShouldNotContainLuhnNumber(minDigits, maxDigits)

The reply contains no sequence of digits within the length range that passes the Luhn checksum. The same checksum validates identifiers other than payment cards, each with its own length, so give the range of the one you are looking for.

The String forms of replyShouldMatch and replyShouldNotMatch compile the expression with Pattern.DOTALL, so . also matches line breaks. The Pattern forms use the pattern as given, with its own flags. A regular expression that does not compile throws PatternSyntaxException when the case is created.

replyShouldNotContain, replyShouldNotMatch, replyShouldNotContainPaymentCard and replyShouldNotContainLuhnNumber state what a reply must not carry. Adversarial tests builds cases from them to check an agent under attack.

Built-in budgets

Label and class Created with Passes when

tool-call-budget
Evaluators.ToolCallBudget

Evaluators.shouldMakeAtMostToolCalls(calls)

The agent made at most this many tool calls.

model-call-budget
Evaluators.ModelCallBudget

Evaluators.shouldMakeAtMostModelCalls(calls)

The agent made at most this many model calls. A reply takes at least one.

token-budget
Evaluators.TokenBudget

Evaluators.shouldUseAtMostTokens(tokens)

Input and output tokens together stay within the limit. Inconclusive when the model reported no tokens.

latency-budget
Evaluators.LatencyBudget

Evaluators.shouldReplyWithin(duration)

The agent command completed within the duration. Inconclusive when the trace carries no timing.

Custom evaluator

A custom evaluator is a named class annotated with @EvalLabel that implements Evaluator. evaluate(…​) receives the case as EvalCase<?> and the Interaction, the reply and everything the runtime traced for the turn, and returns an EvalResult: a pass, a fail with a reason, or an inconclusive result. In the report, its label gets the custom-eval: prefix, see Evaluator labels. An anonymous class or a lambda cannot carry the annotation, so a case refuses it.

RefundWithinTotal is an evaluator that checks that the amount the agent passed to issueRefund does not exceed the total of the order.

/** The refund the agent issued must not exceed the order's total. */
@EvalLabel("refund-within-total") (1)
public final class RefundWithinTotal implements Evaluator {

  private final int totalCents;

  public RefundWithinTotal(int totalCents) {
    this.totalCents = totalCents;
  }

  @Override
  public EvalResult evaluate(EvalCase<?> evalCase, Interaction interaction) {
    var refund = interaction
      .toolCalls()
      .stream()
      .filter(call -> call.name().equals("issueRefund"))
      .findFirst();
    if (refund.isEmpty()) {
      return EvalResult.inconclusive("no refund was issued"); (2)
    }
    var amount = ((Number) refund.get().arguments().get("amountCents")).intValue(); (3)
    return amount <= totalCents
      ? EvalResult.pass()
      : EvalResult.fail("refunded " + amount + " cents of a " + totalCents + " cent order");
  }
}
1 The label. The report shows it as custom-eval:refund-within-total. Gate.passRateShouldBeAtLeast(RefundWithinTotal.class, minimumRate) rates the results of this evaluator.
2 No refund call means nothing to check. Report an inconclusive result rather than a fail; the tools evaluator on the case reports the missing call.
3 Tool arguments arrive as the runtime decoded them from the model’s JSON: a number is a Number, an object is a Map.

It goes into a case next to the built-in evaluators.

@Test
public void aRefundNeverExceedsTheOrderTotal() {
  var refund = EvalCase.of(
    "refund-within-total",
    "Order o_9 arrived broken. I want my money back.",
    Evaluators.shouldCallTool("issueRefund"),
    new RefundWithinTotal(orders.getOrder("o_9").totalCents()) (1)
  );

  var report = new ExperimentRunner(testKit).cases(refund).agent(OrderAgent::ask).run();

  assertThat(report.passed()).withFailMessage(report::render).isTrue();
}
1 The evaluator carries what it needs to know about the case.

Inconclusive results

An evaluator reads only what it names and is inconclusive when that evidence is absent or does not support a verdict. A tool that was never called fails tools and makes tool-order and tool-arguments inconclusive, so one missing call is reported once. An inconclusive result is neither a pass nor a fail. A case with inconclusive results and no failed result passes. The report counts inconclusive results per evaluator label, and Gate.passRateShouldBeAtLeast(evaluatorClass, minimumRate) rates an evaluator over its conclusive results.

See also