Model judges

Feature set: Evaluations Contact our support for access.
This functionality evolves quickly, the behavior and APIs might change between releases without further notice.

A judge decides how well a reply meets a criterion stated as a sentence.

What a judge decides

The built-in evaluators compare the reply and the tool calls against exact values. Some criteria have no exact value. A Judge scores such a criterion between 0 and 1, and the score becomes a result on the case with the label judge.

A judged score can differ between runs of the same reply. Gate a batch on the rate the judge passed rather than asserting on one judged case, see Gating a judged batch.

Asking a model through the TestKit

Judge.modelBased(testKit) returns a ModelBasedJudge that asks a model through JudgeAgent. The verdict is one model call. It is not deterministic and consumes tokens. JudgeAgent is an agent with the component id eval-judge. The TestKit registers it when it starts; it is not part of the deployed service.

The judge asks the model configured in akka.javasdk.agent.model-provider, or the one named with withModel:

var judge = Judge.modelBased(testKit).withModel("eval.judge-model");

A model provider registered for JudgeAgent.class in the TestKit settings takes precedence over both. TestKit.Settings.withModelProvider(JudgeAgent.class, provider) accepts any ModelProvider, so it is another way to give the judge its own model.

Each verdict is decided in a new session with memory disabled, so one verdict does not influence the next.

Stating the criterion

A criterion is one sentence that states what a good reply does, for example "the reply states only what the tools returned". The judge scores the reply against it and returns the score with a one-sentence reason.

shouldSatisfy(criterion) and shouldScoreAtLeast(criterion, threshold) turn the judge into an Evaluator.

  • shouldSatisfy passes at a score of at least 0.5.

  • shouldScoreAtLeast passes at the given threshold.

The judge holds no criterion of its own, so one judge serves any number of them. Each criterion gets one score and one reason. A criterion that combines several properties, such as an apology, an amount and a length, still gets one score, and the reason does not say which property failed. State one property per criterion, so that a failed result names the property that failed.

A criterion a program can decide, such as the length of the reply, belongs in a custom evaluator, not in a judge. Judge has one method, decide(criterion, interaction), so a test can implement it as a lambda that returns a fixed Verdict to run judged cases without a model.

@Test
public void theReplyApologizesAndStatesTheAmount() {
  var judge = Judge.modelBased(testKit); (1)

  var refund = EvalCase.of(
    "judged-refund",
    "Order o_9 arrived broken. I want my money back.",
    Evaluators.shouldCallTool("issueRefund"),
    judge.shouldSatisfy("the reply apologizes and states the refunded amount") (2)
  );

  var report = new ExperimentRunner(testKit).cases(refund).agent(OrderAgent::ask).run();

  assertThat(report.passed()).withFailMessage(report::render).isTrue();
  assertThat(report.results().getFirst().describe()).contains("judge:"); (3)
}
1 A judge bound to the running TestKit. It asks the model configured in application.conf.
2 The criterion, in words. The result has the label judge.
3 The case description contains the judged result as one line. The line starts with the verdict and the evaluator name, then the criterion, the score, the threshold, and the reason the model gave:
PASS judge: the reply apologizes and states the refunded amount: scored 0.90, needed 0.50 — apologizes and refunds 49.99

What the judge sees

The judge receives the criterion and the Interaction: what the agent was asked, the reply, and the tool calls with their arguments, results and errors.

What the agent was asked is the user message the agent sent the model, read from the runtime trace. That is the text the model under test saw, which is not always the case’s command: a command handler builds the user message from the command, and may wrap it or leave fields of it out. When the trace carries no user message, the case’s command is used instead.

ModelBasedJudge renders the sections as text and sends it as the user message. For the judged refund case in Stating the criterion the default user message is:

Criterion:
the reply apologizes and states the refunded amount

The agent was asked:
Order o_9 arrived broken. I want my money back.

The agent replied:
I am sorry the order arrived broken. I have refunded 49.99 to your card.

Tools called, in order:
- getOrder {orderId=o_9} -> {"id":"o_9","status":"delivered","totalCents":4999,"note":""}
- issueRefund {orderId=o_9, amountCents=4999} -> {"orderId":"o_9","amountCents":4999,"reference":"ref_1"}

A tool call that failed shows → failed: and the error instead of a result. A case that called no tool has no Tools called section.

The system message is ModelBasedJudge.DEFAULT_SYSTEM_MESSAGE:

public static final String DEFAULT_SYSTEM_MESSAGE =
    """
    You judge whether an agent's reply meets a stated criterion.

    You will be given the criterion, the input the agent was given, the reply the agent
    produced, and the tools it called to produce that reply. Score how fully the reply meets the
    criterion, between 0 (does not meet it) and 1 (fully meets it). Judge the criterion you
    were given and nothing else: a reply you would have worded differently still meets a
    criterion it satisfies.
    """;

The judge appends ModelBasedJudge.REPLY_FORMAT to the system message, which asks for the reply as a JSON object with a score between 0 and 1 and a reason.

withSystemMessage(text) returns a judge with another system message. The reply format is appended to that message too, so a custom message does not need to ask for it. systemMessage() returns the system message a judge sends, with the reply format appended.

withUserMessage(function) returns a judge that renders the criterion and the Interaction into the user message with the given function, for example to label the sections in the language of the cases.

ModelBasedJudge.defaultUserMessage(criterion, interaction) is the default rendering in What the judge sees. A custom rendering must carry what the system message asks the model to judge.

When the judge is inconclusive

The judge evaluator is inconclusive, and the case is not failed, when:

  • the reply is empty or blank, so there is nothing to judge;

  • the judge throws, for example because the model is unavailable;

  • the score is not between 0 and 1.

The inconclusive result carries the reason and appears in the report as INCONCLUSIVE judge: followed by the criterion and the reason. Inconclusive results describes how the report and the gates count inconclusive results.

Gating a judged batch

With a real model, hold a batch to the rate the judge passed. Gate.passRateShouldBeAtLeast(Evaluators.JudgeEvaluator.class, minimumRate) rates the judge evaluator over the cases where it was conclusive, and fails when it judged no case.

@Test
public void repliesStayFactual() {
  var judge = Judge.modelBased(testKit);
  var factual = judge.shouldScoreAtLeast(
    "the reply states only what the tools returned",
    0.7
  ); (1)

  var report = new ExperimentRunner(testKit)
    .cases(curated())
    .evaluator(factual)
    .agent(OrderAgent::ask)
    .gate(Gate.passRateShouldBeAtLeast(Evaluators.JudgeEvaluator.class, 0.8)) (2)
    .run();

  assertThat(report.passed()).withFailMessage(report::render).isTrue();
}
1 One evaluator with a threshold of 0.7. evaluator(factual) on the runner adds it to every case of the experiment, next to the evaluators each case declares.
2 At least 80% of the judged cases must pass. The other evaluators of each case still apply.

See also