Replaying recorded interactions

Feature set: Evaluations Contact our support for access.
This functionality evolves quickly, the behavior and APIs might change between releases without further notice.

A replayed case sends a recorded user message and expects the agent to use its tools the way production did.

EvalCaseParser reads a JSONL file: one JSON object per line, one interaction per object. Each object becomes one EvalCase. What the case checks depends on which fields the recording carries. This page adds the fields one group at a time: the message and the reply, then the tool calls, then the spend.

The message and the reply

The smallest recording carries the user message and the reply production gave.

{"id":"prod-3","input":"hi, anyone there?","output":"Please give me your order number and I will look it up."}
Field Required Becomes

id

no

The case id. replay-<line> when absent.

input

yes

The command the runner sends to the agent. Text for a command handler that takes a String.

output

no

Nothing. The reply is kept for the reader; the parser does not check it, because a reply worded differently is not a regression.

A case made from these fields alone has no evaluators. It runs, collects the reply and the trace, and appears in the report, but nothing can fail. Add the evaluators yourself: give an evaluator to the runner with evaluator(…​) to run it on every parsed case, or to one case with withEvaluators(…​).

@Test
public void recordedRepliesStayOnTopic() {
  var replayed = EvalCaseParser.parse(recording("/eval/replies.jsonl")); (1)
  var judge = Judge.modelBased(testKit);
  var onTopic = judge.shouldScoreAtLeast(
    "the reply asks for an order number or closes the conversation",
    0.7
  );

  var report = new ExperimentRunner(testKit)
    .cases(replayed)
    .evaluator(Evaluators.shouldNotCallTool("issueRefund")) (2)
    .evaluator(onTopic) (3)
    .agent(OrderAgent::ask)
    .run();

  assertThat(report.passed()).withFailMessage(report::render).isTrue();
}

private static Path recording(String resource) {
  try {
    return Path.of(OrderAgentQualityIntegrationTest.class.getResource(resource).toURI());
  } catch (URISyntaxException e) {
    throw new IllegalStateException(e);
  }
}
1 parse(path) reads the file. These recordings name no tools, so the cases need no bindings.
2 A built-in evaluator on every parsed case. No replayed turn may issue a refund.
3 A judge scores every reply against one criterion.

The file is read once, and every line that does not parse or has no input is listed with its line number in one IllegalArgumentException.

For an agent whose command handler takes its own type, record input as the JSON of that command and read the file with parse(path, commandType), which returns List<EvalCase<C>>.

The tool calls

A recording of a turn that used tools carries each call: the tool name, the arguments the agent passed, and the result production saw.

{"id":"prod-2","input":"Order o_9 arrived broken, please refund it.","toolCalls":[{"name":"getOrder","arguments":{"orderId":"o_9"},"result":{"id":"o_9","status":"delivered","totalCents":4999,"note":""}},{"name":"issueRefund","arguments":{"orderId":"o_9","amountCents":4999},"result":{"orderId":"o_9","amountCents":4999,"reference":"rf_1"}}],"output":"I am sorry about the damage. I have refunded 49.99 to your card, reference rf_1."}

toolCalls lists the calls in the order production made them. Each call has a name, arguments as an object, and the result as recorded JSON. An empty array records a turn without tool calls. The parser keeps each call on the case as a RecordedCall, in EvalCase.recordedCalls(). The parser does not touch the stubs; the runner does, through the bindings described next.

Binding tools to stubs

In the test, a stub serves each tool. For the agent to see what production saw, the stub must return the recorded result. A ToolBindings maps each tool name to a loader that puts the recorded result into the stub serving that tool. Give the bindings to the runner with bindings(…​). Before each case the runner loads the case’s recorded calls through them, in recorded order, then calls the agent. A loader that throws fails the case with a result labeled setup.

The runner refuses to start when a case names a tool without a binding. The error lists each unbound tool with the cases that name it, so a new tool in production cannot be replayed by accident against a stub that does not know it.

/** Adds the order a recorded getOrder call returned. */
public void loadOrder(RecordedCall call) {
  addOrder(call.resultAs(Order.class));
}

/** Adds the refund a recorded issueRefund call returned. */
public void loadRefund(RecordedCall call) {
  addRefund(call.resultAs(Refund.class));
}

RecordedCall.resultAs(type) reads the recorded result into the stub’s type, and argument(name) returns one recorded argument.

@Test
public void replayedTrafficStillHolds() {
  var replayed = EvalCaseParser.parse(recording("/eval/captures.jsonl")); (1)
  var bindings = ToolBindings.builder() (2)
    .bind("getOrder", orders::loadOrder)
    .bind("issueRefund", orders::loadRefund)
    .build();

  var report = new ExperimentRunner(testKit)
    .cases(replayed)
    .bindings(bindings) (3)
    .agent(OrderAgent::ask)
    .gate(Gate.passRateShouldBeAtLeast(0.9)) (4)
    .run();

  assertThat(report.passed()).withFailMessage(report::render).isTrue();
}
1 The cases, with their recorded calls as data.
2 One binding per tool the recordings name.
3 The runner loads each case’s recorded calls through the bindings before the agent is called.
4 The replayed batch runs like a curated one, with a gate for a real model.

What the tool calls assert

The recorded tool calls become evaluators on the case. They describe what production did. They are a baseline, not a statement of correctness.

  • Evaluators.shouldCallTools: every recorded tool is called.

  • Evaluators.shouldCallToolsInOrder: in the recorded relative order.

  • Evaluators.shouldCallToolWith: one per recorded argument, which must arrive with its recorded value.

For the recording of prod-2 the case gets shouldCallTools("getOrder", "issueRefund"), shouldCallToolsInOrder("getOrder", "issueRefund"), and three shouldCallToolWith evaluators: orderId on getOrder, and orderId and amountCents on issueRefund.

The spend

A recording can carry what the turn cost in production.

{"id":"prod-4","input":"Has order o_17 been delivered yet?","toolCalls":[{"name":"getOrder","arguments":{"orderId":"o_17"},"result":{"id":"o_17","status":"delivered","totalCents":1250,"note":""}}],"output":"Order o_17 has been delivered.","modelCalls":2,"tokens":{"input":380,"output":30},"latencyMs":1500}

All three fields are optional, and each must be a positive count.

  • modelCalls: how many model calls the turn took.

  • tokens: either a total, or an object with input and output.

  • latencyMs: how long the turn took, in milliseconds.

What the spend asserts

Each field present becomes a budget evaluator on the case.

  • modelCalls becomes shouldMakeAtMostModelCalls, as recorded.

  • tokens becomes shouldUseAtMostTokens, the recorded total multiplied by the tolerance.

  • latencyMs becomes shouldReplyWithin, the recorded latency multiplied by the same tolerance.

A rerun almost never spends the same as the recording, so the token and latency budgets allow more than the recorded figures. EvalCaseParser.parse(path) uses EvalCaseParser.DEFAULT_TOLERANCE, which is 1.5. parse(path, tolerance) sets another tolerance; 1.0 holds a case to exactly what was recorded, and a value below 1.0 is rejected. Under a TestModelProvider the token budget is inconclusive, because a mocked model reports no tokens.

The replayed batch

The complete file, one interaction per line:

{"id":"prod-1","input":"Where is order o_42?","toolCalls":[{"name":"getOrder","arguments":{"orderId":"o_42"},"result":{"id":"o_42","status":"shipped","totalCents":2599,"note":"leave it with a neighbour"}}],"output":"Order o_42 has been shipped."}
{"id":"prod-2","input":"Order o_9 arrived broken, please refund it.","toolCalls":[{"name":"getOrder","arguments":{"orderId":"o_9"},"result":{"id":"o_9","status":"delivered","totalCents":4999,"note":""}},{"name":"issueRefund","arguments":{"orderId":"o_9","amountCents":4999},"result":{"orderId":"o_9","amountCents":4999,"reference":"rf_1"}}],"output":"I am sorry about the damage. I have refunded 49.99 to your card, reference rf_1."}
{"id":"prod-3","input":"hi, anyone there?","toolCalls":[],"output":"Please give me your order number and I will look it up."}
{"id":"prod-4","input":"Has order o_17 been delivered yet?","toolCalls":[{"name":"getOrder","arguments":{"orderId":"o_17"},"result":{"id":"o_17","status":"delivered","totalCents":1250,"note":""}}],"output":"Order o_17 has been delivered.","modelCalls":2,"tokens":{"input":380,"output":30},"latencyMs":1500}
{"id":"prod-5","input":"thanks, that is all","toolCalls":[],"output":"You are welcome.","modelCalls":1,"tokens":120,"latencyMs":900}

Its report under the mocked model:

5/5 cases passed (100%)
gate: passed — pass rate 1.00 over 5 cases, required 0.90
  tools 3/3
  tool-order 3/3
  tool-arguments 5/5
  model-call-budget 2/2
  token-budget 0/0 (2 inconclusive)
  latency-budget 2/2
spend: 9 model calls, 0 tokens in, 0 out, 30 ms in total, slowest prod-2 at 7 ms, over 5/5 cases with evidence

Three recordings name tools, so three cases have tools and tool-order results, and their five recorded arguments give five tool-arguments results. Two recordings carry spend figures, so two cases have budgets, and their token budgets were inconclusive.