Gates and reports
A gate is what a batch of cases must satisfy over all case results, and the report is what a run returns.
Gates
A Gate is what the batch must satisfy over all case results.
| Gate | Passes when |
|---|---|
|
No case has a failed result. This is the gate that applies when none is given. |
|
The share of cases with no failed result is at least |
|
The share of the evaluator’s results that passed is at least |
|
No case failed while loading its recorded calls or in the agent call. |
|
Both gates hold. The verdict lists both details. |
A failed gate does not throw. Assert on the report, and use the rendered report as the failure message.
Reading the report
run() returns an EvalReport.
| Accessor | Content |
|---|---|
|
Whether the gate passed. |
|
The share of cases with no failed result. |
|
One |
|
The report as text. |
A rendered report of a passing batch:
3/3 cases passed (100%) gate: passed — pass rate 1.00 over 3 cases, required 0.90; tool-arguments rate 1.00 over 2 judged cases, required 1.00; no target failures tools 1/1 tool-arguments 2/2 forbidden-tools 2/2 reply-contains 2/2 tool-order 1/1 reply-matches 1/1 tool-call-budget 1/1 model-call-budget 1/1 spend: 6 model calls, 123 tokens in, 321 out, 550 ms in total, slowest order-status at 489 ms, over 3/3 cases with evidence
Line by line:
-
The cases with no failed result, and the pass rate.
-
The gate’s verdict and its detail. A combined gate lists the detail of each part.
-
One line per evaluator, under its label: the cases it passed over the cases it judged. An evaluator that was inconclusive on some cases shows the count in brackets, for example
token-budget 0/0 (2 inconclusive). -
The spend: model calls, tokens and latency summed over the cases whose trace carried model calls, and the slowest case. The line is absent when no case carried model calls.
-
After these lines, every failed case with its evidence, see Reading a failed case.
Reading a failed case
The rendered report ends with every failed case and its evidence.
This report is from a case that asks where order o_42 is, with an evaluator that expects getOrder to be called with o_43:
0/1 cases passed (0%)
gate: FAILED — failed cases [wrong-order]
tool-arguments 0/1
spend: 2 model calls, 123 tokens in, 321 out, 6 ms in total, slowest wrong-order at 6 ms, over 1/1 cases with evidence
case wrong-order FAILED
reply: Order o_42 is shipped.
tools: getOrder{orderId=o_42}
model: 2 calls, 123 tokens in, 321 out, 6 ms
FAIL tool-arguments: getOrder(orderId) expected o_43, was [o_42]
A failed case prints the reply, the tool calls with their arguments, the model spend, and every result with its detail.
When the agent mapped the model’s text into the reply, a model text: line shows the model’s own text as well.
The reply is cut to one line of at most 200 characters.