<!-- <nav> -->
- [Akka](../../index.html)
- [Testing](../index.html)
- [Evaluation](index.html)

<!-- </nav> -->

# Evaluation

Feature set: Evaluations Contact [our support](https://akka.io/contact) for access.
*This functionality evolves quickly, the behavior and APIs might change between releases without further notice.* Evaluation runs an agent through a set of eval cases and checks each reply, and the tool calls behind it, with evaluators.

## <a href="about:blank#_overview"></a> Overview

Evaluation is part of the Akka TestKit: the `akka.javasdk.testkit.eval` package in the `akka-javasdk-testkit` artifact.
An evaluation is a JUnit test that extends `TestKitSupport`.
The agent runs in the in-process runtime with its tools, session memory and guardrails.
The runtime traces every model call and tool call, and the evaluation reads that trace as evidence.
Nothing is persisted and nothing leaves the test JVM.

Use evaluation for three purposes.

- Measure the quality of an agent against a real model. Run a set of cases as one batch. A gate states what the case results of the batch must satisfy. A real model does not give the same reply twice, so set a gate that requires a minimum pass rate over all cases instead of the default gate, which requires every case to pass.
- Check that an agent still behaves as it did in production. `EvalCaseParser` turns recorded production interactions into cases. Each replayed case expects the agent to call the same tools, in the same order and with the same arguments as in the recording. When the recording carries the model call count, the tokens and the latency, the case also holds the agent to that spend.
- Attack an agent and check that it holds. The user message of the case is an attack, and its evaluators state what a breached reply looks like. See [Adversarial tests](adversarial.html).
Evaluation is not the place for the following.

- The effect one command produces. Use [Unit testing](../unit.html) and the agent’s `TestKit`.
- An endpoint’s HTTP contract, or the wiring between components without an agent. Use [Integration testing](../integration.html).

## <a href="about:blank#_how_a_case_runs"></a> How a case runs

`ExperimentRunner` runs the cases one after the other, in the calling thread.

1. The case’s recorded tool calls, if any, are loaded into the stubs through the runner’s `ToolBindings`.
2. The runner calls the agent’s command handler through the TestKit component client, in a fresh session, with the case’s command.
3. The runner reads the trace the runtime wrote for that session through `TelemetryReader`: the tool calls with their arguments, results and errors, the model calls with their token counts, the guardrail evaluations and the latency.
4. Every evaluator of the case reads the reply and the trace and reports a result: pass, fail, or inconclusive when the evidence it needs is absent.
5. The gate checks the aggregated results and the report is returned.
flowchart LR
    Case[EvalCase] --> Runner[ExperimentRunner]
    Runner -->|command, fresh session| Agent[Agent in the TestKit runtime]
    Agent -->|reply| Runner
    Agent -.->|trace| Telemetry[TelemetryReader]
    Telemetry --> Runner
    Runner --> Evaluators[Built-in evaluators, judges, custom evaluators]
    Evaluators --> Results[Results per case]
    Results --> Gate
    Gate --> Report[EvalReport] A case whose recorded calls fail to load fails with a result labeled `setup`.
A case whose agent call throws fails with a result labeled `target`.
It does not abort the batch.

## <a href="about:blank#_the_parts_of_an_evaluation"></a> The parts of an evaluation

| Type | Role |
| --- | --- |
| `EvalCase` | One case: an id, the command sent to the agent, the recorded tool calls, and the evaluators. See [Eval cases and evaluators](eval-cases.html). |
| `Evaluators` | The built-in evaluators: tool calls, order and arguments, reply text, and budgets. See [Eval cases and evaluators](eval-cases.html). |
| `Evaluator` | A custom check over the reply and the tool calls. See [Writing a custom evaluator](eval-cases.html#custom-evaluators). |
| `ExperimentRunner` | Builds the experiment step by step: the cases with their evaluators and tool bindings, the agent, the gate, then the run. See [Writing the first case](getting-started.html#first-case). |
| `Gate` | What the batch must satisfy over all case results. Without a gate every case must pass. See [Gates and reports](gates-and-reports.html). |
| `EvalReport` | The outcome: whether the gate passed, the pass rate, every case result, and the report as text. See [Gates and reports](gates-and-reports.html). |
| `Judge` | A model that scores the reply against a criterion stated as a sentence. See [Model judges](judges.html). |
| `EvalCaseParser`, `RecordedCall`, `ToolBindings` | Cases derived from recorded interactions, and the stubs their recorded results are loaded into. See [Replaying recorded interactions](replay.html). |
All types are in the `akka.javasdk.testkit.eval` package.
The API documentation is at [akka.javasdk.testkit.eval](../_attachments/testkit/akka/javasdk/testkit/eval/package-summary.html).

## <a href="about:blank#_see_also"></a> See also

- [Getting started with evaluation](getting-started.html)
- [Adversarial tests](adversarial.html)
- [Testing the agent](../../sdk/agents/testing.html)
- [LLM evaluation](../../sdk/agents/llm_eval.html)

<!-- <footer> -->
<!-- <nav> -->
[Integration](../integration.html) [Getting started](getting-started.html)
<!-- </nav> -->

<!-- </footer> -->

<!-- <aside> -->

<!-- </aside> -->