Evaluation

An evaluation scores a registered model against a dataset with a grader. Use evaluations to measure a model after training or to compare it with a baseline. Training and evaluation are separate operations: finishing a run registers its model, and you score that model in a separate step. A score belongs to the evaluation, so you can evaluate a model repeatedly.

Models

List models, then read one:

akka-optimize trained-models list
akka-optimize trained-models get MODEL_ID

get shows the run that produced the model, the model’s lineage, and all of its evaluations.

A completed run registers its final adapter. To register a snapshot from an earlier point in a run, use the following commands:

akka-optimize training snapshots register SNAPSHOT_ID
akka-optimize training snapshots candidate SNAPSHOT_ID

Use the full snapshot ID. Registration retains the weights and exports a serving copy before creating the entry. The trainer accepts the request before registration finishes. candidate reports its progress. Wait for REGISTERED before using the model ID. Registering the same snapshot again returns the same model. You can’t register a snapshot that contains only training state.

Baselines

Compare a candidate with a baseline. Register the base model so that you can evaluate it on the same dataset with the same grader. The following descriptor from the support triage example registers a base model:

{
  "servingModel": {
    "variant": "BASE",
    "baseModel": "unsloth/Llama-3.2-1B-Instruct",
    "revision": "5a8abab4a5d6f164389b1079fb721cfab8d7126c"
  }
}

To register it, run:

akka-optimize trained-models register -f base-model.json

A base-model entry pins a repository commit as part of its identity. Without an explicit revision, registration resolves the head of the repository. Registering a newer commit creates a separate entry, while existing lineage continues to refer to the earlier one. Training uses the commit pinned by the entry, and an evaluation job loads the same commit. A serving process therefore doesn’t select the weights for either operation.

The same command can register an adapter trained elsewhere as a baseline or judge.

A base model doesn’t belong to a workload, so specify the workload for its score when you evaluate it:

BASE=$(akka-optimize trained-models register -f base-model.json -o json --jq .modelId)
BASELINE=$(akka-optimize training evaluations start --model "$BASE" \
  --dataset support-triage-eval-frontier --workload support-triage -o json --jq .evaluationId)
akka-optimize training evaluations wait "$BASELINE" --exit-status

A candidate records the workload that trained it, so it doesn’t require --workload. The trainer rejects a different workload.

Evaluate a model

Start an evaluation, then read its detail, report, and predictions:

akka-optimize training evaluations start --model MODEL_ID --dataset triage-eval --wait --exit-status
akka-optimize training evaluations get EVALUATION_ID
akka-optimize training evaluations report EVALUATION_ID
akka-optimize training evaluations predictions EVALUATION_ID
akka-optimize training evaluations list --model MODEL_ID

training evaluations start takes the following options:

Option Default Description

--model

none

The model to score, by ID. Required.

--dataset

none

The evaluation dataset, by hash or a unique prefix of one. Required.

--grader

json-label-match

The grader: json-label-match, model:MODEL, or http:URL. See Graders and rewards.

--workload

The model’s workload

The workload the score belongs to, by name or ID. A base model belongs to no workload, so set it to score a baseline. For a candidate, it must match the candidate’s workload.

--wait

off

Wait for the evaluation to finish.

--exit-status

off

Exit with a non-zero status if the evaluation fails or is cancelled. Requires --wait.

To cap the length of a scored answer, use an evaluate stage in a pipeline with maxCompletionTokens. See Pipelines.

When the workload uses a grader other than json-label-match, pass it with --grader. training workloads scores tabulates only evaluations whose dataset and grader match the workload’s evaluation settings.

The report shows the mean score, exact-match rate, and per-field accuracy, with the weakest fields first. It marks examples for which the grader returned no verdict as unscored and excludes them from the mean. predictions lists each example’s position in the split, prompt, expected answer, model answer, and score. It marks an unscored example instead of assigning it a zero.

An evaluation fails if it can’t reach the grader. training evaluations cancel stops an evaluation in progress.

Comparisons

A comparison evaluates one source snapshot and selected branch models with one pinned dataset, grader, and completion-token limit. Use a comparison to score branches under the same conditions. The comparison registers the source as a candidate if necessary.

Create a comparison, then read it:

akka-optimize training comparisons create \
  --source SNAPSHOT_ID --candidate MODEL_ID_A --candidate MODEL_ID_B \
  --dataset triage-eval --grader json-label-match --max-completion-tokens 64 \
  -o json --jq .comparisonId
akka-optimize training comparisons get COMPARISON_ID

Full IDs are required. Creation starts durable preparation, and you can read the comparison while its phase is PREPARING or RUNNING. COMPLETED means that every target completed. PARTIAL shows completed results with failed, refused, or uncertain targets. FAILED means that no target completed. An incomplete comparison has no winner.

The contract fingerprint identifies the resolved dataset, grader, and completion limit. A zero limit is uncapped. Results include the configuration differences between targets and their training costs. The comparison counts a selected run once and excludes a selected ancestor from the ancestry total. These values are the costs of complete runs, rather than the cost up to the source checkpoint.

Promote a model

When a candidate scores better than the model that the workload serves, update the references:

akka-optimize training workloads refs set triage candidate MODEL_ID
akka-optimize training workloads refs set triage champion MODEL_ID
akka-optimize training workloads history triage

The history records every change, so you can restore an earlier reference.