Trainer

The trainer runs fine-tuning and evaluation as durable operations on Akka. Use it to manage runs, checkpoints, model registration, and resource cleanup. You can pause and resume a run, branch from a saved snapshot, and compare the resulting models on the same evaluation dataset.

Training and evaluation are separate operations. A completed training run registers its model. You evaluate that model in a separate step, or combine training and evaluation in a repeatable pipeline.

Use the akka CLI with a trainer deployed on the Akka platform. For more information, see Akka Optimize CLI. Concepts defines the terms in these pages. Train a first model shows how to train your first model from start to finish.

The lifecycle

A training cycle takes the following steps:

  1. Select a base model. A run starts from one, and a deployment allows a fixed set of them. See Base models.

  2. Register a dataset. A dataset is immutable and is identified by its content hash. A name is a movable alias. See Datasets.

  3. Create a workload for the task. It specifies the evaluation dataset and grader, and contains references to the base, candidate, and champion models. See Workloads.

  4. Start a run. A run descriptor names the workload, the base model, the training method, and the hyperparameters. See Training runs.

  5. Observe the run’s phase, progress, metrics, and cost. See Observe a run.

  6. Evaluate the resulting model on the workload’s evaluation dataset. Promote it when it scores better than the model the workload serves. See Evaluation.

To try a model with your own requests at any point, create an endpoint for it. See Endpoints. An endpoint isn’t part of the cycle. Training and evaluation don’t require one, and an endpoint reserves devices until you delete it.

Training methods

The trainer supports the following training methods:

Method Learns from

SFT

The assistant messages in the dataset.

GRPO, GSPO

Sampled completions that a configured reward scores. A reward term can call a function from a scoring bundle, call a model, or use the workload’s grader.

A reinforcement run can warm-start from a registered adapter.

Checkpoints and branches

A snapshot is a verified, immutable save from one run attempt. A run keeps one identity across attempts, and resuming it creates another attempt. Branching starts a new run from an exact snapshot of its parent. The new run has its own identity, budget, metrics, and outputs. For more information, see Checkpoints and Branches.

Pipelines

A pipeline submits ordered train and evaluate stages together. A later stage can refer to the output of an earlier stage. For more information, see Pipelines.