Checkpoints
A snapshot is a verified, immutable save from one run attempt. Use snapshots to resume a paused run, start a branch, or register an earlier adapter. A snapshot contains adapter weights, full training state, or both. Full state includes the optimizer, scheduler, random-number generators, and data position. The catalog lists only committed snapshots, so you cannot restore from an interrupted upload.
execution.saveEvery in the run descriptor sets the save interval in optimizer updates. Without it, a run saves only at the end.
List the snapshots of a run, then read one:
akka-optimize training runs snapshots RUN_ID
akka-optimize training snapshots get SNAPSHOT_ID
Pause a run
Pause a run, then read its state:
akka-optimize training runs pause RUN_ID --checkpoint save
akka-optimize training runs get RUN_ID
--checkpoint is required. It states which checkpoint the run stops against:
| Policy | Behavior |
|---|---|
|
Saves at the next optimizer boundary, commits the snapshot, and then stops. |
|
Stops at the newest committed checkpoint. A resumed run repeats updates after that checkpoint. The trainer rejects this policy after a run has trained without any restorable state because the run would otherwise restart from the base model. |
The command returns when the trainer accepts the pause. Saving and device release continue in the background. Follow the run detail until the phase is PAUSED and pause.resourcesReleased is true. The pause object reports the save progress, acknowledged snapshot, deadline, and any failure. If training finishes during the pause, the run completes. Cancellation takes precedence over a pending pause.
The trainer marks a pause as timed out when its deadline passes. It still waits for the job to confirm that it stopped. You can then resume a run that has restorable state. A run that made progress without restorable state fails instead of restarting from the base model.
Resume a run
Resume a paused run and follow it:
akka-optimize training runs resume RUN_ID
akka-optimize training runs wait RUN_ID --exit-status
Resume keeps the run’s identity and restores its selected state in a new attempt. The new attempt requires the same number of devices as the original allocation. To abort the run instead, use training runs cancel.
Retention
For each run, the catalog retains the newest unreferenced full-state snapshot and the number of unreferenced adapter checkpoints set by execution.keepAdapters. A branch or registered candidate retains the artifact kinds that it needs. The catalog does not delete those snapshots while the branch or model exists. It can delete one artifact kind from a snapshot while retaining another.
Support by method
The following table lists what each method supports:
| Method | Support |
|---|---|
Supervised |
Periodic snapshots, |
Reinforcement |
Periodic snapshots, |
For grpo and gspo with a positive klBeta, the trainer rejects a full-state branch if the run started from a warm-start adapter. The restore does not reconstruct the frozen reference adapter. The trainer rejects an unsupported combination and never downgrades it.
A run that starts from another run’s adapter, through a weights branch or a warm start, can set a positive klBeta only on one GPU. On more than one GPU the reference policy is the base model, so a penalty would pull the run toward the base model rather than toward the adapter it started from, and the trainer refuses it.