0.7.0

This release adds multi-node training, where a deployment offers it. The runs, evaluations and base model selections you submit now show more detail while they start. The CLI uses the same progress format for each. More runs and evaluations now complete, and a run or an evaluation that cannot continue reports the reason. You can install the CLI from the downloads for each release.

Headlines

Multi-node runs

A run can now train across more than one node. Where a deployment offers it, set execution.gpus to a supported GPU count. training base-models get lists the GPU counts you can use for each model. Both multi-node and single-node runs produce an adapter, checkpoints and snapshots.

Runs on many GPUs learn as expected

Supervised and reinforcement runs of Gemma 4 models on more than one GPU now learn as they should. Previously, some GPUs in such a run could compute with wrong token positions, so the run reported a higher loss and improved less per step. The effect grew with the number of GPUs. If you trained a Gemma 4 model on more than one GPU with an earlier release, train it again to get the full benefit.

See what a run, an evaluation and a selection are doing

Runs. A run is in PREPARING from when it starts until its first training step. Its stage shows its current activity, such as reading the dataset. If a run prepares rows before training, its progress counts the rows prepared. After its first step, the run is in TRAINING. A run waiting for GPUs has the stage waiting for GPU capacity.

Evaluations. Before its first answer, an evaluation’s stage shows its current activity, such as loading the model. Its progress then counts the examples answered. During grading, the stage shows the number of answers graded, for example grading 40/100.

Base model selections. When the total size is known, a selection’s progress shows the size of the weights fetched, for example 3.2/9.6 GiB (33%).

Following progress. training runs start --wait, training runs wait and the other commands that follow progress show a progress bar when progress is available. They finish with one line that states the outcome, for example Completed 126 steps, loss 0.1042. Each command identifies the item by the short ID shown in listings.

More runs and evaluations finish

Runs now complete in cases that previously failed or stopped. These cases include Gemma 4 12B with multi-turn conversations, Gemma 4 E4B with long sequences, and datasets with more than 2 GiB of text. Evaluations of Gemma 4 models now complete.

A run or an evaluation that cannot continue now fails and reports the reason. An evaluation that runs out of memory reports that reason.

Install the CLI from the release

Each release includes the akka binary for macOS and Linux. Install it with the install script, or download the archive and verify it against the release’s SHA256SUMS file. For instructions, see Akka Optimize CLI. During a long command, the CLI renews your sign-in before it expires.

Other changes

  • A supervised run reuses the rows that an earlier run prepared from the same dataset, base model and loss settings. The run skips that preparation and starts training sooner.

  • If http: graders are unavailable, Optimize refuses a run or an evaluation that uses one when you submit it. Use json-label-match, a model or a bundle instead.

  • When you pause a run and save its progress, the pause completes after the save. This can take longer for a run that saves often.

  • training runs engine-config RUN_ID shows the training settings a run used, including its default values.

  • training runs get shows the batch across all GPUs, for example 16 sequence(s) per update across 8 GPU(s), 1 per GPU per pass.

  • A base model keeps its status when Optimize restarts.

  • Optimize continues to answer requests during an update.

  • Endpoints and judges over a trained model start and answer on every deployment.

  • training runs tensorboard names a run by its short ID.

  • A percentage below ten includes one decimal place, so step 5 of 1000 reads 0.5%.

Breaking changes

  • PREPARING. Every run now enters PREPARING between QUEUED and TRAINING. A script that waits for TRAINING, or treats a phase other than QUEUED and TRAINING as finished, should treat PREPARING as running.

  • Evaluation progress. An evaluation’s progress counts examples. Its unit is examples, and its total is the number of examples. In 0.6, unit was evaluation steps and total was twice the number of examples. The evaluation’s stage now shows grading progress. Existing evaluations keep their previous progress values.

  • Selection progress. When the size is known, a selection’s progress has unit MiB and counts the weights fetched, rather than files.

  • Queued stage. A queued run’s stage reads waiting for GPU capacity, not queued, waiting for GPU capacity.

  • http: graders. If http: graders are unavailable, Optimize refuses a run or an evaluation that uses one with "this deployment’s training jobs cannot reach the service, and an http grader is asked through it: grade with json-label-match, a model or a bundle".

  • Out of memory. An evaluation that runs out of memory fails with a reason that starts "the evaluation job ran out of device memory". A script that matches reason text should match the new sentence.

  • CLI text output. The output from commands that follow progress, training endpoints list, training endpoints get, training runs get and training base-models get has changed. training endpoints list and training endpoints get show the phase in words, such as Awaiting capacity. A script that parses this output should use -o json.