0.5.0

This release adds multi-GPU reinforcement learning under verl, reinforcement learning for Gemma 4, and checkpoint, pause, and branch support under verl. It also adds telemetry for supervised runs and each GPU that a job uses, and failure reasons that distinguish device memory from host memory.

Headlines

Reinforcement learning on multiple GPUs

A grpo or gspo run on a deployment whose reinforcement engine is verl can train on multiple GPUs on one node. training info shows the deployment’s engine in its Engine row. Set execution.gpus to the number of devices. Each device holds a training rank and a rollout engine. During each step, all devices generate their share of the rollouts and then update the policy. Multi-GPU runs support checkpoints, pauses, resumes, and branches.

config.rlHyperparams.promptsPerStep must be divisible by execution.gpus. The trainer checks this when you submit the run. The deployment’s catalog states the qualified GPU counts for each model and method. If the requested count has no matching envelope, the trainer refuses the run and lists the qualified counts, such as at 1 or 4 GPU(s).

rolloutGpus and stalenessThreshold are not supported. The rollout engine runs on the training devices, so leave rolloutGpus out or set it to 0, and count every device in gpus. The trainer refuses either field above zero and says what to set instead.

To see what fits, use the MAX CLAIM column from training capacity. It shows the largest device claim that each pool accepts. training base-models get lists the runs that the deployment admits for a model, including GPU count, memory per GPU, and host memory per node.

akka-optimize training capacity
akka-optimize training base-models get HuggingFaceTB/SmolLM2-135M-Instruct

For more information, see Execution settings.

Checkpoints, pauses, and branches under verl

A reinforcement run under verl publishes periodic snapshots and supports these operations:

  • Pause with --checkpoint latest or --checkpoint save. A save pause saves at the end of the step that the request arrives in, then stops.

  • Resume from the saved checkpoint.

  • Branch from a full-state snapshot, or from a weights snapshot that either engine wrote.

  • Warm-start from a registered adapter with warmStartModelId.

A run that starts from another run’s adapter, through a weights branch or a warm start, must set klBeta to 0 under verl. verl’s reference policy is the base model, so a KL penalty would pull the run toward the base model rather than toward the adapter it started from. The trainer refuses a positive klBeta at submission.

For more information, see Checkpoints and Branches.

Training observability

The trainer and its training jobs publish metrics, spans, and attributes for dashboards and alerts across runs. See Observability for the complete generated reference.

Stalled runs. A training job counts its completed steps. When a job stops responding, its step gauges retain their last values but the completed-step count stops increasing. The Observability page includes an alert query for this condition. A supervised or TRL job that stops completing steps also records every thread’s stack and a device-memory summary beside its metrics.

Supervised runs. A supervised run reports device-memory use and breaks down each step’s peak into weights, adapter, gradients, optimizer state, activations, allocator memory, and driver memory. It reports the duration of each step phase, evaluation, and checkpoint save, along with throughput and model FLOPs utilization (MFU). It counts rows truncated at maxSeqLength, answers truncated at that limit, and dropped rows, and reports the padding share of each batch. runs get prints these counts as Data fit, and the memory split of the step that reserved the most as Memory.

Device memory on TRL. A reinforcement run on TRL publishes the same split of each step’s peak, adding generation, and its reserved memory. The split is available only for a job that uses one device.

Reward graders. A reinforcement run measures how long each reward term takes to score a rollout, how many scores fail and why, and how many tokens a judge uses. When a judge call times out, fails to connect, or returns no verdict, the rollout scores zero for that term and the failure metric states the cause. Previously, the only sign of these failures was a lower reward.

Reward and length distributions. A reinforcement run records the reward of each rollout, in total and for each term, and the length of each completion as histograms. The step gauges give only a mean and a spread. With the histograms, a dashboard can show percentiles, a bimodal reward, and a long tail of completion lengths.

Throughput and efficiency. A reinforcement run, on verl or TRL, publishes the tokens processed by each step, throughput in tokens per second, and model FLOPs utilization (MFU) of the policy update. Previously, these figures were available only in TensorBoard.

Phase timings on TRL. A reinforcement run on TRL publishes akka.optimize.training.rl.step.timing for generation, reward scoring, and the policy update, using the same phase names as verl. Previously, only a run on verl published phase timings.

Additional step metrics. A reinforcement run publishes the shortest and longest completion of each step and the spread of each reward term’s share across the completions. A run on verl also publishes policy loss. A run on TRL publishes the mean length of completions that ended before the maximum, the share of tokens clipped at each policy-ratio bound, and, for a run that declares tools, tool calls per completion and their failure share. Previously, TRL wrote these figures only to metrics.jsonl.

GPU placement. Each process of a training job publishes the pod, node, and GPU that it runs on, so a dashboard can join the cluster’s GPU metrics to a run by the GPU’s UUID. The Observability page includes an example query.

GPU health. Every training job publishes the power, temperature, utilization, memory use, clock frequency, throttle reasons, and memory errors of each GPU that it uses, and counts the XID errors that the driver reports for it. Each series carries the GPU’s UUID. The Observability page describes how to read the throttle reasons.

Capacity over time. The service publishes gauges for the declared, observed, and claimed devices of each GPU pool, and for the claimed devices and quota of each workload. A dashboard can show pool occupancy and queued claims over time. The training capacity command shows a single point in time.

Rollout inspection

A reinforcement run keeps every rollout that it scores, with the prompt, the completion, the score, and the share of each reward term. The runs rollouts command reads them one step at a time. To check that a rising reward comes from better answers and not from reward hacking, read the worst and best rollouts of a step:

akka-optimize training runs rollouts RUN_ID --step 40 --worst 5
akka-optimize training runs rollouts RUN_ID --best 3 -o json

Replace RUN_ID with the ID of the run.

A run that trained on an earlier version has no rollouts.

For more information, see Rollouts.

Reinforcement learning on Gemma 4

An image-text model such as Gemma 4 E4B can take a grpo or gspo run over a text dataset. Previously, such a model supported only sft. As in a supervised run, the adapter updates the language backbone and leaves the vision and audio towers unchanged.

Under TRL, the catalog lists grpo and gspo for such a model. Under verl, it lists them where the deployment enables reinforcement over image-text models. To check which methods a model supports in your deployment, run:

akka-optimize training base-models get google/gemma-4-E4B-it

A reinforcement run cannot train on images. The trainer rejects a grpo or gspo run over a dataset with images, regardless of the base model.

A reinforcement run reports the modules that its adapter covers, as a supervised run does. Another run can use that adapter as its warm start.

For more information, see Reinforcement learning.

Other changes

Failure reasons that say which memory ran out

A run that exhausts memory reports whether device memory or host memory ran out and says what you can change:

  • "the training job ran out of device memory; a shorter sequence, a smaller micro batch or activation recomputation uses less", when the model’s pass does not fit the accelerator.

  • "the training job ran out of host memory; a smaller micro batch or sequence length uses less", when the node’s memory runs out and the job is stopped.

  • "the training job’s rollout engine did not fit its share of device memory; a shorter sequence or a smaller model fits", when a reinforcement run’s rollout engine cannot start in the part of the device it is given.

  • "the training job ran out of device memory handing its weights to the rollout engine; a smaller model or a device with more memory fits", when a reinforcement run under verl holds both copies of the weights on the device while it updates the rollout engine.

  • "the evaluation job ran out of device memory; a lower completion limit uses less", and "the evaluation job ran out of host memory; a device with more host memory fits it", for an evaluation.

When a step of a run fails on every retry, the failure reason names what the run could not do, such as "this run’s job could not be started on the training cluster". When the training cluster does not answer while a run waits for its devices, the reason is "the training cluster could not be reached while this run waited for its devices". Previously, such a run read "step failed after retries" and the phase it was in.

For more information, see Why a run failed.

Larger dataset uploads

A dataset upload streams to storage as it is read, so it can use the deployment’s full upload limit, 256 MiB by default. Previously, an upload over 8 MiB was reset without a reason. A deployment sets the limit with trainer.datasets.max-bytes.

Grafana dashboard for a run

The runs grafana command prints a link to the run on the project’s training dashboard in Grafana. The link covers the run from its start to its end, and it carries a token that lets anyone who has it view the project’s Grafana until the expiry that the command prints.

akka-optimize training runs grafana RUN_ID

For more information, see Grafana.

Resolved configuration on a dry run

The --show-resolved flag on runs start --dry-run and pipelines start --dry-run prints the configuration and execution settings that a submission trains under. The output fills in every default and resolves every dataset reference, so you can check a plan before it uses any devices. The output matches what runs get prints for a started run. With -o json, the output includes the resolved settings as structured data.

To print the resolved settings for a run or a pipeline, run:

akka-optimize training runs start -f run.json --dry-run --show-resolved
akka-optimize training pipelines start -f pipeline.json --dry-run --show-resolved

For a pipeline, the output covers each train stage whose dataset is registered. A stage that trains on the output of an earlier stage has no registered dataset to resolve, so the output omits it.

For more information, see Start a run and Pipelines.

Log in through a named context

akka can use the login of an akka context other than the current one. Name the context with the global --context flag or the AKKA_CONTEXT environment variable. akka exchanges the token at the API host that the context specifies, so one shell can reach a service on another platform. A context name that the configuration file does not contain is refused, and the error lists the contexts that the file contains. The credential environment variables, such as AKKA_ACCESS_TOKEN, still take precedence over any login.

For more information, see Authenticate.

Smaller changes

  • Fixed an issue that caused endpoints create to fail immediately after endpoints delete for the same base model. The deleted endpoint stayed listed for a moment, and the command tried to add the model to it. In this release, the command starts a new endpoint.

  • Fixed an issue that left the peak device memory of a reinforcement run on verl empty. akka.optimize.training.rl.step.memory.allocated, akka.optimize.training.rl.step.memory.reserved, and the peak memory figure of the run report values.

  • Fixed an issue that caused runs metrics --fields maxMemoryGb and the other figures that every engine reports, such as kl, to show no value. With -o json, the output now includes them.

  • Fixed several issues that stopped a reinforcement run on verl from completing on a Ray cluster. A run on verl that writes no step for 30 minutes now stops and fails, instead of waiting indefinitely. A deployment sets the limit with trainer.ray.jobs.verl-stall-limit.

  • Fixed an issue that left akka.optimize.training.run.engine off the spans and series of a reinforcement run, so a dashboard could not tell verl from TRL.

  • runs list and pipelines list show a CREATED column.

  • runs get summarizes the adapter of a run over a multimodal model, for example each layer under model.text_model.layers (down_proj, gate_proj, k_proj, o_proj, q_proj, up_proj, v_proj). Previously, it printed the regular expression that the adapter records. With -o json, targetModules still contains the regular expression.

  • Fixed an issue that caused a reinforcement run on TRL to ignore gradientCheckpointing. The run kept every activation, even when the run plan stated gradient checkpointing or the trainer derived it from maxSeqLength.

  • Fixed an issue that failed every supervised run over a Ministral model. The run now trains on the prompt as the model is served it.

  • Fixed an issue that made a resumed run’s observed step go backwards before the new attempt reported. The step now stays at the checkpoint the run resumed from until the new attempt reports a step.

  • A run that waits for its devices while the training cluster restarts keeps waiting, rather than failing, and a run that fails while it waits releases its devices at once.

  • training capacity no longer counts a Ray head that holds no GPUs as a node, so a cluster scaled to zero reads 0 on 0 nodes.

  • A catalog change reaches the service with the deploy, without a separate restart.

Breaking changes

  • base-models get text output. The runs a model admits are a table under "Runs this deployment admits", not a line for each method. A script that reads them should use -o json, which is unchanged.

  • Failure reason wording. A job stopped for host memory reads "ran out of host memory" where it read "ran out of memory". A script that matches reason text should match the new sentences above.