Observe a run

Run details, metrics, and TensorBoard show the progress and results of a training run. Use them to monitor a run while it trains or investigate it after it finishes.

Run detail

Read a run, or follow it to a terminal phase:

akka-optimize training runs get RUN_ID
akka-optimize training runs wait RUN_ID --exit-status

The run detail reports the phase, training progress, cost, timing, device memory, and any failure reason. When the job ends, it reports the GPU seconds that the run consumed. Until then, the detail marks an estimate based on the training time and the devices requested by the run. The duration separates the time spent waiting for the cluster from the time spent training. The JSON output reports timing for each phase under timing.phaseMillis. A pipeline’s detail reports the cost of each train stage and the total. observedStep counts the updates reported by the worker. savedStep is the selected restorable checkpoint. A branch counts its own updates from zero.

After submission, a job waits for the cluster to schedule it before training. The run detail remains in the QUEUED phase until the job starts. The wait depends on the cluster’s other work. A supervised run then reads PREPARING while it reads and tokenises its dataset and starts its training workers, with rows done of rows total while it tokenises. On a large dataset of long conversations this takes minutes, and the run holds its devices throughout. runs metrics provides the live view of a training run. The run record updates less often, so its progress can lag behind the metrics.

wait streams status updates until the run reaches a terminal phase. With -o json, it streams line-delimited events. A terminal event with the failure reason ends the stream, followed by the full run record. --timeout limits the wait. The run continues after the timeout expires.

Run costs include all attempts and any work repeated after a restore.

Peak memory is the peak device memory for each device. For a supervised run and a reinforcement run on TRL that uses one device, Memory shows what held the device at the peak of the step that reserved the most: the weights, adapter, gradients or generation, activations, optimizer state, what the allocator and the driver held, and what was free. It also shows what the other phase of the step held. For a supervised run, Data fit reports how many rows reached maxSeqLength, how many of them likely lost the end of their target, how many were dropped with nothing left to train on, and the share of the batches that was padding.

Why a run failed

A failed run reports one sentence. That sentence is the whole of what the run says about the failure, and it is one of these:

Reason What to do

A sentence the training runtime refused with, such as "an example ends with a user message"

What the sentence says. The runtime read the run’s inputs and would not train on them.

"the training job ran out of device memory; a shorter sequence, a smaller micro batch or activation recomputation uses less"

A training step exceeded the device’s memory. Lower config.maxSeqLength, lower execution.microBatchSequences on a reinforcement run, or enable config.model.gradientCheckpointing. If the run still fails, select a device profile with more device memory.

"the training job’s rollout engine did not fit its share of device memory; a shorter sequence or a smaller model fits"

The rollout engine could not hold the model weights and the cache for one sequence in its share of device memory. Lower config.maxSeqLength or use a smaller model.

"the training job ran out of device memory handing its weights to the rollout engine; a smaller model or a device with more memory fits"

While the trainer updates the rollout engine, the device holds both copies of the model weights. Use a smaller model, or select a device profile with more device memory.

"the evaluation job ran out of device memory; a lower completion limit uses less"

Lower the evaluation’s completion limit, or score on a device profile with more device memory.

"the evaluation job’s inference engine did not fit its share of device memory; a smaller model or a device with more memory fits"

The evaluation’s inference engine could not hold the model weights and its cache in its share of device memory. The completion limit does not change this. Evaluate a smaller model, or score on a device profile with more device memory.

"the training job ran out of host memory; a smaller micro batch or sequence length uses less"

The worker ran out of host memory, so the operating system stopped the job. Lower config.maxSeqLength and, on a reinforcement run, execution.microBatchSequences. If the run still fails, select a device profile with more host memory.

"the evaluation job ran out of host memory; a device with more host memory fits it"

The evaluation worker ran out of host memory, so the operating system stopped the job. Score on a device profile with more host memory.

"the training cluster could not be reached while this run waited for its devices"

The trainer granted the run its devices, and the training cluster did not answer before the wait for devices ran out. Submit the run again later. If it fails the same way, quote the run ID to your operator.

A sentence that names what the run could not do, such as "this run’s job could not be started on the training cluster"

A step of the run failed on every retry. The sentence names no cause. Submit the run again later. If it fails the same way, quote the run ID to your operator.

"the training job was stopped"

The cluster terminated the job. Start the run again, or resume it if it holds a checkpoint.

"the training job stopped before it finished"

The job ended without reporting what it did. Start the run again.

"the training job failed; the run ID is what support needs"

Nothing in the job reported what went wrong. Quote the run ID.

The reason carries nothing from inside the job: no exception line, no compute backend message, and no module name. You can’t change any of those from here.

Metrics

Read the metrics of a run:

akka-optimize training runs metrics RUN_ID
akka-optimize training runs metrics RUN_ID -o json

The command reports the metrics that the engine logged for each step. It retains every numeric field from the engine under measured. A reinforcement run also reports the contribution of each reward term under rewardTerms. The trainer removes metrics after a restored step, so the series reflects the run’s current progress.

Grafana

Print a link to the run on the project’s training dashboard in Grafana:

akka-optimize training runs grafana RUN_ID

The link covers the run from its start to its end, or to now while it trains. It carries a token that lets anyone who has it view the project’s Grafana until the expiry that the command prints, so do not share it. The command reads the project from --project, the AKKA_PROJECT environment variable, or the akka context, and prints the project that it used. --attempt selects a resubmitted attempt of the run.

Hyperparameters

Read the hyperparameters of a run:

akka-optimize training runs hparams RUN_ID

The command reports the resolved hyperparameters used for training, including every applied default. They become available shortly after the job starts, when the worker writes them.

Engine configuration

Read the configuration the training engine was given:

akka-optimize training runs engine-config RUN_ID
akka-optimize training runs engine-config RUN_ID -o json

The worker writes this before the first step, in two parts. requested is what the worker built from the run’s settings for TRL or verl. resolved is the engine’s whole configuration once the engine has filled in its own defaults, and the adapter as it was attached to the model. The table lists each requested setting and shows a second value where the engine resolved it differently. Use the JSON output to read a default the run did not set, for example --jq .resolved.config.adam_beta1.

TensorBoard

The training engines write TensorBoard event files. During training, the worker uploads these files with reward-term curves and the hyperparameters for the HParams dashboard.

To start TensorBoard for a run, run:

akka-optimize training runs tensorboard RUN_ID

The command uses your credentials to download the files and starts TensorBoard on your machine. TensorBoard has no authentication, so keep it local. Engine curves use the engine’s step counter. A full-state branch can start above zero on these curves, while the run detail counts from zero.

Rollouts

A reinforcement run keeps every rollout that it scores: the prompt, the completion, the score, and the weighted share of each reward term. To check that a rising reward comes from better answers and not from reward hacking, read the rollouts:

akka-optimize training runs rollouts RUN_ID
akka-optimize training runs rollouts RUN_ID --step 40 --worst 5
akka-optimize training runs rollouts RUN_ID --best 3 -o json

The command shows one step. By default, it shows the latest step that the run kept. To show a different step, use --step. --worst N and --best N sort the rollouts of the step by score and show the first N. --limit N shows the first N in scoring order. The table shows the last message of the prompt and truncates the prompt and the completion. -o json and --jq return both in full.

The worker keeps the rollouts of a step after the engine reports that step. If a job stops or fails, it keeps the rollouts that it scored in its last step.

Telemetry

The trainer and its training jobs publish metrics and spans through OpenTelemetry under the akka.optimize.training. prefix. For what they publish and how to use it across runs, see Observability.