Telemetry reference

The names, units, and attributes below are generated from the telemetry registry that the trainer and the training job use to record telemetry. See Observability for how to use them.

Metrics

In the Emitter column, service marks a metric that the trainer service publishes, and job marks a metric that the training job publishes.

GPU pools

Metric Kind Unit Emitter Attributes Description

akka.optimize.training.pool.claims

gauge

{claim}

service

akka.optimize.training.claim.state, akka.optimize.training.instance, akka.optimize.training.pool.accelerator, akka.optimize.training.pool.serving

Number of active claims on a GPU pool, partitioned by claim state.

akka.optimize.training.pool.devices.claimed

gauge

{device}

service

akka.optimize.training.claim.state, akka.optimize.training.instance, akka.optimize.training.pool.accelerator, akka.optimize.training.pool.serving

Number of devices the active claims on a GPU pool hold or are waiting for, partitioned by claim state.

akka.optimize.training.pool.devices.declared

gauge

{device}

service

akka.optimize.training.instance, akka.optimize.training.pool.accelerator, akka.optimize.training.pool.serving

Number of devices the deployment declares a GPU pool holds, which is what work is admitted against.

akka.optimize.training.pool.devices.observed

gauge

{device}

service

akka.optimize.training.instance, akka.optimize.training.pool.accelerator, akka.optimize.training.pool.serving

Number of devices the pool’s cluster reported holding at its last reading. Absent until the cluster has been read.

akka.optimize.training.workload.devices.claimed

gauge

{device}

service

akka.optimize.training.claim.state, akka.optimize.training.instance, akka.optimize.training.workload

Number of devices a workload’s active claims hold or are waiting for across every pool, partitioned by claim state.

akka.optimize.training.workload.devices.quota

gauge

{device}

service

akka.optimize.training.instance, akka.optimize.training.workload

Number of devices a workload may hold at once across every pool.

Training checkpoints

Metric Kind Unit Emitter Attributes Description

akka.optimize.training.checkpoint.operation.duration

histogram

s

job

akka.optimize.training.checkpoint.artifact.kind (only when the operation concerns one artifact kind), akka.optimize.training.checkpoint.failure.reason (only when the outcome is failed), akka.optimize.training.checkpoint.operation, akka.optimize.training.checkpoint.outcome, akka.optimize.training.checkpoint.restore.mode (only on restore observations), akka.optimize.training.checkpoint.storage.mode

Duration of a checkpoint staging, publication, upload, catalog, materialization, or validation attempt.

akka.optimize.training.checkpoint.publications.pending

gauge

{snapshot}

job

Number of staged snapshots still awaiting durable publication by this training job, including an in-flight publication.

akka.optimize.training.checkpoint.transfer

histogram

By

job

akka.optimize.training.checkpoint.artifact.kind (only when the operation concerns one artifact kind), akka.optimize.training.checkpoint.operation, akka.optimize.training.checkpoint.outcome, akka.optimize.training.checkpoint.restore.mode (only on restore observations), akka.optimize.training.checkpoint.storage.mode

Bytes actually uploaded or downloaded during a checkpoint transfer attempt; omitted when no transferred amount is known.

Training cost

Metric Kind Unit Emitter Attributes Description

akka.optimize.training.gpu.time

counter

s

service

akka.optimize.training.gpu.derived, akka.optimize.training.profile, akka.optimize.training.workload

Total GPU compute time consumed by training jobs.

Training devices

Metric Kind Unit Emitter Attributes Description

akka.optimize.training.gpu.clock

gauge

MHz

job

akka.optimize.training.gpu.clock.domain, hw.id (only when the worker process uses an NVIDIA GPU), hw.name, hw.vendor

Frequency of a GPU clock, averaged over the samples taken since the previous export.

akka.optimize.training.gpu.clock.throttle_reasons

gauge

1

job

hw.id (only when the worker process uses an NVIDIA GPU), hw.name, hw.vendor

Every reason NVML gave for holding a GPU’s clocks down in any sample since the previous export, as the bitwise OR of NVML’s clock event reason bits. Zero is a GPU held down by nothing, and 0x1 alone is an idle one. Bit 0x2 is an application clock setting, 0x4 software power cap, 0x8 hardware slowdown, 0x20 software thermal slowdown, 0x40 hardware thermal slowdown, 0x80 hardware power brake. In PromQL, floor(x / 8) % 2 reads the bit 0x8.

akka.optimize.training.gpu.memory.activity

gauge

1

job

hw.id (only when the worker process uses an NVIDIA GPU), hw.name, hw.vendor

Fraction of time the GPU’s memory was being read or written, averaged over the samples taken since the previous export. NVML reports it as memory utilization. It is not how full the memory is, which is hw.gpu.memory.utilization.

akka.optimize.training.gpu.memory.pages.retired

gauge

{page}

job

error.type, hw.id (only when the worker process uses an NVIDIA GPU), hw.name, hw.vendor

Memory pages a GPU has retired over its lifetime, by the kind of error that caused it. Reported by GPUs before Ampere.

akka.optimize.training.gpu.memory.repair.pending

gauge

1

job

hw.id (only when the worker process uses an NVIDIA GPU), hw.name, hw.vendor

1 when a GPU has a row remapping or a page retirement that takes effect only when the GPU is reset, and 0 otherwise. A GPU with a pending repair should be reset before it takes another job.

akka.optimize.training.gpu.memory.rows.remapped

gauge

{row}

job

error.type, hw.id (only when the worker process uses an NVIDIA GPU), hw.name, hw.vendor

Memory rows a GPU has remapped to spare rows over its lifetime, by the kind of error that caused it. Reported by Ampere and later GPUs.

akka.optimize.training.gpu.xid.errors

counter

{error}

job

hw.id (only when the worker process uses an NVIDIA GPU), hw.name, hw.vendor

XIDs, the driver’s codes for a GPU fault, reported for a GPU the job holds since the job started watching it. Published from zero, so an alert on its increase sees the first one. Absent on a GPU whose events NVML does not deliver. Read beside akka.optimize.training.gpu.xid.last, which is the code of the most recent XID. NVIDIA’s XID catalog says what each code means. Some codes are faults in the job’s own kernels, such as an illegal memory access, and others are hardware.

akka.optimize.training.gpu.xid.last

gauge

1

job

hw.id (only when the worker process uses an NVIDIA GPU), hw.name, hw.vendor

The code of the most recent XID reported for a GPU the job holds. Absent until the GPU reports one.

Training jobs

Metric Kind Unit Emitter Attributes Description

akka.optimize.training.devices.held

gauge

{device}

job

Number of accelerator devices currently allocated to this training job, reinforcement or supervised.

akka.optimize.training.job.placement

gauge

{worker}

job

akka.optimize.training.job.rank, akka.optimize.training.job.role, hw.id (only when the worker process uses an NVIDIA GPU), k8s.node.name (only when the job runs on Kubernetes), k8s.pod.name (only when the job runs on Kubernetes)

Always 1, one series for each worker process of verl while it runs, and one for each device a TRL job holds while the job runs. It names the rank, role, pod, node and GPU of the process, and the run’s identity is on its resource, so a node or GPU series joins to the run through it. A worker that exits ends its series.

akka.optimize.training.rl.grader.duration

histogram

s

job

akka.optimize.training.rl.term

Duration of one reward term scoring one rollout, including a failed call. Recorded when the step the rollout belongs to is reported.

akka.optimize.training.rl.grader.failures

counter

{call}

job

akka.optimize.training.rl.grader.failure.cause, akka.optimize.training.rl.term

Number of reward term calls that produced no score, partitioned by cause. A refused judge call and a judge that fails too many calls in a row stop the job instead, and are not counted here.

akka.optimize.training.rl.grader.tokens

counter

{token}

job

akka.optimize.training.rl.grader.token.type, akka.optimize.training.rl.term

Number of tokens a judge term reported for its calls, from the usage block of the judge’s reply. A reply without one adds nothing.

akka.optimize.training.rl.idle.ratio

gauge

1

job

akka.optimize.training.rl.role

Fraction of step duration that the role spent idle waiting for synchronization or data transfer.

akka.optimize.training.rl.rollout.completion.length

histogram

{token}

job

Length in tokens of one generated completion, counted to the last token the model generated. Recorded when the step the rollout belongs to is reported. Recorded on an exponential histogram, so runs with different maximum completion lengths merge into one percentile rather than carrying their own bucket boundaries. akka.optimize.training.rl.step.completion.clipped_ratio reports the truncated share; the histogram alone does not single out the top bucket.

akka.optimize.training.rl.rollout.reward

histogram

1

job

akka.optimize.training.rl.term

Reward of one rollout, once as the weighted sum under the term total and once as each declared term’s weighted share. Recorded when the step the rollout belongs to is reported. The bucket boundaries are fixed. A score is usually in [0, 1], but weights are not normalized and a rule can return any number. The bucket at or below 0 counts the rollouts that scored nothing, and the bucket above 0.99 up to 1 counts those with a full score. The buckets above 1 hold the rest.

akka.optimize.training.rl.step.clip_ratio

gauge

1

job

Proportion of tokens whose policy probability ratio was clipped in the most recent optimizer step.

akka.optimize.training.rl.step.clip_ratio.high

gauge

1

job

Proportion of tokens whose policy probability ratio was clipped at the upper bound in the most recent optimizer step, where a positive advantage would otherwise have pushed the ratio up further.

akka.optimize.training.rl.step.clip_ratio.low

gauge

1

job

Proportion of tokens whose policy probability ratio was clipped at the lower bound in the most recent optimizer step, where a negative advantage would otherwise have pushed the ratio down further.

akka.optimize.training.rl.step.completion.clipped_ratio

gauge

1

job

Proportion of completions generated in the most recent training step that were truncated at the maximum completion length.

akka.optimize.training.rl.step.completion.length

gauge

{token}

job

Mean length in tokens of the completions generated in the most recent training step.

akka.optimize.training.rl.step.completion.length.max

gauge

{token}

job

Longest completion generated in the most recent training step, in tokens.

akka.optimize.training.rl.step.completion.length.min

gauge

{token}

job

Shortest completion generated in the most recent training step, in tokens.

akka.optimize.training.rl.step.completion.length.terminated

gauge

{token}

job

Mean length in tokens of the completions in the most recent training step that ended with an end-of-sequence token rather than being truncated at the maximum completion length. Not recorded for a step in which every completion was truncated.

akka.optimize.training.rl.step.count

counter

{step}

job

Number of steps of the reinforcement learning loop that this job attempt has completed. A step number reported twice in a row counts once. The count starts at zero in each attempt, so a resumed run starts a new series. A series that stops increasing for longer than a checkpoint or a validation pass takes is a stalled loop.

akka.optimize.training.rl.step.duration

histogram

s

job

Duration of a single step in the reinforcement learning training loop.

akka.optimize.training.rl.step.entropy

gauge

1

job

Mean per-token entropy of the completions generated in the most recent training step.

akka.optimize.training.rl.step.flat_groups

gauge

1

job

Proportion of prompt groups evaluated in the latest step that exhibited zero reward variance across completions.

akka.optimize.training.rl.step.grad_norm

gauge

1

job

Gradient norm before clipping in the most recent optimizer step of the reinforcement learning loop.

akka.optimize.training.rl.step.kl

gauge

1

job

KL divergence between the policy and the reference model in the most recent training step. Absent when the run configures no KL penalty.

akka.optimize.training.rl.step.learning_rate

gauge

1

job

Learning rate applied in the most recent optimizer step of the reinforcement learning loop.

akka.optimize.training.rl.step.memory.allocated

gauge

GiBy

job

The most accelerator device memory the training process has held in tensors at once since it started, read at the most recent step. A high-water mark, the same on both engines. Under verl, the most of any of the actor’s ranks.

akka.optimize.training.rl.step.memory.reserved

gauge

GiBy

job

The most accelerator device memory the caching allocator has held at once since the training process started, read at the most recent step. The figure an out-of-memory error is raised against. The difference from the allocated figure is memory kept for reuse or lost to fragmentation. Under verl, the most of any of the actor’s ranks.

akka.optimize.training.rl.step.mfu

gauge

1

job

Model FLOPs utilization of the policy update in the most recent training step: the estimated FLOPs per second the update achieved, as a fraction of the dense bf16 peak of the devices that ran it. The update is the step less generation and scoring. Both engines count 6 FLOPs per parameter per token, a full forward and backward pass, although a LoRA update takes no gradient for the frozen weights and does nearer 4. verl estimates from the model configuration, adds attention over the sequence and the embedding, and reports zero for an architecture it has no estimate for. TRL counts the linear layers of the language model and its output head, and reports nothing on a device with no known peak, such as a CPU.

akka.optimize.training.rl.step.policy_loss

gauge

1

job

Policy gradient loss of the most recent optimizer step of the reinforcement learning loop. Recorded by verl. TRL reports it only with an entropy bonus, which the trainer does not enable, so a TRL run records none.

akka.optimize.training.rl.step.reward

gauge

1

job

Mean reward score evaluated in the most recent reinforcement learning training step.

akka.optimize.training.rl.step.reward.std

gauge

1

job

Standard deviation of the reward across the completions scored in the most recent training step. Zero indicates every completion received the same reward.

akka.optimize.training.rl.step.reward.term

gauge

1

job

akka.optimize.training.rl.term

Mean weighted contribution of one declared reward term across the completions scored in the most recent training step. The contributions of all terms sum to the term named total.

akka.optimize.training.rl.step.reward.term.std

gauge

1

job

akka.optimize.training.rl.term

Standard deviation of one declared reward term’s weighted share across the completions scored in the most recent training step, including the term named total for the weighted sum.

akka.optimize.training.rl.step.reward.zero_fraction

gauge

1

job

Proportion of completions scored in the most recent training step that received a reward of zero.

akka.optimize.training.rl.step.rollouts

gauge

{rollout}

job

Number of completions scored in the most recent training step.

akka.optimize.training.rl.step.throughput

gauge

{token}/s

job

Prompt and completion tokens processed per second per accelerator device in the most recent training step, over the whole step, generation included. A TRL run trains on one device.

akka.optimize.training.rl.step.timing

gauge

s

job

akka.optimize.training.rl.step.phase

Wall-clock duration of one phase of the most recent training step, as measured by the engine. Reported by verl for every phase it has, and by TRL for generation, reward computation and actor update.

akka.optimize.training.rl.step.tokens

gauge

{token}

job

Number of prompt and completion tokens in the batch of the most recent training step. TRL counts the tokens it generated for the step, which is the batch it trains on, and leaves out tool responses inside a completion.

akka.optimize.training.rl.step.tool.call_frequency

gauge

{call}

job

Mean number of tool calls per completion in the most recent training step. Recorded only where the run declares tools.

akka.optimize.training.rl.step.tool.failure_frequency

gauge

1

job

Proportion of tool calls in the most recent training step that failed. Recorded only where the run declares tools.

akka.optimize.training.sft.data.examples

gauge

{example}

job

akka.optimize.training.sft.data.fit

Training rows counted by how they fit maxSeqLength, set once when the job has prepared its dataset. Absent on a packed run and on a run over pages, whose rows TRL does not prepare as token ids.

akka.optimize.training.sft.eval.loss

gauge

1

job

Loss over the held-out evaluation split at the most recent evaluation of the supervised fine-tuning loop.

akka.optimize.training.sft.step.count

counter

{step}

job

Number of steps of the supervised fine-tuning loop that this job attempt has completed. A step number reported twice in a row counts once. The count starts at zero in each attempt, so a resumed run starts a new series. A series that stops increasing for longer than a checkpoint or a validation pass takes is a stalled loop.

akka.optimize.training.sft.step.duration

histogram

s

job

Duration of a single step in the supervised fine-tuning loop.

akka.optimize.training.sft.step.grad_norm

gauge

1

job

Gradient norm before clipping in the most recent optimizer step of the supervised fine-tuning loop.

akka.optimize.training.sft.step.learning_rate

gauge

1

job

Learning rate applied in the most recent optimizer step of the supervised fine-tuning loop.

akka.optimize.training.sft.step.loss

gauge

1

job

Training loss at the most recent step of the supervised fine-tuning loop.

akka.optimize.training.sft.step.memory.allocated

gauge

GiBy

job

The most accelerator device memory the training process has held in tensors at once since it started, read at the most recent supervised fine-tuning step. With more than one device, the figure is for the device that held the most. Means the same as the reinforcement figure of the same name.

akka.optimize.training.sft.step.memory.reserved

gauge

GiBy

job

The most accelerator device memory the caching allocator has held at once since the training process started, read at the most recent supervised fine-tuning step. With more than one device, the figure is for the device that held the most. Means the same as the reinforcement figure of the same name.

akka.optimize.training.sft.step.mfu

gauge

1

job

Model FLOPs utilization of the most recent supervised fine-tuning step: the estimated FLOPs per second the step achieved, as a fraction of the dense bf16 peak of the devices that ran it. Counted as the reinforcement figure on TRL is: 6 FLOPs per parameter of the language model’s linear layers and output head per token, although a LoRA update takes no gradient for the frozen weights and does nearer 4. Absent on a device with no known peak, such as a CPU.

akka.optimize.training.sft.step.padding

gauge

1

job

Share of the token positions in the batches of the most recent supervised fine-tuning step that were padding. Absent on a packed run, which is collated without padding.

akka.optimize.training.sft.step.throughput

gauge

{token}/s

job

Tokens trained on per second per accelerator device in the most recent supervised fine-tuning step, padding left out. Read beside the padding share of a batch: padded tokens cost time and are not counted.

akka.optimize.training.sft.step.timing

gauge

s

job

akka.optimize.training.sft.step.phase

Wall-clock duration of one phase of the supervised loop, the most recent time it ran: the input, forward and backward passes, optimizer step and whole step of the most recent update, and the most recent evaluation and checkpoint save.

akka.optimize.training.sft.step.token_accuracy

gauge

1

job

Mean next-token prediction accuracy over the batch of the most recent supervised fine-tuning step.

akka.optimize.training.sft.step.tokens

gauge

{token}

job

Number of tokens in the batch of the most recent supervised fine-tuning step, padding left out.

akka.optimize.training.step.memory.usage

gauge

GiBy

job

akka.optimize.training.memory.component

What held the memory of the device at the peak of the most recent training step, one series per part. The phase that did not hold the peak, update or generation, reports zero, so the parts add up to the device’s memory. Under verl, each worker measures its own device, and the series are those of the rank that reserved the most. A TRL job reports them only when it sees one device. Weights, adapter, gradients and optimizer state are counted from the tensors that hold them. Activations and generation are what the step’s peak held above the memory in use when the step or its generation began. The allocator’s share is the step’s peak reserved memory less its peak allocated memory. What is outside the allocator is read from the driver at the end of the step.

Training pipelines

Metric Kind Unit Emitter Attributes Description

akka.optimize.training.pipeline.active

gauge

{pipeline}

service

akka.optimize.training.instance

Number of training pipelines currently active.

akka.optimize.training.pipeline.stage.duration

histogram

s

service

akka.optimize.training.pipeline.stage.outcome, akka.optimize.training.pipeline.stage.type

Execution duration of an individual training pipeline stage.

Training runs

Metric Kind Unit Emitter Attributes Description

akka.optimize.training.run.active

gauge

{run}

service

akka.optimize.training.instance, akka.optimize.training.run.phase

Number of active training runs, partitioned by phase.

akka.optimize.training.run.duration

histogram

s

service

akka.optimize.training.run.outcome

Total end-to-end execution duration of a training run.

akka.optimize.training.run.failed

counter

{run}

service

akka.optimize.training.run.phase (only if the failing phase is known)

Total number of failed training runs, partitioned by the phase in which failure occurred.

akka.optimize.training.run.phase.duration

histogram

s

service

akka.optimize.training.run.phase

Duration spent in a specific training run phase before transitioning.

akka.optimize.training.run.phase.entered

counter

{entry}

service

akka.optimize.training.run.phase

Total number of times training runs have entered each lifecycle phase.

Training service operations

Metric Kind Unit Emitter Attributes Description

akka.optimize.training.service.operation.attempts

counter

{attempt}

service

akka.optimize.training.pause.policy (only on pause observations), akka.optimize.training.service.failure.reason (only when the outcome is not succeeded), akka.optimize.training.service.operation, akka.optimize.training.service.outcome, akka.optimize.training.service.stage

Number of accepted requests and actual external pause, catalog, and branch preparation attempts observed by the service.

akka.optimize.training.service.operation.duration

histogram

s

service

akka.optimize.training.pause.policy (only on pause observations), akka.optimize.training.service.failure.reason (only when the outcome is not succeeded), akka.optimize.training.service.operation, akka.optimize.training.service.outcome, akka.optimize.training.service.stage

Elapsed time from durable acceptance to a pause acknowledgement, resource release, or terminal preparation result.

Spans

Span Kind Attributes Description

restore

internal

akka.optimize.training.checkpoint.artifact.kind (only when the operation concerns one artifact kind), akka.optimize.training.checkpoint.restore.mode (only on restore observations), akka.optimize.training.checkpoint.snapshot.id (only after a snapshot has been identified), akka.optimize.training.checkpoint.source.attempt (only on restore spans), akka.optimize.training.checkpoint.source.engine_step (only on restore spans), akka.optimize.training.checkpoint.source.run_id (only on restore spans), akka.optimize.training.checkpoint.storage.mode, akka.optimize.training.job.submission_id, akka.optimize.training.pipeline.id (only if a pipeline stage started the run), akka.optimize.training.pipeline.stage.name (only if a pipeline stage started the run), akka.optimize.training.run.attempt, akka.optimize.training.run.base_model, akka.optimize.training.run.engine, akka.optimize.training.run.group_size (only on a reinforcement run), akka.optimize.training.run.id, akka.optimize.training.run.learning_rate, akka.optimize.training.run.lora_rank, akka.optimize.training.run.method, akka.optimize.training.run.settings_sha, akka.optimize.training.workload

Restores exact training state for a later attempt or adapter weights for a branch origin.

ship-checkpoint

internal

akka.optimize.training.checkpoint.snapshot.id (only after a snapshot has been identified), akka.optimize.training.checkpoint.storage.mode, akka.optimize.training.job.submission_id, akka.optimize.training.pipeline.id (only if a pipeline stage started the run), akka.optimize.training.pipeline.stage.name (only if a pipeline stage started the run), akka.optimize.training.run.attempt, akka.optimize.training.run.base_model, akka.optimize.training.run.engine, akka.optimize.training.run.group_size (only on a reinforcement run), akka.optimize.training.run.id, akka.optimize.training.run.learning_rate, akka.optimize.training.run.lora_rank, akka.optimize.training.run.method, akka.optimize.training.run.settings_sha, akka.optimize.training.workload

Publishes one staged checkpoint and waits for durable catalog acknowledgement when configured.

stage-inputs

internal

akka.optimize.training.job.submission_id, akka.optimize.training.pipeline.id (only if a pipeline stage started the run), akka.optimize.training.pipeline.stage.name (only if a pipeline stage started the run), akka.optimize.training.run.attempt, akka.optimize.training.run.base_model, akka.optimize.training.run.engine, akka.optimize.training.run.group_size (only on a reinforcement run), akka.optimize.training.run.id, akka.optimize.training.run.learning_rate, akka.optimize.training.run.lora_rank, akka.optimize.training.run.method, akka.optimize.training.run.settings_sha, akka.optimize.training.workload

Fetches and stages the training dataset and base model weights to local storage.

train

internal

akka.optimize.training.job.submission_id, akka.optimize.training.pipeline.id (only if a pipeline stage started the run), akka.optimize.training.pipeline.stage.name (only if a pipeline stage started the run), akka.optimize.training.run.attempt, akka.optimize.training.run.base_model, akka.optimize.training.run.engine, akka.optimize.training.run.group_size (only on a reinforcement run), akka.optimize.training.run.id, akka.optimize.training.run.learning_rate, akka.optimize.training.run.lora_rank, akka.optimize.training.run.method, akka.optimize.training.run.settings_sha, akka.optimize.training.workload

Executes the training loop. A reinforcement learning run performs sample generation, reward evaluation, and optimizer steps; a supervised fine-tuning run performs optimizer steps over the dataset.

upload-adapter

internal

akka.optimize.training.job.submission_id, akka.optimize.training.pipeline.id (only if a pipeline stage started the run), akka.optimize.training.pipeline.stage.name (only if a pipeline stage started the run), akka.optimize.training.run.attempt, akka.optimize.training.run.base_model, akka.optimize.training.run.engine, akka.optimize.training.run.group_size (only on a reinforcement run), akka.optimize.training.run.id, akka.optimize.training.run.learning_rate, akka.optimize.training.run.lora_rank, akka.optimize.training.run.method, akka.optimize.training.run.settings_sha, akka.optimize.training.workload

Exports and uploads the final trained model adapter weights to artifact storage.

Attributes

akka.optimize.training.checkpoint.artifact.kind

The checkpoint payload represented by the observation.

Value Meaning

trainingState

Engine state needed to resume optimizer and schedule continuity.

weights

Adapter weights used to initialize a branch without optimizer continuity.

akka.optimize.training.checkpoint.failure.reason

A bounded category for a failed checkpoint operation.

Value Meaning

validation

Checkpoint content or its declared identity failed validation.

transfer

Durable storage transfer failed.

catalog

Catalog reservation, commit, or acknowledgement failed.

conflict

An immutable destination already contained different bytes.

internal

The operation failed for another local runtime reason.

akka.optimize.training.checkpoint.operation

The checkpoint operation being measured.

Value Meaning

stage

Copies a newly committed native checkpoint into protected local staging.

publish

Publishes one staged snapshot through durable catalog acknowledgement.

upload

Transfers one artifact kind to its durable storage location.

catalogReserve

Reserves a snapshot through the catalog callback before upload.

catalogCommit

Waits for durable catalog acknowledgement after upload.

materialize

Fetches a restore manifest and its payload into a temporary local directory.

validate

Validates a restore manifest and exact descriptor before payload transfer.

akka.optimize.training.checkpoint.outcome

Whether the measured checkpoint operation completed successfully.

Value Meaning

succeeded

The operation completed successfully.

failed

The operation failed before reaching its completion boundary.

akka.optimize.training.checkpoint.restore.mode

Why the job restores the selected checkpoint artifact.

Value Meaning

resumedAttempt

A later attempt resumes training state produced by the same run.

branchOrigin

A child run starts from an artifact produced by another run.

akka.optimize.training.checkpoint.snapshot.id

Content-derived identifier of the snapshot involved in the operation.

Type string, for example f3a82e7d4b390b66.

akka.optimize.training.checkpoint.source.attempt

Attempt that produced a restored snapshot.

Type int.

akka.optimize.training.checkpoint.source.engine_step

Engine step at which the restored snapshot was produced.

Type int.

akka.optimize.training.checkpoint.source.run_id

Run that produced a restored snapshot.

Type string, for example parent-run.

akka.optimize.training.checkpoint.storage.mode

The storage transport used for checkpoint artifacts.

Value Meaning

objectStorage

Payloads are transferred through object storage.

sharedFilesystem

Payloads are copied through a shared filesystem.

akka.optimize.training.claim.state

Whether an active claim holds its devices or is waiting for them.

Value Meaning

GRANTED

The claim holds its devices.

QUEUED

The claim is waiting for devices to be free.

akka.optimize.training.gpu.clock.domain

The clock a frequency was measured on.

Value Meaning

sm

The streaming multiprocessor clock.

memory

The memory clock.

akka.optimize.training.gpu.derived

Whether the GPU duration was estimated from job start and end timestamps rather than reported directly by the hardware or backend.

Type boolean.

akka.optimize.training.instance

The instance that reported the measurement, the pod it runs as. Carried by the gauges every instance reports alike, so that each instance writes a series of its own and a max across instances has series to aggregate over. Not named service.instance.id, which is a resource attribute that belongs to the runtime.

Type string, for example trainer-7d9f8b6c4-x2k8q.

akka.optimize.training.job.rank

The rank of the worker process in the engine’s worker group, from zero to the group’s size minus one. For a job that trains in its own process, the position of the device among the devices that process holds.

Type int, for example 0, 7.

akka.optimize.training.job.role

The roles the worker process holds, under the engine’s own name for them. verl runs the actor, the rollout engine and the reference model in one process per device, so the value names every role that process holds.

Value Meaning

actor

Computes the policy update.

rollout

Generates rollouts.

ref

Holds the reference model the KL term is measured against.

actor_rollout

The actor and the rollout engine. An adapter run takes its reference from the base model inside the actor.

actor_rollout_ref

The actor, the rollout engine and a separate reference model.

trainer

The one process of a job that trains in its own process, a TRL reinforcement or supervised job. It does all of the job’s work.

akka.optimize.training.job.submission_id

Unique job or submission identifier assigned by the compute cluster for this attempt.

Type string, for example commit-router-7f3a-2.

akka.optimize.training.memory.component

What a part of a GPU’s memory held, at the peak of a training step. The parts other than free sum to what the step’s peak held of the device, where the peak was during the update. A step that generates holds the generation and not the gradients and activations at its peak.

Value Meaning

weights

The model’s frozen weights and buffers. On a quantized base model, the packed weights.

adapter

The trainable adapter weights.

gradients

The gradients of the trainable weights, held from the backward pass to the optimizer step.

optimizer

The optimizer’s state, such as Adam’s two moments, for the trainable weights.

activations

What the forward and backward passes held on top of the weights, gradients and optimizer state: activations kept for the backward pass, and temporary buffers.

generation

What generating rollouts held on top of the weights and optimizer state: the KV cache and the activations of generation. Reinforcement runs on TRL only.

other

Tensors held at the start of a step that are none of the above, such as a quantized model’s scales and cached position tables.

allocator

Memory the caching allocator reserved and did not hand out: kept for reuse, or unusable because it is fragmented.

outside

Device memory in use that the caching allocator does not hold: the CUDA context, library workspaces, and any other process on the device.

free

Device memory nothing held at the step’s peak.

akka.optimize.training.pause.policy

The checkpoint policy selected for a pause request.

Value Meaning

save

Save a fresh full-state checkpoint before stopping.

latest

Stop using the latest already committed checkpoint.

akka.optimize.training.pipeline.id

Unique identifier of the pipeline that started the training run.

Type string, for example support-triage-2026-09-01.

akka.optimize.training.pipeline.stage.name

Name of the pipeline stage that started the training run, as declared in the pipeline definition.

Type string, for example train, train-round-2.

akka.optimize.training.pipeline.stage.outcome

Terminal outcome of the pipeline stage execution.

Value Meaning

completed

The stage completed successfully and produced its expected output.

failed

The stage encountered an error and failed to complete.

akka.optimize.training.pipeline.stage.type

The category or type of pipeline stage executed.

Value Meaning

curate

Filtering and selecting high-quality training examples from raw data.

train

Fine-tuning model parameters on the curated dataset.

evaluate

Evaluating model quality against benchmark test suites or grader models.

akka.optimize.training.pool.accelerator

The device profile the pool states capacity for, or any where the deployment states one capacity for every device it has.

Type string, for example any, nvidia-l4, nvidia-h100-80gb.

akka.optimize.training.pool.serving

Whether the pool holds the devices the serving endpoints draw from, rather than the devices the training jobs draw from.

Type boolean.

akka.optimize.training.profile

The backend profile configuring the training cluster and evaluation serving environment.

Type string, for example ray, stub.

akka.optimize.training.rl.grader.failure.cause

Why a reward term’s grader call produced no score. The rollout is scored zero for that term.

Value Meaning

timeout

The judge did not answer within the call’s timeout.

unavailable

The judge could not be reached, or it replied with something other than a chat completion.

unreadable

The judge replied, and the reply carries no verdict in the expected form.

ungradable

The row’s gold answer is not a JSON label, so the label match has nothing to compare against.

akka.optimize.training.rl.grader.token.type

Whether a count of judge tokens was read or written by the judge. The values match the OpenTelemetry gen_ai.token.type attribute.

Value Meaning

input

Tokens in the prompt the judge was sent.

output

Tokens the judge generated in its reply.

akka.optimize.training.rl.role

The functional role of the process within a disaggregated reinforcement learning loop.

Value Meaning

trainer

The trainer process maintaining model weights and computing optimizer gradient updates.

rollout

The rollout generation worker generating policy rollouts and sample completions.

akka.optimize.training.rl.step.phase

The phase of a training step a duration was measured over, using the engine’s own name for it. verl reports generation, reward computation, advantage estimation, old log-probability computation, actor update, checkpoint save, and the whole step. TRL reports generation, reward computation and actor update, under the same three names.

Type string, for example gen, reward, update_actor, save_checkpoint, step.

akka.optimize.training.rl.term

The declared reward term a measurement belongs to: a rule by its file name, a judge by its model name, the built-in label match, or total for the weighted sum of the shares.

Type string, for example rule:score.py, model:judge-4b, json-label-match, total.

akka.optimize.training.run.attempt

Sequential attempt number for the training run, incremented when a paused or restarted run resumes.

Type int, for example 1, 2.

akka.optimize.training.run.base_model

The base model the run trains an adapter for.

Type string, for example Qwen/Qwen3-4B-Instruct.

akka.optimize.training.run.engine

The training engine executing the run.

Type string, for example trl, verl.

akka.optimize.training.run.group_size

Number of completions generated per prompt in a reinforcement learning run.

Type int, for example 8.

akka.optimize.training.run.id

Unique identifier of the training run.

Type string, for example commit-router-7f3a.

akka.optimize.training.run.learning_rate

The learning rate configured for the run.

Type double, for example 0.000001.

akka.optimize.training.run.lora_rank

The rank of the LoRA adapter the run trains.

Type int, for example 16.

akka.optimize.training.run.method

The training method, either a reinforcement learning algorithm or SFT for supervised fine-tuning.

Type string, for example GRPO, GSPO, SFT.

akka.optimize.training.run.outcome

Terminal outcome of a completed, failed, or cancelled training run.

Value Meaning

COMPLETED

The run completed successfully and produced a trained model artifact.

FAILED

The run failed due to an error.

CANCELLED

The run was cancelled before completion.

akka.optimize.training.run.phase

The current lifecycle execution phase of a training run.

Value Meaning

RESOLVING_DATASET

Reading the registered dataset’s manifest to determine how much the run will train on.

AWAITING_BASE_MODEL

Waiting for the run’s base model to finish preparing, so the weights the job loads are in the cluster’s cache.

DEPLOYING_GRADER

Deploying the evaluation model or grader service used for scoring.

AWAITING_GRADER

Waiting for the grader endpoint to become healthy and ready to serve requests.

AWAITING_CAPACITY

Waiting for the cluster to hold the devices the run claimed, before the job that needs them is submitted.

STARTING_JOB

Submitting and scheduling the training job on the compute backend.

QUEUED

The training job is submitted and waiting for the cluster to schedule it.

PREPARING

The training job is running and preparing, reading and tokenising its data or starting its workers, before its first step.

TRAINING

The training job is actively executing.

PAUSING

Stopping the active job and releasing compute resources while transitioning to PAUSED.

PAUSED

The run is suspended with compute resources released, awaiting resumption.

RECORDING

Registering the trained model artifact and recording final evaluation metrics.

CLEANING_UP

Tearing down ephemeral resources before entering a terminal phase.

COMPLETED

The run completed successfully and produced a trained model artifact.

FAILED

The run failed due to an error.

CANCELLED

The run was cancelled by a user or external request before completion.

akka.optimize.training.run.settings_sha

Content hash of the run’s full hyperparameters. Runs with identical settings on the same dataset share the same value.

Type string, for example 57c93d025d0260af.

akka.optimize.training.service.failure.reason

A bounded category for an unsuccessful service operation.

Value Meaning

validation

Input or artifact validation failed.

conflict

Existing durable state conflicts with the requested operation.

retention

A source could not be retained or a retained artifact could not be deleted.

provider

The training or evaluation provider failed or refused an operation.

storage

Artifact export or deletion failed.

catalog

A catalog entity call failed or was refused.

timeout

The operation exhausted its configured deadline or recovery attempts.

internal

The operation failed for another internal reason.

akka.optimize.training.service.operation

The durable service operation being observed.

Value Meaning

pause

Stops a training attempt and confirms that its resources were released.

branchPreparation

Retains a source snapshot and starts its child training run.

candidatePreparation

Retains, exports, and registers a snapshot as a candidate model.

snapshotCommit

Verifies and commits a snapshot to the catalog.

snapshotDelete

Obtains a retention decision and deletes a snapshot artifact.

snapshotRetention

Takes or gives back an operator’s hold on a snapshot artifact.

comparisonPreparation

Prepares a comparison source and launches target evaluations.

akka.optimize.training.service.outcome

The bounded result of a service operation or attempt.

Value Meaning

succeeded

The operation reached its intended completion boundary.

failed

The operation failed with a known result.

refused

A component definitively rejected the operation.

timedOut

The logical deadline elapsed before confirmation arrived.

uncertain

Recovery ended while an external call could still complete.

blocked

Snapshot retention references intentionally prevented deletion.

akka.optimize.training.service.stage

The bounded stage within the service operation.

Value Meaning

request

Accepts a new logical operation.

acknowledge

Accepts the worker acknowledgement for a requested save.

release

Gives back what an operation holds, whether provider resources or a snapshot reference in the catalog.

retain

Retains a snapshot artifact for dependent work or on an operator’s request.

start

Starts a child workflow or other durable process.

export

Copies an artifact to the serving location.

register

Writes a durable registry entry.

verify

Verifies uploaded snapshot manifests and payload declarations.

commit

Commits verified metadata to the snapshot catalog.

delete

Deletes an artifact from storage.

launch

Starts one evaluation target in a comparison.

notify

Tells a dependent component what this operation produced.

akka.optimize.training.sft.data.fit

How a count of training rows fits the run’s maxSeqLength.

Value Meaning

total

Every training row.

at_max_seq_length

Rows that reached maxSeqLength once tokenised: every row that was cut, and any that fit exactly.

targets_cut

Rows at maxSeqLength whose last token is still one the run trains on, so their target was most likely cut short. None where the loss policy tokenised the rows, because it cuts no target.

dropped

Rows left with no token to train on once cut, which TRL drops.

akka.optimize.training.sft.step.phase

The phase of the supervised loop a duration was measured over.

Value Meaning

input

Fetching and collating the micro-batches of one optimizer update, and moving them to the device. It runs before the step and is not part of it. The first update after a resume reports no input, because its fetch also loads every batch the resume skips.

forward_backward

The forward and backward passes over those micro-batches, up to and including clipping the gradients.

optimizer

The optimizer step.

step

The whole update, from the start of its first forward pass to the end of the step, not counting its input. It is the span throughput and MFU are derived over.

evaluation

One pass over the validation dataset. It runs between steps and is not part of one.

save_checkpoint

Writing one checkpoint. It runs between steps and is not part of one.

akka.optimize.training.workload

Identifier for the logical workload or domain associated with the training run.

Type string, for example commit-messages, support-triage.

error.type

Whether a memory error was corrected. An uncorrected error corrupts what the job computes.

Value Meaning

corrected

Detected and corrected by ECC. Single-bit.

uncorrected

Detected and not corrected. Double-bit.

hw.id

The UUID of the GPU the worker process uses, as NVML reports it. Ray’s node series carry the same value as GpuUuid and dcgm-exporter’s as UUID. The OpenTelemetry hardware semantic conventions identify a GPU by this attribute.

Type string, for example GPU-36e1567d-37ed-051e-f8ff-df807517b396.

hw.name

The product name of the GPU, as NVML reports it.

Type string, for example NVIDIA H100 80GB HBM3.

hw.vendor

The vendor of the GPU.

Type string, for example NVIDIA.

k8s.node.name

The name of the Kubernetes node the worker process runs on. The OpenTelemetry semantic convention of the same name.

Type string, for example gke-gpu-h100-pool-1a2b.

k8s.pod.name

The name of the Kubernetes pod the worker process runs in. The OpenTelemetry semantic convention of the same name.

Type string, for example trainer-ray-gpu-worker-7xk2p.