Telemetry reference
The names, units, and attributes below are generated from the telemetry registry that the trainer and the training job use to record telemetry. See Observability for how to use them.
Metrics
In the Emitter column, service marks a metric that the trainer service publishes, and job marks a metric that the training job publishes.
GPU pools
| Metric | Kind | Unit | Emitter | Attributes | Description |
|---|---|---|---|---|---|
|
gauge |
|
service |
|
Number of active claims on a GPU pool, partitioned by claim state. |
|
gauge |
|
service |
|
Number of devices the active claims on a GPU pool hold or are waiting for, partitioned by claim state. |
|
gauge |
|
service |
|
Number of devices the deployment declares a GPU pool holds, which is what work is admitted against. |
|
gauge |
|
service |
|
Number of devices the pool’s cluster reported holding at its last reading. Absent until the cluster has been read. |
|
gauge |
|
service |
|
Number of devices a workload’s active claims hold or are waiting for across every pool, partitioned by claim state. |
|
gauge |
|
service |
|
Number of devices a workload may hold at once across every pool. |
Training checkpoints
| Metric | Kind | Unit | Emitter | Attributes | Description |
|---|---|---|---|---|---|
|
histogram |
|
job |
|
Duration of a checkpoint staging, publication, upload, catalog, materialization, or validation attempt. |
|
gauge |
|
job |
Number of staged snapshots still awaiting durable publication by this training job, including an in-flight publication. |
|
|
histogram |
|
job |
|
Bytes actually uploaded or downloaded during a checkpoint transfer attempt; omitted when no transferred amount is known. |
Training cost
| Metric | Kind | Unit | Emitter | Attributes | Description |
|---|---|---|---|---|---|
|
counter |
|
service |
|
Total GPU compute time consumed by training jobs. |
Training devices
| Metric | Kind | Unit | Emitter | Attributes | Description |
|---|---|---|---|---|---|
|
gauge |
|
job |
|
Frequency of a GPU clock, averaged over the samples taken since the previous export. |
|
gauge |
|
job |
|
Every reason NVML gave for holding a GPU’s clocks down in any sample since the previous export, as the bitwise OR of NVML’s clock event reason bits. Zero is a GPU held down by nothing, and 0x1 alone is an idle one. Bit 0x2 is an application clock setting, 0x4 software power cap, 0x8 hardware slowdown, 0x20 software thermal slowdown, 0x40 hardware thermal slowdown, 0x80 hardware power brake. In PromQL, |
|
gauge |
|
job |
|
Fraction of time the GPU’s memory was being read or written, averaged over the samples taken since the previous export. NVML reports it as memory utilization. It is not how full the memory is, which is |
|
gauge |
|
job |
|
Memory pages a GPU has retired over its lifetime, by the kind of error that caused it. Reported by GPUs before Ampere. |
|
gauge |
|
job |
|
1 when a GPU has a row remapping or a page retirement that takes effect only when the GPU is reset, and 0 otherwise. A GPU with a pending repair should be reset before it takes another job. |
|
gauge |
|
job |
|
Memory rows a GPU has remapped to spare rows over its lifetime, by the kind of error that caused it. Reported by Ampere and later GPUs. |
|
counter |
|
job |
|
XIDs, the driver’s codes for a GPU fault, reported for a GPU the job holds since the job started watching it. Published from zero, so an alert on its increase sees the first one. Absent on a GPU whose events NVML does not deliver. Read beside |
|
gauge |
|
job |
|
The code of the most recent XID reported for a GPU the job holds. Absent until the GPU reports one. |
Training jobs
| Metric | Kind | Unit | Emitter | Attributes | Description |
|---|---|---|---|---|---|
|
gauge |
|
job |
Number of accelerator devices currently allocated to this training job, reinforcement or supervised. |
|
|
gauge |
|
job |
|
Always 1, one series for each worker process of verl while it runs, and one for each device a TRL job holds while the job runs. It names the rank, role, pod, node and GPU of the process, and the run’s identity is on its resource, so a node or GPU series joins to the run through it. A worker that exits ends its series. |
|
histogram |
|
job |
Duration of one reward term scoring one rollout, including a failed call. Recorded when the step the rollout belongs to is reported. |
|
|
counter |
|
job |
|
Number of reward term calls that produced no score, partitioned by cause. A refused judge call and a judge that fails too many calls in a row stop the job instead, and are not counted here. |
|
counter |
|
job |
|
Number of tokens a judge term reported for its calls, from the usage block of the judge’s reply. A reply without one adds nothing. |
|
gauge |
|
job |
Fraction of step duration that the role spent idle waiting for synchronization or data transfer. |
|
|
histogram |
|
job |
Length in tokens of one generated completion, counted to the last token the model generated. Recorded when the step the rollout belongs to is reported. Recorded on an exponential histogram, so runs with different maximum completion lengths merge into one percentile rather than carrying their own bucket boundaries. |
|
|
histogram |
|
job |
Reward of one rollout, once as the weighted sum under the term |
|
|
gauge |
|
job |
Proportion of tokens whose policy probability ratio was clipped in the most recent optimizer step. |
|
|
gauge |
|
job |
Proportion of tokens whose policy probability ratio was clipped at the upper bound in the most recent optimizer step, where a positive advantage would otherwise have pushed the ratio up further. |
|
|
gauge |
|
job |
Proportion of tokens whose policy probability ratio was clipped at the lower bound in the most recent optimizer step, where a negative advantage would otherwise have pushed the ratio down further. |
|
|
gauge |
|
job |
Proportion of completions generated in the most recent training step that were truncated at the maximum completion length. |
|
|
gauge |
|
job |
Mean length in tokens of the completions generated in the most recent training step. |
|
|
gauge |
|
job |
Longest completion generated in the most recent training step, in tokens. |
|
|
gauge |
|
job |
Shortest completion generated in the most recent training step, in tokens. |
|
|
gauge |
|
job |
Mean length in tokens of the completions in the most recent training step that ended with an end-of-sequence token rather than being truncated at the maximum completion length. Not recorded for a step in which every completion was truncated. |
|
|
counter |
|
job |
Number of steps of the reinforcement learning loop that this job attempt has completed. A step number reported twice in a row counts once. The count starts at zero in each attempt, so a resumed run starts a new series. A series that stops increasing for longer than a checkpoint or a validation pass takes is a stalled loop. |
|
|
histogram |
|
job |
Duration of a single step in the reinforcement learning training loop. |
|
|
gauge |
|
job |
Mean per-token entropy of the completions generated in the most recent training step. |
|
|
gauge |
|
job |
Proportion of prompt groups evaluated in the latest step that exhibited zero reward variance across completions. |
|
|
gauge |
|
job |
Gradient norm before clipping in the most recent optimizer step of the reinforcement learning loop. |
|
|
gauge |
|
job |
KL divergence between the policy and the reference model in the most recent training step. Absent when the run configures no KL penalty. |
|
|
gauge |
|
job |
Learning rate applied in the most recent optimizer step of the reinforcement learning loop. |
|
|
gauge |
|
job |
The most accelerator device memory the training process has held in tensors at once since it started, read at the most recent step. A high-water mark, the same on both engines. Under verl, the most of any of the actor’s ranks. |
|
|
gauge |
|
job |
The most accelerator device memory the caching allocator has held at once since the training process started, read at the most recent step. The figure an out-of-memory error is raised against. The difference from the allocated figure is memory kept for reuse or lost to fragmentation. Under verl, the most of any of the actor’s ranks. |
|
|
gauge |
|
job |
Model FLOPs utilization of the policy update in the most recent training step: the estimated FLOPs per second the update achieved, as a fraction of the dense bf16 peak of the devices that ran it. The update is the step less generation and scoring. Both engines count 6 FLOPs per parameter per token, a full forward and backward pass, although a LoRA update takes no gradient for the frozen weights and does nearer 4. verl estimates from the model configuration, adds attention over the sequence and the embedding, and reports zero for an architecture it has no estimate for. TRL counts the linear layers of the language model and its output head, and reports nothing on a device with no known peak, such as a CPU. |
|
|
gauge |
|
job |
Policy gradient loss of the most recent optimizer step of the reinforcement learning loop. Recorded by verl. TRL reports it only with an entropy bonus, which the trainer does not enable, so a TRL run records none. |
|
|
gauge |
|
job |
Mean reward score evaluated in the most recent reinforcement learning training step. |
|
|
gauge |
|
job |
Standard deviation of the reward across the completions scored in the most recent training step. Zero indicates every completion received the same reward. |
|
|
gauge |
|
job |
Mean weighted contribution of one declared reward term across the completions scored in the most recent training step. The contributions of all terms sum to the term named |
|
|
gauge |
|
job |
Standard deviation of one declared reward term’s weighted share across the completions scored in the most recent training step, including the term named |
|
|
gauge |
|
job |
Proportion of completions scored in the most recent training step that received a reward of zero. |
|
|
gauge |
|
job |
Number of completions scored in the most recent training step. |
|
|
gauge |
|
job |
Prompt and completion tokens processed per second per accelerator device in the most recent training step, over the whole step, generation included. A TRL run trains on one device. |
|
|
gauge |
|
job |
Wall-clock duration of one phase of the most recent training step, as measured by the engine. Reported by verl for every phase it has, and by TRL for generation, reward computation and actor update. |
|
|
gauge |
|
job |
Number of prompt and completion tokens in the batch of the most recent training step. TRL counts the tokens it generated for the step, which is the batch it trains on, and leaves out tool responses inside a completion. |
|
|
gauge |
|
job |
Mean number of tool calls per completion in the most recent training step. Recorded only where the run declares tools. |
|
|
gauge |
|
job |
Proportion of tool calls in the most recent training step that failed. Recorded only where the run declares tools. |
|
|
gauge |
|
job |
Training rows counted by how they fit |
|
|
gauge |
|
job |
Loss over the held-out evaluation split at the most recent evaluation of the supervised fine-tuning loop. |
|
|
counter |
|
job |
Number of steps of the supervised fine-tuning loop that this job attempt has completed. A step number reported twice in a row counts once. The count starts at zero in each attempt, so a resumed run starts a new series. A series that stops increasing for longer than a checkpoint or a validation pass takes is a stalled loop. |
|
|
histogram |
|
job |
Duration of a single step in the supervised fine-tuning loop. |
|
|
gauge |
|
job |
Gradient norm before clipping in the most recent optimizer step of the supervised fine-tuning loop. |
|
|
gauge |
|
job |
Learning rate applied in the most recent optimizer step of the supervised fine-tuning loop. |
|
|
gauge |
|
job |
Training loss at the most recent step of the supervised fine-tuning loop. |
|
|
gauge |
|
job |
The most accelerator device memory the training process has held in tensors at once since it started, read at the most recent supervised fine-tuning step. With more than one device, the figure is for the device that held the most. Means the same as the reinforcement figure of the same name. |
|
|
gauge |
|
job |
The most accelerator device memory the caching allocator has held at once since the training process started, read at the most recent supervised fine-tuning step. With more than one device, the figure is for the device that held the most. Means the same as the reinforcement figure of the same name. |
|
|
gauge |
|
job |
Model FLOPs utilization of the most recent supervised fine-tuning step: the estimated FLOPs per second the step achieved, as a fraction of the dense bf16 peak of the devices that ran it. Counted as the reinforcement figure on TRL is: 6 FLOPs per parameter of the language model’s linear layers and output head per token, although a LoRA update takes no gradient for the frozen weights and does nearer 4. Absent on a device with no known peak, such as a CPU. |
|
|
gauge |
|
job |
Share of the token positions in the batches of the most recent supervised fine-tuning step that were padding. Absent on a packed run, which is collated without padding. |
|
|
gauge |
|
job |
Tokens trained on per second per accelerator device in the most recent supervised fine-tuning step, padding left out. Read beside the padding share of a batch: padded tokens cost time and are not counted. |
|
|
gauge |
|
job |
Wall-clock duration of one phase of the supervised loop, the most recent time it ran: the input, forward and backward passes, optimizer step and whole step of the most recent update, and the most recent evaluation and checkpoint save. |
|
|
gauge |
|
job |
Mean next-token prediction accuracy over the batch of the most recent supervised fine-tuning step. |
|
|
gauge |
|
job |
Number of tokens in the batch of the most recent supervised fine-tuning step, padding left out. |
|
|
gauge |
|
job |
What held the memory of the device at the peak of the most recent training step, one series per part. The phase that did not hold the peak, update or generation, reports zero, so the parts add up to the device’s memory. Under verl, each worker measures its own device, and the series are those of the rank that reserved the most. A TRL job reports them only when it sees one device. Weights, adapter, gradients and optimizer state are counted from the tensors that hold them. Activations and generation are what the step’s peak held above the memory in use when the step or its generation began. The allocator’s share is the step’s peak reserved memory less its peak allocated memory. What is outside the allocator is read from the driver at the end of the step. |
Training pipelines
| Metric | Kind | Unit | Emitter | Attributes | Description |
|---|---|---|---|---|---|
|
gauge |
|
service |
Number of training pipelines currently active. |
|
|
histogram |
|
service |
|
Execution duration of an individual training pipeline stage. |
Training runs
| Metric | Kind | Unit | Emitter | Attributes | Description |
|---|---|---|---|---|---|
|
gauge |
|
service |
|
Number of active training runs, partitioned by phase. |
|
histogram |
|
service |
Total end-to-end execution duration of a training run. |
|
|
counter |
|
service |
|
Total number of failed training runs, partitioned by the phase in which failure occurred. |
|
histogram |
|
service |
Duration spent in a specific training run phase before transitioning. |
|
|
counter |
|
service |
Total number of times training runs have entered each lifecycle phase. |
Training service operations
| Metric | Kind | Unit | Emitter | Attributes | Description |
|---|---|---|---|---|---|
|
counter |
|
service |
|
Number of accepted requests and actual external pause, catalog, and branch preparation attempts observed by the service. |
|
histogram |
|
service |
|
Elapsed time from durable acceptance to a pause acknowledgement, resource release, or terminal preparation result. |
Spans
Training job resource
Resource attributes identifying the training run and job submission across all signals emitted by a training job.
-
akka.optimize.training.pipeline.id(only if a pipeline stage started the run) -
akka.optimize.training.pipeline.stage.name(only if a pipeline stage started the run) -
akka.optimize.training.run.group_size(only on a reinforcement run)
Attributes
akka.optimize.training.checkpoint.artifact.kind
The checkpoint payload represented by the observation.
| Value | Meaning |
|---|---|
|
Engine state needed to resume optimizer and schedule continuity. |
|
Adapter weights used to initialize a branch without optimizer continuity. |
akka.optimize.training.checkpoint.failure.reason
A bounded category for a failed checkpoint operation.
| Value | Meaning |
|---|---|
|
Checkpoint content or its declared identity failed validation. |
|
Durable storage transfer failed. |
|
Catalog reservation, commit, or acknowledgement failed. |
|
An immutable destination already contained different bytes. |
|
The operation failed for another local runtime reason. |
akka.optimize.training.checkpoint.operation
The checkpoint operation being measured.
| Value | Meaning |
|---|---|
|
Copies a newly committed native checkpoint into protected local staging. |
|
Publishes one staged snapshot through durable catalog acknowledgement. |
|
Transfers one artifact kind to its durable storage location. |
|
Reserves a snapshot through the catalog callback before upload. |
|
Waits for durable catalog acknowledgement after upload. |
|
Fetches a restore manifest and its payload into a temporary local directory. |
|
Validates a restore manifest and exact descriptor before payload transfer. |
akka.optimize.training.checkpoint.outcome
Whether the measured checkpoint operation completed successfully.
| Value | Meaning |
|---|---|
|
The operation completed successfully. |
|
The operation failed before reaching its completion boundary. |
akka.optimize.training.checkpoint.restore.mode
Why the job restores the selected checkpoint artifact.
| Value | Meaning |
|---|---|
|
A later attempt resumes training state produced by the same run. |
|
A child run starts from an artifact produced by another run. |
akka.optimize.training.checkpoint.snapshot.id
Content-derived identifier of the snapshot involved in the operation.
Type string, for example f3a82e7d4b390b66.
akka.optimize.training.checkpoint.source.attempt
Attempt that produced a restored snapshot.
Type int.
akka.optimize.training.checkpoint.source.engine_step
Engine step at which the restored snapshot was produced.
Type int.
akka.optimize.training.checkpoint.source.run_id
Run that produced a restored snapshot.
Type string, for example parent-run.
akka.optimize.training.checkpoint.storage.mode
The storage transport used for checkpoint artifacts.
| Value | Meaning |
|---|---|
|
Payloads are transferred through object storage. |
|
Payloads are copied through a shared filesystem. |
akka.optimize.training.claim.state
Whether an active claim holds its devices or is waiting for them.
| Value | Meaning |
|---|---|
|
The claim holds its devices. |
|
The claim is waiting for devices to be free. |
akka.optimize.training.gpu.clock.domain
The clock a frequency was measured on.
| Value | Meaning |
|---|---|
|
The streaming multiprocessor clock. |
|
The memory clock. |
akka.optimize.training.gpu.derived
Whether the GPU duration was estimated from job start and end timestamps rather than reported directly by the hardware or backend.
Type boolean.
akka.optimize.training.instance
The instance that reported the measurement, the pod it runs as. Carried by the gauges every instance reports alike, so that each instance writes a series of its own and a max across instances has series to aggregate over. Not named service.instance.id, which is a resource attribute that belongs to the runtime.
Type string, for example trainer-7d9f8b6c4-x2k8q.
akka.optimize.training.job.rank
The rank of the worker process in the engine’s worker group, from zero to the group’s size minus one. For a job that trains in its own process, the position of the device among the devices that process holds.
Type int, for example 0, 7.
akka.optimize.training.job.role
The roles the worker process holds, under the engine’s own name for them. verl runs the actor, the rollout engine and the reference model in one process per device, so the value names every role that process holds.
| Value | Meaning |
|---|---|
|
Computes the policy update. |
|
Generates rollouts. |
|
Holds the reference model the KL term is measured against. |
|
The actor and the rollout engine. An adapter run takes its reference from the base model inside the actor. |
|
The actor, the rollout engine and a separate reference model. |
|
The one process of a job that trains in its own process, a TRL reinforcement or supervised job. It does all of the job’s work. |
akka.optimize.training.job.submission_id
Unique job or submission identifier assigned by the compute cluster for this attempt.
Type string, for example commit-router-7f3a-2.
akka.optimize.training.memory.component
What a part of a GPU’s memory held, at the peak of a training step. The parts other than free sum to what the step’s peak held of the device, where the peak was during the update. A step that generates holds the generation and not the gradients and activations at its peak.
| Value | Meaning |
|---|---|
|
The model’s frozen weights and buffers. On a quantized base model, the packed weights. |
|
The trainable adapter weights. |
|
The gradients of the trainable weights, held from the backward pass to the optimizer step. |
|
The optimizer’s state, such as Adam’s two moments, for the trainable weights. |
|
What the forward and backward passes held on top of the weights, gradients and optimizer state: activations kept for the backward pass, and temporary buffers. |
|
What generating rollouts held on top of the weights and optimizer state: the KV cache and the activations of generation. Reinforcement runs on TRL only. |
|
Tensors held at the start of a step that are none of the above, such as a quantized model’s scales and cached position tables. |
|
Memory the caching allocator reserved and did not hand out: kept for reuse, or unusable because it is fragmented. |
|
Device memory in use that the caching allocator does not hold: the CUDA context, library workspaces, and any other process on the device. |
|
Device memory nothing held at the step’s peak. |
akka.optimize.training.pause.policy
The checkpoint policy selected for a pause request.
| Value | Meaning |
|---|---|
|
Save a fresh full-state checkpoint before stopping. |
|
Stop using the latest already committed checkpoint. |
akka.optimize.training.pipeline.id
Unique identifier of the pipeline that started the training run.
Type string, for example support-triage-2026-09-01.
akka.optimize.training.pipeline.stage.name
Name of the pipeline stage that started the training run, as declared in the pipeline definition.
Type string, for example train, train-round-2.
akka.optimize.training.pipeline.stage.outcome
Terminal outcome of the pipeline stage execution.
| Value | Meaning |
|---|---|
|
The stage completed successfully and produced its expected output. |
|
The stage encountered an error and failed to complete. |
akka.optimize.training.pipeline.stage.type
The category or type of pipeline stage executed.
| Value | Meaning |
|---|---|
|
Filtering and selecting high-quality training examples from raw data. |
|
Fine-tuning model parameters on the curated dataset. |
|
Evaluating model quality against benchmark test suites or grader models. |
akka.optimize.training.pool.accelerator
The device profile the pool states capacity for, or any where the deployment states one capacity for every device it has.
Type string, for example any, nvidia-l4, nvidia-h100-80gb.
akka.optimize.training.pool.serving
Whether the pool holds the devices the serving endpoints draw from, rather than the devices the training jobs draw from.
Type boolean.
akka.optimize.training.profile
The backend profile configuring the training cluster and evaluation serving environment.
Type string, for example ray, stub.
akka.optimize.training.rl.grader.failure.cause
Why a reward term’s grader call produced no score. The rollout is scored zero for that term.
| Value | Meaning |
|---|---|
|
The judge did not answer within the call’s timeout. |
|
The judge could not be reached, or it replied with something other than a chat completion. |
|
The judge replied, and the reply carries no verdict in the expected form. |
|
The row’s gold answer is not a JSON label, so the label match has nothing to compare against. |
akka.optimize.training.rl.grader.token.type
Whether a count of judge tokens was read or written by the judge. The values match the OpenTelemetry gen_ai.token.type attribute.
| Value | Meaning |
|---|---|
|
Tokens in the prompt the judge was sent. |
|
Tokens the judge generated in its reply. |
akka.optimize.training.rl.role
The functional role of the process within a disaggregated reinforcement learning loop.
| Value | Meaning |
|---|---|
|
The trainer process maintaining model weights and computing optimizer gradient updates. |
|
The rollout generation worker generating policy rollouts and sample completions. |
akka.optimize.training.rl.step.phase
The phase of a training step a duration was measured over, using the engine’s own name for it. verl reports generation, reward computation, advantage estimation, old log-probability computation, actor update, checkpoint save, and the whole step. TRL reports generation, reward computation and actor update, under the same three names.
Type string, for example gen, reward, update_actor, save_checkpoint, step.
akka.optimize.training.rl.term
The declared reward term a measurement belongs to: a rule by its file name, a judge by its model name, the built-in label match, or total for the weighted sum of the shares.
Type string, for example rule:score.py, model:judge-4b, json-label-match, total.
akka.optimize.training.run.attempt
Sequential attempt number for the training run, incremented when a paused or restarted run resumes.
Type int, for example 1, 2.
akka.optimize.training.run.base_model
The base model the run trains an adapter for.
Type string, for example Qwen/Qwen3-4B-Instruct.
akka.optimize.training.run.engine
The training engine executing the run.
Type string, for example trl, verl.
akka.optimize.training.run.group_size
Number of completions generated per prompt in a reinforcement learning run.
Type int, for example 8.
akka.optimize.training.run.id
Unique identifier of the training run.
Type string, for example commit-router-7f3a.
akka.optimize.training.run.learning_rate
The learning rate configured for the run.
Type double, for example 0.000001.
akka.optimize.training.run.lora_rank
The rank of the LoRA adapter the run trains.
Type int, for example 16.
akka.optimize.training.run.method
The training method, either a reinforcement learning algorithm or SFT for supervised fine-tuning.
Type string, for example GRPO, GSPO, SFT.
akka.optimize.training.run.outcome
Terminal outcome of a completed, failed, or cancelled training run.
| Value | Meaning |
|---|---|
|
The run completed successfully and produced a trained model artifact. |
|
The run failed due to an error. |
|
The run was cancelled before completion. |
akka.optimize.training.run.phase
The current lifecycle execution phase of a training run.
| Value | Meaning |
|---|---|
|
Reading the registered dataset’s manifest to determine how much the run will train on. |
|
Waiting for the run’s base model to finish preparing, so the weights the job loads are in the cluster’s cache. |
|
Deploying the evaluation model or grader service used for scoring. |
|
Waiting for the grader endpoint to become healthy and ready to serve requests. |
|
Waiting for the cluster to hold the devices the run claimed, before the job that needs them is submitted. |
|
Submitting and scheduling the training job on the compute backend. |
|
The training job is submitted and waiting for the cluster to schedule it. |
|
The training job is running and preparing, reading and tokenising its data or starting its workers, before its first step. |
|
The training job is actively executing. |
|
Stopping the active job and releasing compute resources while transitioning to PAUSED. |
|
The run is suspended with compute resources released, awaiting resumption. |
|
Registering the trained model artifact and recording final evaluation metrics. |
|
Tearing down ephemeral resources before entering a terminal phase. |
|
The run completed successfully and produced a trained model artifact. |
|
The run failed due to an error. |
|
The run was cancelled by a user or external request before completion. |
akka.optimize.training.run.settings_sha
Content hash of the run’s full hyperparameters. Runs with identical settings on the same dataset share the same value.
Type string, for example 57c93d025d0260af.
akka.optimize.training.service.failure.reason
A bounded category for an unsuccessful service operation.
| Value | Meaning |
|---|---|
|
Input or artifact validation failed. |
|
Existing durable state conflicts with the requested operation. |
|
A source could not be retained or a retained artifact could not be deleted. |
|
The training or evaluation provider failed or refused an operation. |
|
Artifact export or deletion failed. |
|
A catalog entity call failed or was refused. |
|
The operation exhausted its configured deadline or recovery attempts. |
|
The operation failed for another internal reason. |
akka.optimize.training.service.operation
The durable service operation being observed.
| Value | Meaning |
|---|---|
|
Stops a training attempt and confirms that its resources were released. |
|
Retains a source snapshot and starts its child training run. |
|
Retains, exports, and registers a snapshot as a candidate model. |
|
Verifies and commits a snapshot to the catalog. |
|
Obtains a retention decision and deletes a snapshot artifact. |
|
Takes or gives back an operator’s hold on a snapshot artifact. |
|
Prepares a comparison source and launches target evaluations. |
akka.optimize.training.service.outcome
The bounded result of a service operation or attempt.
| Value | Meaning |
|---|---|
|
The operation reached its intended completion boundary. |
|
The operation failed with a known result. |
|
A component definitively rejected the operation. |
|
The logical deadline elapsed before confirmation arrived. |
|
Recovery ended while an external call could still complete. |
|
Snapshot retention references intentionally prevented deletion. |
akka.optimize.training.service.stage
The bounded stage within the service operation.
| Value | Meaning |
|---|---|
|
Accepts a new logical operation. |
|
Accepts the worker acknowledgement for a requested save. |
|
Gives back what an operation holds, whether provider resources or a snapshot reference in the catalog. |
|
Retains a snapshot artifact for dependent work or on an operator’s request. |
|
Starts a child workflow or other durable process. |
|
Copies an artifact to the serving location. |
|
Writes a durable registry entry. |
|
Verifies uploaded snapshot manifests and payload declarations. |
|
Commits verified metadata to the snapshot catalog. |
|
Deletes an artifact from storage. |
|
Starts one evaluation target in a comparison. |
|
Tells a dependent component what this operation produced. |
akka.optimize.training.sft.data.fit
How a count of training rows fits the run’s maxSeqLength.
| Value | Meaning |
|---|---|
|
Every training row. |
|
Rows that reached |
|
Rows at |
|
Rows left with no token to train on once cut, which TRL drops. |
akka.optimize.training.sft.step.phase
The phase of the supervised loop a duration was measured over.
| Value | Meaning |
|---|---|
|
Fetching and collating the micro-batches of one optimizer update, and moving them to the device. It runs before the step and is not part of it. The first update after a resume reports no input, because its fetch also loads every batch the resume skips. |
|
The forward and backward passes over those micro-batches, up to and including clipping the gradients. |
|
The optimizer step. |
|
The whole update, from the start of its first forward pass to the end of the step, not counting its input. It is the span throughput and MFU are derived over. |
|
One pass over the validation dataset. It runs between steps and is not part of one. |
|
Writing one checkpoint. It runs between steps and is not part of one. |
akka.optimize.training.workload
Identifier for the logical workload or domain associated with the training run.
Type string, for example commit-messages, support-triage.
error.type
Whether a memory error was corrected. An uncorrected error corrupts what the job computes.
| Value | Meaning |
|---|---|
|
Detected and corrected by ECC. Single-bit. |
|
Detected and not corrected. Double-bit. |
hw.id
The UUID of the GPU the worker process uses, as NVML reports it. Ray’s node series carry the same value as GpuUuid and dcgm-exporter’s as UUID. The OpenTelemetry hardware semantic conventions identify a GPU by this attribute.
Type string, for example GPU-36e1567d-37ed-051e-f8ff-df807517b396.
hw.name
The product name of the GPU, as NVML reports it.
Type string, for example NVIDIA H100 80GB HBM3.