Observability

The trainer and its training jobs publish metrics and spans through OpenTelemetry. Use them to build dashboards and alerts that cover every run in a deployment. To inspect a single run, see Observe a run.

Signal sources

Two processes publish training telemetry:

  • The trainer service publishes run and pipeline lifecycle events, GPU time, and the operations that it performs for a run. These metrics carry the workload but not the run ID, so the number of series grows with the number of workloads, not the number of runs.

  • A training job publishes what it measures during training: figures for each step, checkpoint operations, and spans. Every signal from a job carries the training job resource attributes, which identify the run, attempt, workload, pipeline stage, and training settings.

The reference later on this page names the process that emits each metric. The trainer passes each job the collector endpoint that trainer.telemetry.collector-endpoint sets, and the trace context of the submission. As a result, the spans of a job join the trace that started the run.

The run record is the authoritative account of a run. Use telemetry for trends and alerts. Metric export and run persistence are not atomic, so a worker failure during export can lose or duplicate an observation. Traces and structured logs carry exact identities. Metric attributes do not.

Alert on a stalled run

When a training job stops responding, for example on a stuck collective operation, it keeps exporting its step gauges at their last value, so the gauges look healthy. To detect a stalled job, alert on the step counters instead. akka.optimize.training.rl.step.count and akka.optimize.training.sft.step.count increase by one for each step that the job completes, and the job exports the total at every interval while it runs. If the total stops increasing, the job has not completed a step.

In Prometheus, through a collector that uses remote write, the counters are named akka_optimize_training_rl_step_count_total and akka_optimize_training_sft_step_count_total. The following query matches a reinforcement job that has not completed a step in 15 minutes:

increase(akka_optimize_training_rl_step_count_total[15m]) == 0

Base the window on the usual step time of the run. When a job exits, its series stops, so the query no longer matches it. A job that has not completed its first step has no series.

Each attempt has its own count. A resumed run is a new attempt, so its count starts at zero under a new akka.optimize.training.run.attempt value. A step that the job repeats after a restore counts again, so a restore does not look like a stall.

Reward graders

A reinforcement run scores every rollout with each reward term that it declares. A judge term calls its judge once for each rollout. The job publishes the following metrics, partitioned by the term in akka.optimize.training.rl.term:

  • akka.optimize.training.rl.grader.duration records how long one term takes to score one rollout, including failed calls.

  • akka.optimize.training.rl.grader.failures counts the scores that fail, partitioned by akka.optimize.training.rl.grader.failure.cause.

  • akka.optimize.training.rl.grader.tokens counts the input and output tokens that the judge reports in its reply.

A failed score counts as zero for its term, so a judge that times out or returns no verdict lowers the reward. The failure count distinguishes that case from a policy that is not improving. The following table describes each cause:

Cause What happened

timeout

The judge did not respond in time.

unavailable

The job could not connect to the judge, or the judge did not respond with a chat completion.

unreadable

The judge responded without a verdict. Check the judge’s prompt and the model that it runs.

ungradable

The gold answer in the row is not a JSON label, so the label match term has nothing to compare. Check the dataset.

If a judge refuses the job’s credentials or fails too many consecutive calls, the job stops and does not count those calls. The job reports the calls for a step when it reports that step.

Reward and completion-length distributions

The gauges for each step give the mean, spread, and zero fraction of the reward, the mean and spread of each term’s share, and the mean, shortest, and longest completion length. The following histograms record every rollout, so a dashboard can show percentiles, a bimodal reward, or a long tail of completion lengths:

  • akka.optimize.training.rl.rollout.reward records the reward of each rollout under the term total, and the weighted share of each declared term under the name of that term.

  • akka.optimize.training.rl.rollout.completion.length records the length of each completion in tokens.

The reward buckets are fixed: 0, 0.1 to 0.9 in steps of 0.1, then 0.99, 1, 2, and 5. The bucket at or below 0 holds the rollouts that scored nothing, and the bucket from 0.99 to 1 holds the full scores. Weights are not normalized, so a reward can exceed 1.

The completion length is recorded on an exponential histogram, so the runs in one query merge into one set of percentiles, whatever their maximum completion length. The histogram does not single out the completions truncated at the maximum length. akka.optimize.training.rl.step.completion.clipped_ratio reports that share for each step. If it increases, the policy is producing longer completions.

Throughput and efficiency

A reinforcement run publishes the following metrics for each step:

  • akka.optimize.training.rl.step.tokens reports the number of tokens that the step processed.

  • akka.optimize.training.rl.step.throughput reports those tokens per second.

  • akka.optimize.training.rl.step.mfu reports the model FLOPs utilization (MFU) of the policy update, from 0 to 1.

Use these metrics to compare training speed across runs, models, or pool shapes. The two engines estimate FLOPs differently, so compare MFU between runs on the same engine. verl estimates from the model configuration, and reports an MFU of zero for a model architecture without a FLOPs estimate. TRL counts the linear layers of the language model and its output head, and publishes no MFU on a device with no known peak, such as a CPU.

Find the GPUs a run uses

Each worker process of a reinforcement run on verl publishes akka.optimize.training.job.placement, a gauge with the value 1, for as long as the process runs. A supervised run and a reinforcement run on TRL train in one process, which publishes one series for each GPU that it holds, under the role trainer. Like every signal from a job, the series carries the run ID. It also carries the rank and role, the pod and node that the process runs on, and the UUID of its GPU in hw.id. Use it to join the cluster’s GPU and node metrics to a run.

The GPU UUID is the only key that the Ray GPU metrics, the job’s metrics, and a DCGM exporter share. A GPU index is not a join key, because Ray counts the devices that a container sees and DCGM counts the devices on the host. The following query adds the role and rank of each worker in a run to Ray’s GPU utilization:

ray_node_gpus_utilization
  * on(GpuUuid) group_left(akka_optimize_training_job_role, akka_optimize_training_job_rank)
    label_replace(akka_optimize_training_job_placement{akka_optimize_training_run_id="$run"},
                  "GpuUuid", "$1", "hw_id", "(.+)")

Filter the placement to one run. When a GPU passes to another attempt or run, the old placement stays in the query lookback for a few minutes, and an unfiltered join finds two series for one GPU.

Rollout engine

On verl, each rollout server publishes vLLM’s own metrics: requests running and waiting, KV cache usage, prompt and generation tokens, preemptions, time to first token, whether the engine is awake, and more. vLLM documents each one. They are published through Ray, so their names start with ray_ and have _ where vLLM has :. For example, vllm:num_requests_running is ray_vllm_num_requests_running, and the counter vllm:generation_tokens is ray_vllm_generation_tokens_total. A run on TRL publishes none of them.

Each series carries the run ID and attempt, as the job’s own metrics do, and the pod of the rollout server:

ray_vllm_kv_cache_usage_perc{akka_optimize_training_run_id="$run"}

The engine sleeps while the actor updates, so ray_vllm_engine_sleep_state{sleep_state="awake"} falls to 0 in the update phase of each step. A step shorter than the scrape interval does not show.

GPU health

A supervised run and a reinforcement run on TRL read each GPU that the job uses through NVML every five seconds. On verl, each worker process reads the GPU that it holds, the one its placement names. Power, GPU utilization, memory activity, and clock frequency report the mean of the samples since the previous export. The throttle reasons report every reason in those samples. The other figures report the latest sample.

Where the OpenTelemetry hardware conventions define a figure, the job uses that name, for example hw.power, hw.temperature, hw.gpu.utilization, and hw.gpu.memory.usage. The other figures are under akka.optimize.training.gpu. Every series carries the GPU’s UUID in hw.id, which joins it to the Ray GPU metrics.

akka.optimize.training.gpu.clock.throttle_reasons is a bitmask of NVML’s clock event reasons. A value greater than 1 means that something other than idleness held the clocks down. To read one reason, such as hardware slowdown (0x8), use floor(x / 8) % 2 in PromQL. The reference lists every bit.

hw.gpu.memory.utilization is the fraction of the memory in use. The fraction of time that the memory was read or written, which NVML calls memory utilization, is akka.optimize.training.gpu.memory.activity.

An XID is the driver’s code for a GPU fault. The job waits for XIDs on each GPU it holds, counts them in akka.optimize.training.gpu.xid.errors, and keeps the code of the most recent one in akka.optimize.training.gpu.xid.last. It also writes each XID and its GPU to its output. The count is published from zero, so an alert on increase() sees the first XID. Some codes are faults in the job’s own kernels, such as an illegal memory access, and others are hardware faults. See NVIDIA’s XID reference for what each code means.

If the job cannot read NVML, it logs one line and trains without these figures. A verl worker whose GPU cannot be read publishes its placement without hw.id, and none of these figures.

Capacity over time

The akka-optimize training capacity command reports what holds the GPU pools at the moment you run it. The service publishes the same figures as gauges, so a dashboard can show occupancy and queued claims over time. The following table maps each gauge to a column of the command’s output:

Gauge Attributes training capacity column

akka.optimize.training.pool.devices.declared

pool, instance

DECLARED

akka.optimize.training.pool.devices.observed

pool, instance

OBSERVED. The gauge has no value until the service reads the cluster.

akka.optimize.training.pool.devices.claimed

pool, claim state, instance

The devices on the pool that claims hold or wait for.

akka.optimize.training.pool.claims

pool, claim state, instance

The claims on the pool.

akka.optimize.training.workload.devices.claimed

workload, claim state, instance

The devices that the claims of a workload hold or wait for.

akka.optimize.training.workload.devices.quota

workload, instance

The quota of a workload.

Two attributes identify a pool. akka.optimize.training.pool.serving is true for the pool that serving endpoints use. akka.optimize.training.pool.accelerator is the device profile, or any for a deployment that states one capacity for all of its devices. The claim state, akka.optimize.training.claim.state, is GRANTED for held devices and QUEUED for devices that a claim waits for.

The gauges carry no run, evaluation, or endpoint ID. To find which claims hold a pool, use training capacity.

Every instance of the service reports the same values, each under its own akka.optimize.training.instance, the pod that reported it. Aggregate across instances with max, not sum. For example, the following query reports the devices that claims hold on each pool:

max without (akka_optimize_training_instance) (akka_optimize_training_pool_devices_claimed{akka_optimize_training_claim_state="GRANTED"})

After the last claim on a pool or of a workload ends, an instance keeps reporting that pool or workload at 0 until the instance restarts. An idle pool reads 0 and does not disappear from a dashboard.

Reference

See Telemetry reference for the generated names, units, and attributes of every metric and span that the trainer and the training job use to record telemetry.