Training runs
A run adapts a base model with one dataset and configuration. Use runs to train a model, inspect its progress, and manage its lifecycle. A run keeps the same identity across attempts.
The descriptor
training runs start accepts a JSON or YAML descriptor. The following descriptor from the support triage example defines a complete supervised run:
{
"kind": "run",
"workload": "support-triage",
"baseModel": "unsloth/Llama-3.2-1B-Instruct",
"config": {
"dataset": "support-triage-train",
"method": "sft",
"hyperparams": {"lora": {"rank": 8}, "optimizer": {"learningRate": 0.0001}, "training": {"epochs": 2}},
"maxSeqLength": 2048,
"model": {"precision": "BF16"}
},
"execution": {"gpus": 1, "saveEvery": 20, "keepAdapters": 4},
"description": "two epochs at rank 8",
"configName": "r8-e2"
}
The descriptor has the following top-level fields:
| Field | Description |
|---|---|
|
|
|
The workload for the run, by name or ID. Required. |
|
The base model repository. Optional when the workload’s |
|
The configuration: what model is trained. See Configuration. |
|
The execution settings: how the run is executed. See Execution settings. |
|
A grader for the run’s evaluations. It overrides the workload’s grader. See Graders and rewards. |
|
A registered adapter that a reinforcement run starts from, instead of the base model. |
|
An ID for this attempt to start the run. The trainer derives the run ID from it, so a request sent again after a dropped connection reaches the run the first request started. The CLI generates a new one for each command. State one to make every rerun of a descriptor reach the same run. |
|
Run a smoke check instead of the run. See Smoke check. |
|
Annotations. |
Configuration
Every field of config affects the trained model. The trainer hashes the configuration and uses the hash as its identifier. Two runs with the same configuration and base model use the same training recipe. The start response lists previous runs of that recipe with their scores.
config has the following fields:
| Field | Description |
|---|---|
|
The training dataset, specified by name or hash. Required. |
|
|
|
Supervised hyperparameters, for |
|
Reinforcement hyperparameters, for |
|
The scorer for each sampled completion in a reinforcement run. Defaults to the run’s grader. See Graders and rewards. |
|
The configuration hash of the model that a reinforcement run starts from. The trainer sets it from |
|
How the base model is loaded and what the forward pass keeps. See Model settings. |
|
The longest example that the job trains on, in tokens. Longer examples are truncated. Default 2048. |
|
How the run reads its data, including the dataset it validates on. Optional. See Data. |
|
Which of a conversation’s tokens the run takes its loss over. Optional, and refused on a reinforcement run. See Loss. |
|
Settings the run gives the training engine directly. Optional. See Engine settings. |
An omitted field uses its default. A configuration that specifies a default and one that omits it produce the same hash.
Supervised hyperparameters
hyperparams has four groups: lora, optimizer, scheduler, and training. A refusal names a field by its group, for example optimizer.weightDecay. Every field is part of the configuration hash.
State as much or as little of a group as you want. Each field you leave out takes the default for the loop that runs, so stating one field of a group does not change the rest.
The lora group describes the shape of the adapter:
| Field | Default | Description |
|---|---|---|
|
8 |
The adapter’s rank. The endpoint that serves the model must be able to load an adapter of this rank. |
|
Twice the rank |
The scale at which the update is applied. |
|
0 |
The probability that an adapted activation is dropped during training. Must be at least 0 and less than 1. |
|
|
The modules that the adapter covers, as a regular expression that matches whole module paths. |
The optimizer group describes what performs the update:
| Field | Default | Description |
|---|---|---|
|
|
The algorithm. Only |
|
0.0001 |
The peak learning rate. |
|
0 |
The decoupled decay applied to every adapted weight. Must not be negative. |
|
1 |
The norm that the gradient is clipped to before the update. |
The scheduler group describes how the learning rate changes over the run:
| Field | Default | Description |
|---|---|---|
|
|
|
|
0 |
The fraction of the run’s updates spent rising to the learning rate. Must be at least 0 and less than 1. |
|
0 |
The fraction of the learning rate that the schedule decays to, instead of zero. Only |
The training group describes the loop:
| Field | Default | Description |
|---|---|---|
|
3 |
The number of passes over the dataset. |
|
None |
A cap on optimizer updates. A positive value stops the run before its epochs complete. |
|
Sized from |
The number of sequences in one forward and backward pass on one device. The run is refused if this value times |
|
Sized from |
The number of passes summed into one update. Multiplied by |
|
100 |
The number of optimizer updates between evaluations on the validation dataset. A run that names no validation dataset evaluates nothing, whatever this states. |
|
42 |
The seed that initializes the adapter, orders the data, and drives dropout. Zero takes the default, so state another number to pin a seed of your own. |
If a run states only one of microBatchPerGpu and gradientAccumulation, the service sizes the other so that the update keeps the requested batch. A run that states neither is split to the largest pass whose tokens fit 4096 and accumulated back to 8 sequences in one update. Both halves are resolved before the run is hashed, so the numbers a run trains at are part of its configuration hash.
Every number in these groups takes its default when it is left out or stated as zero. A negative value is refused, naming the field, rather than read as a value that was left out.
Data
data says how the run reads the data it trains on. It has the following fields:
| Field | Default | Description |
|---|---|---|
|
none |
A second registered dataset that the run evaluates on while it trains, by name or hash. Submission refuses a dataset that is not registered, and one that is the training dataset. |
|
|
What happens to a row longer than |
|
|
Whether short rows are packed into one sequence. Packing makes a step cover more rows and makes two runs over the same data see different batches. |
|
|
Whether the rows are shuffled before the first epoch. The order is drawn from |
There is no format field, because a dataset is chat JSONL. There is no chatTemplate field either, because the template is the tokenizer’s: a run that supplied its own would train a model that serves under a different one.
With a validation dataset set, the run evaluates every validationFrequencySteps updates. The validation loss and its perplexity are written to the run’s metrics beside the training loss, and runs get shows the last reading. Read the pair rather than either alone: a training loss that falls while the validation loss does not is a run memorizing its data.
Loss
loss says which of a conversation’s tokens the run takes its loss over. System and user tokens are never trained on, whatever this states. It has the following fields:
| Field | Default | Description |
|---|---|---|
|
|
Which assistant messages are trained on: |
|
|
Whether a selected assistant message’s tool-call tokens count. It composes with |
|
|
Whether a tool result’s own tokens count. A result is what a tool said, not what the model should learn to say. |
|
none |
A dataset column that, where a row carries it, is the mask as given. The rest of this group is then not consulted. |
Submission refuses a loss or a data group on a reinforcement run, which takes its loss from the reward over a rollout and reads its dataset as rollout prompts. A group that writes out nothing but the defaults asks for nothing and is accepted. A branch to a reinforcement method inherits neither group, so branching a supervised run that stated one is not refused for something the branch never wrote.
Reinforcement hyperparameters
rlHyperparams has the following fields, and the same lora, optimizer, and scheduler groups as hyperparams. Three defaults differ under this loop, each of them what the reinforcement engine already did: optimizer.learningRate is 0.00001, optimizer.weightDecay is 0.01, and scheduler.kind is constant. The reinforcement engines train the adapter without a dropout and follow a constant or cosine schedule, so a lora.dropout above zero or a linear schedule is refused.
| Field | Default | Description |
|---|---|---|
|
1 |
The number of passes over the prompts. |
|
None |
A cap on optimizer updates. A positive value overrides |
|
4 |
The number of completions sampled per prompt. The trainer computes group statistics from these completions. |
|
8 |
The number of prompts rolled out per step. A step generates |
|
0.04 |
The KL penalty toward the reference policy. Zero is no penalty. |
|
0.7 |
The sampling temperature. Zero is greedy. |
|
256 |
The longest completion sampled, in tokens. |
|
0 |
Not supported. The trainer refuses a value above zero: every step trains on-policy, on rollouts generated from the current weights. Leave it out or set it to 0. |
The trainer preserves the configured values of klBeta, temperature, maxSteps, and stalenessThreshold because zero is valid for each field. The other fields use their defaults when omitted or set to zero.
Model settings
The model object has the following fields:
| Field | Default | Description |
|---|---|---|
|
|
How the trainer loads the base model’s weights: |
|
|
Whether the backward pass recomputes activations instead of keeping them. This trades compute for the memory that a long-context run needs. If omitted, the value is derived from |
|
|
The kernel the attention layers run: |
|
The deployment’s, usually |
The precision that verl holds the trained model’s weights in: |
Engine settings
engineSettings gives verl settings that change only memory and speed. Use them to fit a new model to its devices before a setting becomes a run setting of its own. It has the following fields:
| Field | Description |
|---|---|
|
|
|
Each setting’s value, by its name. Required. |
A run can set the following settings, where its deployment allows each one:
| Setting | Default | Description |
|---|---|---|
|
|
Fetches the next layer’s weights while the current layer computes. A step can be faster, and it needs more memory. |
|
|
Frees each layer’s gathered weights after the forward pass. |
|
|
Pads each pass up to a multiple of |
|
1024 |
The multiple |
|
|
Splits a long prompt’s prefill into chunks. Reinforcement runs only. |
|
|
Reuses the cache for prompts that start alike, such as the completions of one group. Reinforcement runs only. |
|
|
Runs the rollout engine without CUDA graphs. It starts faster and holds less memory, and it generates more slowly. Reinforcement runs only. |
|
The step’s rollouts |
The most sequences the rollout engine generates at once. Reinforcement runs only. |
|
The larger of |
The most tokens the rollout engine processes in one batch. Reinforcement runs only. |
The following example sets two of them:
"engineSettings": {
"engine": "verl",
"settings": {"engine.forward_prefetch": true, "rollout.enforce_eager": true}
}
The settings are part of the configuration hash, and training runs get shows them with the configuration. A run that states them trains under verl. The trainer refuses it, before it claims devices, where the base model’s catalog entry does not list verl.
The trainer also refuses a run, naming the setting, when the setting:
-
is not one the deployment allows,
-
is for the rollout engine and the run is supervised, or
-
has a value of the wrong kind, such as a number for
engine.forward_prefetch.
A setting the run’s plan already decides, a setting that names a path or code, and the seed are never allowed.
A branch keeps its parent’s engine settings, less any its own method has no use for, such as a rollout setting on a supervised branch.
Execution settings
execution has the following fields:
| Field | Default | Description |
|---|---|---|
|
1 |
The number of devices that the training job claims. For a reinforcement run under verl, all devices must be on one node. Use a value that the deployment’s catalog admits. |
|
0 |
Not supported. The trainer refuses a value above zero: the rollout engine runs on the training devices. Leave it out or set it to 0, and count every device in |
|
chosen by the job |
The maximum number of sequences per device in one forward and backward pass. It changes only peak memory. Accumulation combines the same rollouts in the same optimizer step. |
|
off |
The periodic checkpoint interval, in optimizer updates. When it is off, the run writes only one checkpoint at the end. An interrupted run therefore has no checkpoint from which to restart. |
|
5 |
The number of unreferenced adapter checkpoints that the catalog keeps for the run. |
|
chosen by the catalog |
The accelerator to train on, as the catalog names it. Omit it when a deployment has only one kind of device. If the named device has no qualified capability for the model, the trainer refuses the run and names the devices that do. |
|
chosen by the catalog |
The training engine, |
|
The deployment’s, usually 0.5 |
The share of each device that a reinforcement run’s rollout engine takes for its weights and its cache. See Memory settings. |
|
The deployment’s, usually |
Whether the trained model’s weights move to host memory between passes. See Memory settings. |
|
none |
The host memory that the run’s loader needs from the node. The trainer checks this value against the named device’s own host memory and refuses the run at submission if the value is higher. It is not passed to the job: the pod’s memory request is the cluster’s to set. Omit it when the device states no host memory, which accepts any value. |
Execution settings are recorded with the run and are not part of the configuration hash.
Memory settings
Three settings trade memory against speed or learning. Each one applies only under verl. A run that states none of them trains at the deployment’s values.
| Setting | Trade |
|---|---|
|
A larger share gives the rollout engine more room for its cache, so it generates more completions at once. It leaves less memory for training. A model that fills most of the device needs a larger share to load at all. A reinforcement run only, since a supervised run has no rollout engine. |
|
|
|
|
The deployment decides which of these a run can state and the range of each. The trainer refuses a run before it claims devices, and names the setting, when:
-
the deployment does not let a run state the setting,
-
memoryShareis outside the deployment’s range, or the run is supervised, -
weightPrecisionis not one the deployment allows, or -
no engine that the catalog lists for the run can honor the setting.
A run that states a setting trains under verl. TRL honors none of the three, so a run whose catalog entry lists only TRL is refused.
A run records the value of each setting when it is created. A later change to the deployment’s values does not change the run, a branch of it, or a retry. training runs get shows the recorded values for a verl run and marks the ones the run stated. training runs engine-config shows the values the job passed to verl.
Where a reinforcement run generates rollouts
A reinforcement run generates its rollouts on the same devices it trains on. Each of the gpus devices holds a training rank and a rollout engine, and a step takes turns: every device generates its share of the step’s rollouts, then every device takes part in the update. To run on four GPUs, set gpus to 4 and leave rolloutGpus out. A layout that sets devices aside for generation, such as three for training and one for rollouts, is not supported.
The deployment states the most devices a run can use, and counts gpus and rolloutGpus together against it. The catalog states the counts each model is qualified at on each device. A run at any other count is refused, and the refusal names the counts that are qualified.
Start a run
Validate a descriptor, start the run, and wait for it to finish:
akka-optimize training runs start -f run.json --dry-run
akka-optimize training runs start -f run.json -o json --jq .runId
akka-optimize training runs wait RUN_ID --exit-status
--dry-run validates the descriptor against the service without training. Without it, the command validates the descriptor and asks for confirmation. --force skips the prompt. Every submission starts a new run, so submitting the same descriptor twice trains it twice.
Add --show-resolved to --dry-run to see the configuration and execution settings the run would actually train under, every default filled in and every dataset and reward reference read:
akka-optimize training runs start -f run.json --dry-run --show-resolved
This reports the same figures training runs get reports once the run exists, so it answers what a descriptor that named few fields resolves to before anything is spent. -o json includes the resolved configuration as structured data.
Smoke check
A smoke check trains on the first rows of the dataset until the model has memorized them. It then checks that the model reproduces them. A model that cannot reproduce rows it has seen many times has a problem in its data, masking, or tokenizer rather than a hard task. The check finds that in minutes, before a full run spends hours.
A smoke check takes everything else from the descriptor. Add smoke to the descriptor to run one:
| Field | Default | Description |
|---|---|---|
|
200 |
The rows taken from the start of the dataset. These are also the rows scored afterwards. |
|
2000 |
The cap on optimizer updates. Raise it with |
|
0.98 |
The reproduction score a working run reaches. It must not be above 1. |
training runs get shows the verdict. training runs wait --exit-status exits with a non-zero status when the check fails.
Base models
A run can specify only a base model that the deployment allows and that someone has selected. To list the allowed models and select one, see Base models.
If a descriptor specifies a model that the deployment doesn’t allow, the trainer rejects the run. Specify a model from the list, or ask your operator to allow the model that you need.
Before it admits a run, the trainer also checks the adapter rank and sequence length against what the deployment supports.
Methods
Supervised fine-tuning
sft learns from the assistant messages in the dataset. Use it when every example has a known correct output, such as classification, extraction, or another task with a manually written or curated dataset.
Reinforcement learning
grpo and gspo sample groupSize completions per prompt and score each completion with the reward. They update the model toward completions that score above the group mean. Use them when you can check but can’t enumerate correct output, such as executable code, SQL that must return the correct rows, or an answer with a verifiable final value.
A reward with only a strict correctness term gives a cold model no signal. Every completion in a group scores zero, so the advantage is zero. Add a format term so that the model has a gradient before it produces a correct answer. Alternatively, warm-start from a supervised adapter with warmStartModelId.
You can train an image-text model such as Gemma 4 E4B with grpo or gspo over a text dataset under TRL. As in a supervised run, the adapter updates the language backbone and leaves the vision and audio towers unchanged. Under verl, your deployment must enable grpo and gspo for image-text models; otherwise, the model supports only sft. Run base-models get with the model’s repository to see the methods that your deployment supports. Reinforcement runs cannot train on images. The trainer rejects a grpo or gspo run whose dataset contains images, regardless of the base model.
The following descriptor, from the GSM8K example, is a descriptor for a reinforcement run:
{
"kind": "run",
"workload": "gsm8k",
"configName": "grpo-100-steps",
"baseModel": "Qwen/Qwen3-0.6B",
"config": {
"dataset": "gsm8k-train",
"method": "grpo",
"rlHyperparams": {
"epochs": 1,
"maxSteps": 100,
"groupSize": 8,
"klBeta": 0.001,
"temperature": 1.0,
"maxCompletionLength": 512,
"lora": {
"rank": 32
},
"optimizer": {
"learningRate": 1e-5
}
},
"reward": {
"terms": [
{"kind": "rule", "weight": 1.0, "bundle": "gsm8k-scoring", "module": "reward"}
]
},
"maxSeqLength": 1536
},
"execution": {"gpus": 1, "saveEvery": 10, "keepAdapters": 4}
}
Lifecycle
To read a run, list open runs, and cancel a run, run the following commands:
akka-optimize training runs get RUN_ID
akka-optimize training runs list --phase open
akka-optimize training runs cancel RUN_ID
The run detail reports the phase, progress, cost, timing, and any failure reason. observedStep counts the updates reported by the worker. savedStep is the selected restorable checkpoint. The difference between these values is the work that a restore repeats.
A run reports one of the following phases. A first attempt passes through them in order and skips phases that it doesn’t need. A resumed run re-enters at AWAITING_BASE_MODEL, DEPLOYING_GRADER or AWAITING_CAPACITY. If the cluster stops a job, its run moves from TRAINING to PAUSING.
| Phase | What the run is doing |
|---|---|
|
Reading the dataset’s manifest. |
|
Waiting for the run’s base model to finish preparing. |
|
Starting the judge specified by a reinforcement run’s reward and waiting for it to respond. |
|
Waiting for the cluster to allocate the devices requested by the run. |
|
Submitting the training job. |
|
Waiting for the cluster to schedule the submitted job. The wait depends on the cluster’s other work. |
|
The job is running and has not trained a step yet: reading the dataset, tokenising its rows, or starting its training workers. The stage names which, and tokenising counts rows done of rows total. A supervised run reports it; a reinforcement run goes from |
|
The job is running. |
|
Stopping the job and releasing its resources until the run resumes. |
|
Registering the trained adapter as a model. |
|
Releasing the run’s resources before a terminal phase. |
|
Terminal. |
A completed run registers its final adapter as a model. training runs get shows the model. Cancelling a run releases its devices but retains its committed snapshots, which you can still register or use as branch sources.
For training runs pause and resume, see Checkpoints. For training runs branch, see Branches.