0.6.0
This release chooses how each run trains from its plan, and a supervised run can train on several GPUs of one node. Where endpoints and runs share a pool of GPUs, a queued run, evaluation or endpoint reports what it waits for. By default, an endpoint that receives no request for an hour is stopped. A reinforcement run over a text model spends less of each step handing its weights to rollout generation.
Headlines
The run plan decides how a run trains
When you submit a run, the trainer chooses how to train it from its plan: the method, the GPU count, the precision, the model, the dataset, other settings such as a KL penalty, and what the run restores. A supervised run and a reinforcement run can each train on several GPUs of one node, on a GPU count that training base-models get lists for the model.
The trainer refuses a run that it cannot train, and the reason says what the run asked for, such as "this deployment does not train grpo with a KL penalty over an inherited adapter on 4 GPU(s)". training base-models get lists the runs that you can start for a model.
A branch is decided in the same way, with two limits for a trainingState branch. It continues the optimizer state that its parent wrote, so it trains the way its parent did, and it trains on the GPU count that its snapshot was saved on. A weights branch can train on another GPU count that the model supports.
Supervised runs on several GPUs
Set execution.gpus above 1 to train a supervised run across that many GPUs of one node. The run uses the same adapter, checkpoints and snapshots as a run on one GPU. A run resumes, and a trainingState branch trains, only on the GPU count that the snapshot was saved on. A weights branch can use another GPU count. A supervised run on one GPU trains as before.
Endpoints and runs on one pool of GPUs
Where endpoints and runs share a pool of GPUs, an endpoint can wait for a GPU that a run holds, and a run can wait for a GPU that an endpoint holds. Within a pool, the trainer considers waiting claims in the order they were queued. A claim that its workload’s quota blocks is passed over, and a later claim of another workload can start. A claim that does not fit in the pool holds back the claims queued after it, so a large claim is not passed over indefinitely by smaller ones.
What a waiting claim waits on. Read a queued run, evaluation or endpoint to see the claims that hold GPUs of its pool and the claims queued ahead of it. Each entry has its kind, ID and GPU count. An endpoint’s entry also has its base model and how long it has been idle. training runs get, training evaluations get, training endpoints get and a run’s progress line print it, for example held by endpoint ENDPOINT_ID (Qwen/Qwen3-8B, idle 42m); queued ahead: run RUN_ID (4 GPU). With -o json, it is waitingOn, under scheduling for a run or an evaluation.
Idle endpoints stop. By default, an endpoint that receives no request for one hour is stopped, and its GPUs are released. The hour runs from the last request sent through the trainer, or from when the trainer first saw the endpoint answering. Requests sent to the endpoint’s address directly do not count. While an endpoint is READY, training endpoints get shows when it will stop unless it receives a request first. A stopped endpoint is TERMINATED, and its detail says why. A judge endpoint is not stopped while the run that uses it is still running.
Faster reinforcement steps, and larger models
A reinforcement run hands only its adapter to rollout generation after each update, over text and image-text models alike. Previously, every step sent every weight of the base model with the adapter folded in, which on a model the size of Gemma 4 E4B took most of the step. Sending the adapter also needs far less GPU memory, so a model near the size of its GPUs, such as Gemma 4 12B at fp32 on two or four GPUs, no longer runs out of memory at the first update.
You can resume and branch from snapshots written by 0.5.
Other changes
Endpoints that outlast a restart
A READY endpoint returns to STARTING when the process that serves it is lost and starts again. The endpoint keeps its GPUs. It has no address until it is ready again, so a request to it is refused with its phase. The endpoint’s start timeout runs again from that moment, and the endpoint fails if it is not ready before the timeout ends. Previously, the endpoint failed at once.
Breaking changes
-
Idle endpoints. By default, an endpoint that receives no request for one hour is stopped. To keep one up, send it a request before the stop time that
training endpoints getshows. A request that arrives as the endpoint stops is refused with "endpoint ENDPOINT_ID was stopped before this request reached it". -
Full-state branches across GPU counts. A
trainingStatebranch on a GPU count other than its snapshot’s is refused with "this snapshot’s training state was saved across N GPU(s), and this run trains on M. Restore its weights instead". Branch with--restore weightsto change the GPU count. -
training info. The output no longer has anEnginerow, and-o jsonno longer hasrlEngine, because how a run trains is decided for each run.training runs getshows the engine a run trains under asEngine, and-o jsonhas it asengine. -
Refusal wording. A branch from a snapshot that the deployment cannot read reads "this snapshot’s training state is in a layout this deployment cannot read as a branch source. Restore its weights instead", or "this snapshot’s weights are in a layout this deployment cannot read as a branch source". A branch comparison over a snapshot that it cannot load reads "branch comparisons cannot evaluate snapshot SNAPSHOT_ID yet: its weights are in a layout a comparison does not load". A script that matches reason text should match the new sentences.