Pipelines
A pipeline submits an ordered sequence of curate, train, and evaluate stages. Use one to run a repeatable training and evaluation process in a single submission. A later stage can refer to an earlier stage’s output as @STAGE_NAME wherever it specifies a dataset or model. These references include a train stage’s config.dataset or warmStartModelId, an evaluate stage’s model, and the model in a model: grader. Each stage has a durable identity, so an interrupted pipeline resumes without repeating completed work.
Validate, start, read, resume, and cancel a pipeline:
akka-optimize training pipelines start -f pipeline.json --dry-run
akka-optimize training pipelines start -f pipeline.json --wait --exit-status
akka-optimize training pipelines get PIPELINE_ID
akka-optimize training pipelines resume PIPELINE_ID
akka-optimize training pipelines cancel PIPELINE_ID
--dry-run runs the same checks a submission runs and prints what each stage resolved to, without starting the pipeline. It rejects a stage that specifies an unknown base model or dataset, the same refusal a submission gives. Add --show-resolved to also print a train stage’s resolved configuration and execution settings, every default filled in, for a stage whose dataset is already registered:
akka-optimize training pipelines start -f pipeline.json --dry-run --show-resolved
A stage that references an earlier stage’s output, such as @curated, has nothing registered to resolve defaults against yet, so --show-resolved prints nothing for it. get reports the status of every stage. resume continues a stopped pipeline from the unfinished stage.
A pipeline that trains its own judge
The following pipeline trains a judge, gates it on its own score, and then uses it to grade a supervised run and a reinforcement run:
{
"kind": "pipeline",
"stages": [
{
"type": "train", "name": "judge", "workload": "commit-judge",
"baseModel": "Qwen/Qwen3-0.6B",
"config": {"dataset": "commit-judge-train",
"hyperparams": {"lora": {"rank": 8}, "optimizer": {"learningRate": 0.0001}, "training": {"epochs": 3}}}
},
{
"type": "evaluate", "name": "judge-scored",
"model": "@judge", "dataset": "commit-judge-eval", "minMeanScore": 0.7
},
{
"type": "train", "name": "sft", "workload": "commit-messages",
"baseModel": "Qwen/Qwen3-0.6B", "grader": "model:@judge",
"config": {"dataset": "commit-messages-train",
"hyperparams": {"lora": {"rank": 8}, "optimizer": {"learningRate": 0.0001}, "training": {"epochs": 3}}}
},
{
"type": "train", "name": "rl", "workload": "commit-messages",
"baseModel": "Qwen/Qwen3-0.6B", "warmStartModelId": "@sft", "grader": "model:@judge",
"config": {"dataset": "commit-messages-train", "method": "grpo",
"rlHyperparams": {"groupSize": 4, "maxCompletionLength": 64}}
},
{
"type": "evaluate", "name": "rl-scored",
"model": "@rl", "dataset": "commit-messages-eval", "grader": "model:@judge"
}
]
}
The pipeline trains the judge first and checks its score before using it to grade anything. The reinforcement stage warm-starts from the supervised adapter and uses the judge for its reward.
The descriptor
A pipeline descriptor has the following top-level fields:
| Field | Description |
|---|---|
|
|
|
The stages, in the order they run. Required. See Stages. |
|
An ID for this attempt to start the pipeline. The trainer derives the pipeline ID from it, so a request sent again after a dropped connection reaches the pipeline the first request started. The CLI generates a new one for each command. |
|
An annotation. It is not part of the pipeline’s identity. |
Stages
Every stage has the following fields:
| Field | Description |
|---|---|
|
|
|
A name that is unique in the pipeline. Required. A later stage refers to the stage’s output as |
The curate stage
A curate stage judges each exchange in a set of uploaded sessions and keeps the ones that score at or above a threshold. Its output is the dataset of the kept exchanges. It has the following fields:
| Field | Default | Description |
|---|---|---|
|
none |
The uploaded sessions to read, by hash. Required. |
|
none |
The format of the sessions: |
|
|
The language of the exchanges to keep. |
|
A general data-quality instruction |
The instruction the judge scores each exchange with. The judge answers with a score from 0 to 1. |
|
0.7 |
The lowest score an exchange is kept at. From 0 to 1. |
|
none |
The system message for a kept exchange that has none of its own. Required. |
|
none |
An annotation. |
The train stage
A train stage contains the fields of a run descriptor, except kind and submissionId, plus name. baseModel is optional when the workload’s base reference is set. For the run descriptor, see Training runs.
The evaluate stage
An evaluate stage has the following fields:
| Field | Description |
|---|---|
|
The model to score, by ID or as |
|
A base model to score as it is, by repository. Use it to measure the model before training. The serving deployment must already run it. |
|
The workload the score belongs to, by name or ID. A trained model records its workload, so set this for a |
|
The evaluation dataset. Without it, the stage scores the model’s training dataset, which measures reproduction rather than generalization. |
|
The grader. Default |
|
A gate. The pipeline fails when the mean score is below this value. |
|
A ceiling on the length of one scored answer. Without it, the answer length is not capped. |
|
An annotation. |
The gate includes unscored examples in its decision. It passes if the score meets the threshold when unscored examples count as zero. It fails if the score misses the threshold when unscored examples count as one. Otherwise, it rejects the undecided result.