Pipelines

A pipeline submits an ordered sequence of curate, train, and evaluate stages. Use one to run a repeatable training and evaluation process in a single submission. A later stage can refer to an earlier stage’s output as @STAGE_NAME wherever it specifies a dataset or model. These references include a train stage’s config.dataset or warmStartModelId, an evaluate stage’s model, and the model in a model: grader. Each stage has a durable identity, so an interrupted pipeline resumes without repeating completed work.

Validate, start, read, resume, and cancel a pipeline:

akka-optimize training pipelines start -f pipeline.json --dry-run
akka-optimize training pipelines start -f pipeline.json --wait --exit-status
akka-optimize training pipelines get PIPELINE_ID
akka-optimize training pipelines resume PIPELINE_ID
akka-optimize training pipelines cancel PIPELINE_ID

--dry-run runs the same checks a submission runs and prints what each stage resolved to, without starting the pipeline. It rejects a stage that specifies an unknown base model or dataset, the same refusal a submission gives. Add --show-resolved to also print a train stage’s resolved configuration and execution settings, every default filled in, for a stage whose dataset is already registered:

akka-optimize training pipelines start -f pipeline.json --dry-run --show-resolved

A stage that references an earlier stage’s output, such as @curated, has nothing registered to resolve defaults against yet, so --show-resolved prints nothing for it. get reports the status of every stage. resume continues a stopped pipeline from the unfinished stage.

A pipeline that trains its own judge

The following pipeline trains a judge, gates it on its own score, and then uses it to grade a supervised run and a reinforcement run:

{
  "kind": "pipeline",
  "stages": [
    {
      "type": "train", "name": "judge", "workload": "commit-judge",
      "baseModel": "Qwen/Qwen3-0.6B",
      "config": {"dataset": "commit-judge-train",
                 "hyperparams": {"lora": {"rank": 8}, "optimizer": {"learningRate": 0.0001}, "training": {"epochs": 3}}}
    },
    {
      "type": "evaluate", "name": "judge-scored",
      "model": "@judge", "dataset": "commit-judge-eval", "minMeanScore": 0.7
    },
    {
      "type": "train", "name": "sft", "workload": "commit-messages",
      "baseModel": "Qwen/Qwen3-0.6B", "grader": "model:@judge",
      "config": {"dataset": "commit-messages-train",
                 "hyperparams": {"lora": {"rank": 8}, "optimizer": {"learningRate": 0.0001}, "training": {"epochs": 3}}}
    },
    {
      "type": "train", "name": "rl", "workload": "commit-messages",
      "baseModel": "Qwen/Qwen3-0.6B", "warmStartModelId": "@sft", "grader": "model:@judge",
      "config": {"dataset": "commit-messages-train", "method": "grpo",
                 "rlHyperparams": {"groupSize": 4, "maxCompletionLength": 64}}
    },
    {
      "type": "evaluate", "name": "rl-scored",
      "model": "@rl", "dataset": "commit-messages-eval", "grader": "model:@judge"
    }
  ]
}

The pipeline trains the judge first and checks its score before using it to grade anything. The reinforcement stage warm-starts from the supervised adapter and uses the judge for its reward.

The descriptor

A pipeline descriptor has the following top-level fields:

Field Description

kind

pipeline. Required.

stages

The stages, in the order they run. Required. See Stages.

submissionId

An ID for this attempt to start the pipeline. The trainer derives the pipeline ID from it, so a request sent again after a dropped connection reaches the pipeline the first request started. The CLI generates a new one for each command.

description

An annotation. It is not part of the pipeline’s identity.

Stages

Every stage has the following fields:

Field Description

type

curate, train, or evaluate. Required.

name

A name that is unique in the pipeline. Required. A later stage refers to the stage’s output as @NAME.

The curate stage

A curate stage judges each exchange in a set of uploaded sessions and keeps the ones that score at or above a threshold. Its output is the dataset of the kept exchanges. It has the following fields:

Field Default Description

sessionHashes

none

The uploaded sessions to read, by hash. Required.

source

none

The format of the sessions: oasst1, otel, or canonical. Required.

lang

en

The language of the exchanges to keep. canonical sessions carry no language, so this does not apply to them.

judgeInstruction

A general data-quality instruction

The instruction the judge scores each exchange with. The judge answers with a score from 0 to 1.

threshold

0.7

The lowest score an exchange is kept at. From 0 to 1.

systemPrompt

none

The system message for a kept exchange that has none of its own. Required.

description

none

An annotation.

The train stage

A train stage contains the fields of a run descriptor, except kind and submissionId, plus name. baseModel is optional when the workload’s base reference is set. For the run descriptor, see Training runs.

The evaluate stage

An evaluate stage has the following fields:

Field Description

model

The model to score, by ID or as @STAGE_NAME. Set exactly one of model and baseModel.

baseModel

A base model to score as it is, by repository. Use it to measure the model before training. The serving deployment must already run it.

workload

The workload the score belongs to, by name or ID. A trained model records its workload, so set this for a baseModel.

dataset

The evaluation dataset. Without it, the stage scores the model’s training dataset, which measures reproduction rather than generalization.

grader

The grader. Default json-label-match.

minMeanScore

A gate. The pipeline fails when the mean score is below this value.

maxCompletionTokens

A ceiling on the length of one scored answer. Without it, the answer length is not capped.

description

An annotation.

The gate includes unscored examples in its decision. It passes if the score meets the threshold when unscored examples count as zero. It fails if the score misses the threshold when unscored examples count as one. Otherwise, it rejects the undecided result.

Annotations

description on any stage and configName on a train stage are annotations. The pipeline identity excludes them, so two documents that differ only in annotations identify the same pipeline.