Concepts
The trainer uses the following terms for its data, operations, and artifacts. Use these definitions when you read the trainer pages or CLI output.
- Dataset
-
An immutable set of chat-format examples, identified by the content hash of its examples. A name is a movable alias for a hash. A run records the resolved hash, so moving the name later doesn’t change its training data. You register training and evaluation datasets in the same way but keep them separate. For more information, see Datasets.
- Workload
-
One task that a model is specialized for. A workload groups the runs and models for the task. It specifies the evaluation dataset and grader for every candidate, and contains three model references:
base,candidate, andchampion. For more information, see Workloads. - Base model
-
The pretrained model that a run adapts. A base model is named by its repository, such as
Qwen/Qwen3-0.6B, and pinned to a repository commit when it is registered. - Run
-
One training operation over one dataset, from one base model, under one configuration. A run keeps one identity across attempts. For more information, see Training runs.
- Attempt
-
One execution of a run. Resuming a paused run creates another attempt of the same run, with the same identity, metrics, and cost accounting.
- Configuration
-
The part of a run descriptor that determines the trained model: dataset, method, hyperparameters, reward, precision, and sequence length. Two runs with the same configuration and base model are comparable. When you start a run, the trainer reports earlier runs with the same configuration.
- Execution settings
-
The part of a run descriptor that controls how the run executes: devices, checkpoint interval, and retained adapters. Execution settings don’t change the trained model.
- Method
-
How a run learns.
sftlearns from the assistant messages in the dataset.grpoandgspolearn from sampled completions that a reward scores. - Grader
-
The scorer that compares a finished model’s answer with an evaluation example. A workload has one grader, and an evaluation can specify another. A grader can match JSON labels, use a registered model as a judge, call an HTTP evaluator that you operate, or call a function in a scoring bundle. For more information, see Graders and rewards.
- Reward
-
The scorer for every sampled completion of a reinforcement run, defined as a weighted sum of terms. A term can match JSON labels, call a scoring bundle function, ask a model, or use the run’s grader.
- Scoring bundle
-
Python modules and their data files in a zip file. Its content hash identifies the bundle, and you can assign it a name. A reward term or grader specifies a function in the bundle. For more information, see Scoring bundles.
- Snapshot
-
A verified, immutable save from one run attempt. A snapshot contains adapter weights, full training state, or both. Only committed snapshots appear in lists, so you can’t restore from an interrupted upload. For more information, see Checkpoints.
- Branch
-
A new run started from an exact snapshot of a parent run. A branch has its own identity, budget, metrics, and outputs. For more information, see Branches.
- Model
-
A registered artifact that the trainer can serve and evaluate. It can be the final adapter of a run, a snapshot registered as a candidate, or a base model registered as a baseline. For more information, see Evaluation.
- Evaluation
-
One scoring of one model against one dataset with one grader. A score belongs to an evaluation, not to a model, so you can evaluate a model repeatedly.
- Comparison
-
An evaluation of one source snapshot and selected branch models with one pinned dataset, grader, and completion-token limit.
- Pipeline
-
Ordered
trainandevaluatestages submitted together. A later stage refers to an earlier stage’s output as@STAGE_NAME. For more information, see Pipelines.