Concepts

The trainer uses the following terms for its data, operations, and artifacts. Use these definitions when you read the trainer pages or CLI output.

Dataset

An immutable set of chat-format examples, identified by the content hash of its examples. A name is a movable alias for a hash. A run records the resolved hash, so moving the name later doesn’t change its training data. You register training and evaluation datasets in the same way but keep them separate. For more information, see Datasets.

Workload

One task that a model is specialized for. A workload groups the runs and models for the task. It specifies the evaluation dataset and grader for every candidate, and contains three model references: base, candidate, and champion. For more information, see Workloads.

Base model

The pretrained model that a run adapts. A base model is named by its repository, such as Qwen/Qwen3-0.6B, and pinned to a repository commit when it is registered.

Run

One training operation over one dataset, from one base model, under one configuration. A run keeps one identity across attempts. For more information, see Training runs.

Attempt

One execution of a run. Resuming a paused run creates another attempt of the same run, with the same identity, metrics, and cost accounting.

Configuration

The part of a run descriptor that determines the trained model: dataset, method, hyperparameters, reward, precision, and sequence length. Two runs with the same configuration and base model are comparable. When you start a run, the trainer reports earlier runs with the same configuration.

Execution settings

The part of a run descriptor that controls how the run executes: devices, checkpoint interval, and retained adapters. Execution settings don’t change the trained model.

Method

How a run learns. sft learns from the assistant messages in the dataset. grpo and gspo learn from sampled completions that a reward scores.

Grader

The scorer that compares a finished model’s answer with an evaluation example. A workload has one grader, and an evaluation can specify another. A grader can match JSON labels, use a registered model as a judge, call an HTTP evaluator that you operate, or call a function in a scoring bundle. For more information, see Graders and rewards.

Reward

The scorer for every sampled completion of a reinforcement run, defined as a weighted sum of terms. A term can match JSON labels, call a scoring bundle function, ask a model, or use the run’s grader.

Scoring bundle

Python modules and their data files in a zip file. Its content hash identifies the bundle, and you can assign it a name. A reward term or grader specifies a function in the bundle. For more information, see Scoring bundles.

Snapshot

A verified, immutable save from one run attempt. A snapshot contains adapter weights, full training state, or both. Only committed snapshots appear in lists, so you can’t restore from an interrupted upload. For more information, see Checkpoints.

Branch

A new run started from an exact snapshot of a parent run. A branch has its own identity, budget, metrics, and outputs. For more information, see Branches.

Model

A registered artifact that the trainer can serve and evaluate. It can be the final adapter of a run, a snapshot registered as a candidate, or a base model registered as a baseline. For more information, see Evaluation.

Evaluation

One scoring of one model against one dataset with one grader. A score belongs to an evaluation, not to a model, so you can evaluate a model repeatedly.

Comparison

An evaluation of one source snapshot and selected branch models with one pinned dataset, grader, and completion-token limit.

Pipeline

Ordered train and evaluate stages submitted together. A later stage refers to an earlier stage’s output as @STAGE_NAME. For more information, see Pipelines.