Graders and rewards

A grader scores a finished model’s answer against an evaluation example. A reward scores every sampled completion at every step of a reinforcement run. Use graders to evaluate models and rewards to guide reinforcement training. A grader runs once per example, while a reward runs once per rollout. A scorer that is affordable as a grader can therefore be too expensive as a reward, so you configure them separately.

Graders

You configure a grader with a specification string. A workload’s evaluation settings specify one, and a run descriptor specifies one for its evaluations. training evaluations start --grader and a pipeline’s evaluate stage specify a grader for an individual evaluation. When no grader is specified, the trainer uses json-label-match.

The following table lists the grader specifications:

Specification Scores by

json-label-match

Per-field exact match of the JSON in the model’s answer against the JSON in the expected answer. Code fences and prose around the JSON are tolerated. This is the default.

model:MODEL_ID

A registered trained model that the trainer provisions as a judge for the duration of the evaluation. The judge’s system prompt comes from its training dataset, so a base model can’t be a judge.

http:URL

An evaluator that you operate. For each example, it receives an envelope with input, output, and expected. It returns JSON with a score between 0 and 1, or a boolean ok or passed, and can include an explanation. An unreachable evaluator fails the evaluation.

bundle:BUNDLE/MODULE or bundle:BUNDLE/MODULE:FUNCTION

A function in a scoring bundle that the trainer serves for the duration of the evaluation. FUNCTION defaults to compute_score. See Scoring bundles.

If the grader returns no verdict for an example, the example is unscored. training evaluations predictions marks it, and the mean excludes it.

json-label-match grades a structured answer. To grade free text, use a judge model, an HTTP evaluator, or a bundle function.

Rewards

A reinforcement run’s config.reward is a list of weighted terms. The reward for a completion is the weighted sum of these terms. The trainer doesn’t normalize weights because rescaling them would change the advantage. A run with no declared reward uses its grader’s kind as the only term.

The following example declares a bundle function and a label match as reward terms:

"reward": {
  "terms": [
    {"kind": "rule", "weight": 1.0, "bundle": "spider-scoring", "module": "reward", "function": "compute_score"},
    {"kind": "json-label-match", "weight": 0.2}
  ]
}

The following table lists the reward term kinds:

Kind Scores by

json-label-match

A per-field match against the example’s assistant message, like the grader of the same name.

rule

A function in a scoring bundle that runs inside the training job. bundle is a name or hash, module is the module in import form, and function is the function. The function defaults to compute_score. This kind has the lowest cost per rollout.

model

A registered model that the trainer starts as a judge at an endpoint for the run’s workload. The endpoint remains available for the entire run. The run’s grader specifies the model to deploy. The endpoint reserves trainer.serving.ray.gpus-per-endpoint devices from the workload’s GPU quota in addition to the devices requested by the run, so the workload needs quota for both. The term identifies the model by configuration hash. A pipeline that trains its judge in an earlier stage therefore refers to it as model:@judge. A model term always runs at an endpoint, so its placement is fixed.

endpoint

A model served inside the cluster at url under the name model. The job posts each completion to this model.

grader

The run’s grader, whatever kind it is.

The bundle function

A rule term and a bundle: grader call the same Python function:

def compute_score(data_source, solution_str, ground_truth, extra_info=None) -> float:
    # Return a score for solution_str.

solution_str contains the completion, and ground_truth contains the example’s assistant message. When the trainer serves the function as a grader, extra_info["input"] contains the prompt. The function returns the score, to which the trainer applies the term’s weight. If the function raises an exception, the evaluation records an error for that example.

Format before correctness

A cold model with a strict correctness reward can produce no correct completions. Every completion in a group then receives the same score, which makes the advantage zero and provides no training signal. Add a second term that rewards the answer format, such as a parseable JSON object or a final line in the expected shape. This term provides a gradient before the model produces correct answers. The correctness term dominates after the model starts producing them.