Graders and rewards
A grader scores a finished model’s answer against an evaluation example. A reward scores every sampled completion at every step of a reinforcement run. Use graders to evaluate models and rewards to guide reinforcement training. A grader runs once per example, while a reward runs once per rollout. A scorer that is affordable as a grader can therefore be too expensive as a reward, so you configure them separately.
Graders
You configure a grader with a specification string. A workload’s evaluation settings specify one, and a run descriptor specifies one for its evaluations. training evaluations start --grader and a pipeline’s evaluate stage specify a grader for an individual evaluation. When no grader is specified, the trainer uses json-label-match.
The following table lists the grader specifications:
| Specification | Scores by |
|---|---|
|
Per-field exact match of the JSON in the model’s answer against the JSON in the expected answer. Code fences and prose around the JSON are tolerated. This is the default. |
|
A registered trained model that the trainer provisions as a judge for the duration of the evaluation. The judge’s system prompt comes from its training dataset, so a base model can’t be a judge. |
|
An evaluator that you operate. For each example, it receives an envelope with |
|
A function in a scoring bundle that the trainer serves for the duration of the evaluation. |
If the grader returns no verdict for an example, the example is unscored. training evaluations predictions marks it, and the mean excludes it.
json-label-match grades a structured answer. To grade free text, use a judge model, an HTTP evaluator, or a bundle function.
Rewards
A reinforcement run’s config.reward is a list of weighted terms. The reward for a completion is the weighted sum of these terms. The trainer doesn’t normalize weights because rescaling them would change the advantage. A run with no declared reward uses its grader’s kind as the only term.
The following example declares a bundle function and a label match as reward terms:
"reward": {
"terms": [
{"kind": "rule", "weight": 1.0, "bundle": "spider-scoring", "module": "reward", "function": "compute_score"},
{"kind": "json-label-match", "weight": 0.2}
]
}
The following table lists the reward term kinds:
| Kind | Scores by |
|---|---|
|
A per-field match against the example’s assistant message, like the grader of the same name. |
|
A function in a scoring bundle that runs inside the training job. |
|
A registered model that the trainer starts as a judge at an endpoint for the run’s workload. The endpoint remains available for the entire run. The run’s grader specifies the model to deploy. The endpoint reserves |
|
A model served inside the cluster at |
|
The run’s grader, whatever kind it is. |
The bundle function
A rule term and a bundle: grader call the same Python function:
def compute_score(data_source, solution_str, ground_truth, extra_info=None) -> float:
# Return a score for solution_str.
solution_str contains the completion, and ground_truth contains the example’s assistant message. When the trainer serves the function as a grader, extra_info["input"] contains the prompt. The function returns the score, to which the trainer applies the term’s weight. If the function raises an exception, the evaluation records an error for that example.
Format before correctness
A cold model with a strict correctness reward can produce no correct completions. Every completion in a group then receives the same score, which makes the advantage zero and provides no training signal. Add a second term that rewards the answer format, such as a parseable JSON object or a final line in the expected shape. This term provides a gradient before the model produces correct answers. The correctness term dominates after the model starts producing them.