GSM8K: reinforcement learning with a scoring bundle

This example trains a model on grade-school word problems. Each gold answer is a reasoning chain that ends in ## NUMBER. Use this example for a task whose answers can be checked but not enumerated. The model samples its own solutions and receives a reward when the final number is correct. You provide the reward function in a scoring bundle, and the same function grades the finished model.

GSM8K is a standard control for training with verifiable rewards. Published curves provide a reference. If the reward doesn’t improve in this example, check the training recipe before investigating the task.

The example directory contains the data preparation script, scoring bundle, run descriptor, and walkthrough.sh. The walkthrough runs each step on this page against the trainer at TRAINER_URL. To download the directory, see the GSM8K archive.

Before you begin

  • A trainer deployed in your Akka project, with its address in TRAINER_URL. akka service get trainer shows the route hostname. See Akka Optimize CLI.

  • An authenticated CLI session. Run akka auth login to authenticate akka.

  • Python 3. The data script uses the standard library only.

  • Network access to datasets-server.huggingface.co, where the data script reads the rows. The API limits the request rate, so fetching a large ROWS value takes several minutes. The script caches rows, so subsequent runs use the cached data.

  • A base model that your trainer allows. Run akka training base-models list to list the models, and akka training base-models select <repository> to select one. See Base models.

The walkthrough reads ROWS and EVAL_ROWS from the environment. Start with the default ROWS=200 EVAL_ROWS=50. These values show the reward changing over 100 steps; they aren’t intended to produce a useful model.

The data

prepare_dataset.py fetches openai/gsm8k from the Hugging Face datasets-server API. It uses only the Python standard library. Training rows come from the beginning of the train split, and evaluation rows come from the beginning of the test split.

To prepare the data, run:

python3 prepare_dataset.py --rows "$ROWS" --eval-rows "$EVAL_ROWS"

To register the datasets, run:

akka-optimize datasets create -f data/train.jsonl --name gsm8k-train
akka-optimize datasets create -f data/eval.jsonl --name gsm8k-eval

The system prompt requests the final answer on a separate last line in the form ## NUMBER. The reward reads this format.

The scoring bundle

The bundle is one module, reward.py, with one function:

import re

FINAL = re.compile(r"####\s*(-?[\d,]+(?:\.\d+)?)")

# What a well-formed answer is worth before anyone asks whether it is right. Small against the
# correctness term, so a model cannot do better by emitting the marker and nothing else.
FORMAT_WEIGHT = 0.1
CORRECT_WEIGHT = 1.0


def final_answer(text):
    """The last '#### <number>' in a completion, as a number. None where there is none."""
    matches = FINAL.findall(text or "")
    if not matches:
        return None
    try:
        return float(matches[-1].replace(",", ""))
    except ValueError:
        return None


def compute_score(data_source=None, solution_str=None, ground_truth=None, extra_info=None):
    answered = final_answer(solution_str)
    if answered is None:
        return 0.0
    expected = final_answer(ground_truth)
    if expected is not None and answered == expected:
        return FORMAT_WEIGHT + CORRECT_WEIGHT
    return FORMAT_WEIGHT

The reward combines two terms. A cold model with only a strict correctness reward can produce no correct answers. Every completion in a group then receives the same score, which makes the advantage zero and provides no training signal. The format term gives a small score to a well-formed ## NUMBER line before checking the number. This provides a gradient before the model answers correctly. The score is small enough that emitting only the marker never scores higher than answering.

Push the directory as a named bundle. A run records the resolved hash when it is admitted, so changing the code under that name doesn’t change accepted work.

To push the bundle, run:

akka-optimize training bundles push scoring/ --name gsm8k-scoring

The workload

The workload uses the same function as its grader. The trainer serves the function for the duration of each evaluation. Create the workload and set its evaluation settings:

akka-optimize training workloads get gsm8k >/dev/null 2>&1 ||
  akka-optimize training workloads create gsm8k
akka-optimize training workloads set-evaluation gsm8k --dataset gsm8k-eval --grader bundle:gsm8k-scoring/reward:compute_score

The trainer rejects create when the name exists, so the walkthrough runs it only when the workload is missing. The walkthrough runs set-evaluation every time. As a result, it applies this dataset and grader to a workload left partially configured by an earlier pass.

Select the base model for the run. A run can specify only a model that the deployment allows and that someone has selected:

akka-optimize training base-models select Qwen/Qwen3-0.6B --wait

--wait returns when the weights are cached. Without it, the run waits for the weights. Selecting an already selected model succeeds, so you can repeat the walkthrough.

The run

The run descriptor specifies the bundle as its reward:

{
  "kind": "run",
  "workload": "gsm8k",
  "configName": "grpo-100-steps",
  "baseModel": "Qwen/Qwen3-0.6B",
  "config": {
    "dataset": "gsm8k-train",
    "method": "grpo",
    "rlHyperparams": {
      "epochs": 1,
      "maxSteps": 100,
      "groupSize": 8,
      "klBeta": 0.001,
      "temperature": 1.0,
      "maxCompletionLength": 512,
      "lora": {
        "rank": 32
      },
      "optimizer": {
        "learningRate": 1e-5
      }
    },
    "reward": {
      "terms": [
        {"kind": "rule", "weight": 1.0, "bundle": "gsm8k-scoring", "module": "reward"}
      ]
    },
    "maxSeqLength": 1536
  },
  "execution": {"gpus": 1, "saveEvery": 10, "keepAdapters": 4}
}

The base model must be one that your deployment allows and that someone has selected; see Base models. If the deployment doesn’t allow Qwen/Qwen3-0.6B, replace the model in run.json. The results will differ.

The reward term specifies the bundle and module. By convention, the function is compute_score. The remaining settings follow published GRPO recipes for this task with a LoRA adapter. The learning rate is an order of magnitude higher than the full-weight baseline because only the adapter changes. Reports show that rank 32 lets a 0.5B model converge like full fine-tuning. klBeta is near zero because the reference policy is the model’s starting point, which offers little useful behavior to preserve for this task. At eight prompts per step, 100 steps make four passes over a 200-row subset. This short run demonstrates a change in the reward rather than producing a useful model.

Start the run:

akka-optimize training runs start -f run.json --dry-run
RUN=$(akka-optimize training runs start -f run.json -o json --jq .runId)
akka-optimize training runs wait "$RUN" --exit-status

What to look at first

Check zeroVarianceGroups before the reward. training runs metrics RUN_ID reports the proportion of groups in a step whose completions all received the same score. Those groups contribute no gradient. A run that remains near 1.0 isn’t learning, even if the reward curve doesn’t show the problem. The proportion should fall in the first few steps as the format term begins to separate completions. The reward should then start to increase.

training runs tensorboard RUN_ID shows the same curves with the per-term contributions.

Evaluate

To evaluate the model, run:

MODEL=$(akka-optimize training runs get "$RUN" -o json --jq .modelId)
EVALUATION=$(akka-optimize training evaluations start --model "$MODEL" --dataset gsm8k-eval \
  --grader bundle:gsm8k-scoring/reward:compute_score -o json --jq .evaluationId)
akka-optimize training evaluations wait "$EVALUATION" --exit-status
akka-optimize training evaluations report "$EVALUATION"

The evaluation specifies the bundle grader explicitly. The trainer serves it in a separate job for the duration of the evaluation, then stops the job. The report’s mean score is the proportion of held-out problems answered with the correct final number.