Support triage: narrow classification graded by exact match

This example classifies a retail banking support message as one of 77 intents. Use it to try a narrow, high-volume task that can be checked without a human. The gold answer is a fixed label, so exact matching determines correctness without using one model to grade another.

The example directory contains the data preparation script, pipeline descriptor, base model descriptor, and walkthrough.sh. The walkthrough runs each step on this page against the trainer at TRAINER_URL. To download the directory, see the support triage archive.

Before you begin

  • A trainer deployed in your Akka project, with its address in TRAINER_URL. akka service get trainer shows the route hostname. See Akka Optimize CLI.

  • An authenticated CLI session. Run akka auth login to authenticate akka.

  • Python 3. The data script uses the standard library only.

  • Network access to datasets-server.huggingface.co, where the data script reads the rows. The API limits the request rate, so fetching a large ROWS value takes several minutes. The script caches rows, so subsequent runs use the cached data.

  • A base model that your trainer allows. Run akka training base-models list to list the models, and akka training base-models select <repository> to select one. See Base models.

The walkthrough reads ROWS and EVAL_ROWS from the environment. Start with ROWS=500 EVAL_ROWS=100. This data is enough for a 1B model to learn the task, and training takes a few minutes on one GPU after the run receives a device. The defaults are 2,000 and 400.

Why this task

The task has the following properties:

Property How the task meets it

Gradeable without a judge

77 fixed labels, exact match.

Narrow and repetitive

One intent per message, no free text.

Small output

{"intent": "card_arrival"} rather than prose.

Openly available

mteb/banking77, MIT license.

The data

prepare_dataset.py fetches mteb/banking77 from the Hugging Face datasets-server API. It uses only the Python standard library. The script writes chat-format JSONL with the customer’s message as the user turn and a one-field JSON label as the assistant turn.

To prepare the data, run:

python3 prepare_dataset.py --rows "$ROWS" --eval-rows "$EVAL_ROWS"

The script writes three files:

File Purpose

data/train.jsonl

The rows that the model trains on.

data/eval-tuned.jsonl

Held-out rows under the same short system prompt that the training data uses. The tuned model is scored on this file.

data/eval-frontier.jsonl

The same rows and gold labels under a system prompt that lists all 77 intents. An untuned model needs this prompt on every request to answer at all. The baseline is scored on this file.

Both evaluation files contain identical rows and gold labels. Only the model and its prompt differ in the comparison. A tuned model avoids sending the approximately 1,900 extra prompt characters in every request.

The number of rows matters more than the number of epochs. With 500 rows across 77 intents, each class has about six examples. A classifier loses quality faster from too few examples per class than from fewer passes over those examples. To improve the result, increase ROWS before increasing epochs in the descriptor.

Register the datasets and the workload

To register the datasets, run:

akka-optimize datasets create -f data/train.jsonl --name support-triage-train
akka-optimize datasets create -f data/eval-tuned.jsonl --name support-triage-eval
akka-optimize datasets create -f data/eval-frontier.jsonl --name support-triage-eval-frontier

To create the workload, run:

akka-optimize training workloads get support-triage >/dev/null 2>&1 ||
  akka-optimize training workloads create support-triage
akka-optimize training workloads set-evaluation support-triage --dataset support-triage-eval --grader json-label-match

The workload uses the tuned prompt file as its evaluation dataset and the default json-label-match grader. The trainer rejects creation when the workload name already exists, so the walkthrough creates the workload only when it is missing. It then sets the evaluation settings every time. As a result, set-evaluation applies this dataset and grader to a workload left in any state by an earlier pass.

Before the first run, select the base model for the pipeline. A run can specify only a model that the deployment allows and that someone has selected:

akka-optimize training base-models select unsloth/Llama-3.2-1B-Instruct --wait

--wait returns when the weights are cached. Without it, the first run waits for the weights. Selecting an already selected model succeeds, so you can repeat the walkthrough.

Train and score

The descriptor defines a two-stage pipeline: a supervised run followed by an evaluation of its model.

{
  "kind": "pipeline",
  "stages": [
    {
      "type": "train",
      "name": "sft",
      "configName": "sft-r8-e2",
      "workload": "support-triage",
      "baseModel": "unsloth/Llama-3.2-1B-Instruct",
      "config": {
        "dataset": "support-triage-train",
        "hyperparams": {"lora": {"rank": 8}, "optimizer": {"learningRate": 0.0001}, "training": {"epochs": 2}},
        "maxSeqLength": 2048
      },
      "execution": {"gpus": 1}
    },
    {
      "type": "evaluate",
      "name": "scored",
      "model": "@sft",
      "dataset": "support-triage-eval",
      "maxCompletionTokens": 32
    }
  ]
}

The base model must be one that your deployment allows and that someone has selected; see Base models. If the deployment doesn’t allow unsloth/Llama-3.2-1B-Instruct, replace the model in pipeline.json and base-model.json. The results will differ. maxCompletionTokens limits the length of each scored answer. A label is short, but a model that hasn’t learned the format can otherwise continue generating.

To start the pipeline, run:

PIPELINE=$(akka-optimize training pipelines start -f pipeline.json -o json --jq .pipelineId)
akka-optimize training pipelines wait "$PIPELINE" --exit-status
akka-optimize training pipelines get "$PIPELINE"

pipelines get reports each stage, its run or evaluation, and the evaluation score.

The baseline

Register the base model to score the untuned model on the same rows with the complete taxonomy in the prompt. The descriptor pins a repository commit. Without a pinned commit, registration resolves the head of the repository at that time:

{
  "servingModel": {
    "variant": "BASE",
    "baseModel": "unsloth/Llama-3.2-1B-Instruct",
    "revision": "5a8abab4a5d6f164389b1079fb721cfab8d7126c"
  }
}

To register the baseline and score it, run:

BASE=$(akka-optimize trained-models register -f base-model.json -o json --jq .modelId)
BASELINE=$(akka-optimize training evaluations start --model "$BASE" \
  --dataset support-triage-eval-frontier --workload support-triage -o json --jq .evaluationId)
akka-optimize training evaluations wait "$BASELINE" --exit-status

Read the result

To read the scores, run:

akka-optimize training workloads scores support-triage
akka-optimize training evaluations report "$BASELINE"

workloads scores tabulates the tuned run under the workload’s evaluation settings. The baseline uses a different dataset, the complete prompt file, so read its evaluation report separately. Compare the two mean scores. Both models answer the same rows against the same gold labels. The tuned model stores the taxonomy in its weights, while the baseline receives it in every request.

Base model size affects the result more than rows or epochs. At 1B parameters, a short run learns the task. At 135M parameters, a short run completes the process but learns almost nothing.