Support triage: narrow classification graded by exact match
This example classifies a retail banking support message as one of 77 intents. Use it to try a narrow, high-volume task that can be checked without a human. The gold answer is a fixed label, so exact matching determines correctness without using one model to grade another.
The example directory contains the data preparation script, pipeline descriptor, base model descriptor, and walkthrough.sh. The walkthrough runs each step on this page against the trainer at TRAINER_URL. To download the directory, see the support triage archive.
Before you begin
-
A trainer deployed in your Akka project, with its address in
TRAINER_URL.akka service get trainershows the route hostname. See Akka Optimize CLI. -
An authenticated CLI session. Run
akka auth loginto authenticateakka. -
Python 3. The data script uses the standard library only.
-
Network access to
datasets-server.huggingface.co, where the data script reads the rows. The API limits the request rate, so fetching a largeROWSvalue takes several minutes. The script caches rows, so subsequent runs use the cached data. -
A base model that your trainer allows. Run
akka training base-models listto list the models, andakka training base-models select <repository>to select one. See Base models.
The walkthrough reads ROWS and EVAL_ROWS from the environment. Start with ROWS=500 EVAL_ROWS=100. This data is enough for a 1B model to learn the task, and training takes a few minutes on one GPU after the run receives a device. The defaults are 2,000 and 400.
Why this task
The task has the following properties:
| Property | How the task meets it |
|---|---|
Gradeable without a judge |
77 fixed labels, exact match. |
Narrow and repetitive |
One intent per message, no free text. |
Small output |
|
Openly available |
|
The data
prepare_dataset.py fetches mteb/banking77 from the Hugging Face datasets-server API. It uses only the Python standard library. The script writes chat-format JSONL with the customer’s message as the user turn and a one-field JSON label as the assistant turn.
To prepare the data, run:
python3 prepare_dataset.py --rows "$ROWS" --eval-rows "$EVAL_ROWS"
The script writes three files:
| File | Purpose |
|---|---|
|
The rows that the model trains on. |
|
Held-out rows under the same short system prompt that the training data uses. The tuned model is scored on this file. |
|
The same rows and gold labels under a system prompt that lists all 77 intents. An untuned model needs this prompt on every request to answer at all. The baseline is scored on this file. |
Both evaluation files contain identical rows and gold labels. Only the model and its prompt differ in the comparison. A tuned model avoids sending the approximately 1,900 extra prompt characters in every request.
The number of rows matters more than the number of epochs. With 500 rows across 77 intents, each class has about six examples. A classifier loses quality faster from too few examples per class than from fewer passes over those examples. To improve the result, increase ROWS before increasing epochs in the descriptor.
Register the datasets and the workload
To register the datasets, run:
akka-optimize datasets create -f data/train.jsonl --name support-triage-train
akka-optimize datasets create -f data/eval-tuned.jsonl --name support-triage-eval
akka-optimize datasets create -f data/eval-frontier.jsonl --name support-triage-eval-frontier
To create the workload, run:
akka-optimize training workloads get support-triage >/dev/null 2>&1 ||
akka-optimize training workloads create support-triage
akka-optimize training workloads set-evaluation support-triage --dataset support-triage-eval --grader json-label-match
The workload uses the tuned prompt file as its evaluation dataset and the default json-label-match grader. The trainer rejects creation when the workload name already exists, so the walkthrough creates the workload only when it is missing. It then sets the evaluation settings every time. As a result, set-evaluation applies this dataset and grader to a workload left in any state by an earlier pass.
Before the first run, select the base model for the pipeline. A run can specify only a model that the deployment allows and that someone has selected:
akka-optimize training base-models select unsloth/Llama-3.2-1B-Instruct --wait
--wait returns when the weights are cached. Without it, the first run waits for the weights. Selecting an already selected model succeeds, so you can repeat the walkthrough.
Train and score
The descriptor defines a two-stage pipeline: a supervised run followed by an evaluation of its model.
{
"kind": "pipeline",
"stages": [
{
"type": "train",
"name": "sft",
"configName": "sft-r8-e2",
"workload": "support-triage",
"baseModel": "unsloth/Llama-3.2-1B-Instruct",
"config": {
"dataset": "support-triage-train",
"hyperparams": {"lora": {"rank": 8}, "optimizer": {"learningRate": 0.0001}, "training": {"epochs": 2}},
"maxSeqLength": 2048
},
"execution": {"gpus": 1}
},
{
"type": "evaluate",
"name": "scored",
"model": "@sft",
"dataset": "support-triage-eval",
"maxCompletionTokens": 32
}
]
}
The base model must be one that your deployment allows and that someone has selected; see Base models. If the deployment doesn’t allow unsloth/Llama-3.2-1B-Instruct, replace the model in pipeline.json and base-model.json. The results will differ. maxCompletionTokens limits the length of each scored answer. A label is short, but a model that hasn’t learned the format can otherwise continue generating.
To start the pipeline, run:
PIPELINE=$(akka-optimize training pipelines start -f pipeline.json -o json --jq .pipelineId)
akka-optimize training pipelines wait "$PIPELINE" --exit-status
akka-optimize training pipelines get "$PIPELINE"
pipelines get reports each stage, its run or evaluation, and the evaluation score.
The baseline
Register the base model to score the untuned model on the same rows with the complete taxonomy in the prompt. The descriptor pins a repository commit. Without a pinned commit, registration resolves the head of the repository at that time:
{
"servingModel": {
"variant": "BASE",
"baseModel": "unsloth/Llama-3.2-1B-Instruct",
"revision": "5a8abab4a5d6f164389b1079fb721cfab8d7126c"
}
}
To register the baseline and score it, run:
BASE=$(akka-optimize trained-models register -f base-model.json -o json --jq .modelId)
BASELINE=$(akka-optimize training evaluations start --model "$BASE" \
--dataset support-triage-eval-frontier --workload support-triage -o json --jq .evaluationId)
akka-optimize training evaluations wait "$BASELINE" --exit-status
Read the result
To read the scores, run:
akka-optimize training workloads scores support-triage
akka-optimize training evaluations report "$BASELINE"
workloads scores tabulates the tuned run under the workload’s evaluation settings. The baseline uses a different dataset, the complete prompt file, so read its evaluation report separately. Compare the two mean scores. Both models answer the same rows against the same gold labels. The tuned model stores the taxonomy in its weights, while the baseline receives it in every request.
Base model size affects the result more than rows or epochs. At 1B parameters, a short run learns the task. At 135M parameters, a short run completes the process but learns almost nothing.