Datasets

A dataset is an immutable set of training or evaluation examples. Its content hash identifies it, and a name is a movable alias for that hash. A run records the resolved hash when it is admitted, so moving the name later doesn’t change its training data.

Use datasets to provide examples for training and evaluation. Keep training and evaluation datasets separate. A workload scores every candidate with the evaluation dataset, so an example that appears in both datasets inflates the score.

Format

A dataset uses chat-format JSONL. Each line contains one JSON object with a messages array of system, user, and assistant messages. The following example shows one row:

{"messages":[
  {"role":"system","content":"Extract the invoice total as JSON."},
  {"role":"user","content":"Invoice 1042, 3 items, total due 1290.00 EUR"},
  {"role":"assistant","content":"{\"total\":1290}"}
]}

Supervised fine-tuning teaches the model to produce the assistant message. A reinforcement run samples its own completions and uses the assistant message as the gold answer when the reward requires one. During an evaluation, the grader compares the model’s answer with the assistant message.

Conversations and tool calls

A row is a whole conversation rather than one exchange. Roles are system, user, assistant, and tool. A row ends with an assistant message. An assistant message may carry tool_calls, and each call is answered by a tool message naming the tool_call_id it answers. A row declares the tools it may call in a row-level tools list, which the chat template renders.

The following example is one three-turn row with a tool call:

{"tools":[{"type":"function","function":{
   "name":"get_weather",
   "description":"Look up the current weather for a city",
   "parameters":{"type":"object","properties":{"city":{"type":"string"}},"required":["city"]}}}],
 "messages":[
  {"role":"system","content":"You are a weather assistant."},
  {"role":"user","content":"What is the weather in Wellington?"},
  {"role":"assistant","content":"","tool_calls":[{"id":"call_1","type":"function",
     "function":{"name":"get_weather","arguments":{"city":"Wellington"}}}]},
  {"role":"tool","tool_call_id":"call_1","name":"get_weather","content":"18C and windy"},
  {"role":"assistant","content":"It is 18C and windy in Wellington."}
]}

A system message is optional. Chat data commonly carries none, and a chat template renders such a conversation without one, so a row without it registers. Inventing a system prompt to satisfy a validator would train the model on a sentence nobody wrote.

Registration refuses a row it cannot render or mask, naming the line and the reason:

  • a row that does not end with an assistant message

  • a row with no user or no assistant message

  • a role other than the four above

  • a system message that is not at the head of the conversation

  • a tool message that follows no assistant message carrying tool_calls

  • a tool message whose tool_call_id the preceding assistant message did not call

  • a tool call that no following tool message answers

  • an assistant message carrying both tool_calls and text

  • an image part outside a user message

The last refusal is the one that is not obvious. A chat template renders a calling assistant message as its calls, then the results that answer them, then its text, so text on a calling message moves once the result arrives. The runtime attributes tokens to messages by rendering the conversation one message at a time and taking each message’s tokens as the growth past the previous render, and that attribution is wrong where the earlier text moves. Put the text in an assistant message of its own before the call.

A dataset’s diagnostics report how many rows are multi-turn, how many tool calls the corpus makes, and how many rows declare tools. None of the three raises a finding. What a multi-turn or tool-calling row is worth is the run’s loss policy to state. See Training runs.

Register a dataset

Register a dataset, then read and list datasets:

akka-optimize datasets create -f train.jsonl --name triage-train
akka-optimize datasets get triage-train
akka-optimize datasets list

create prints the content hash. Registering the same examples again returns the same hash. The trainer accepts an upload of up to 256 MiB by default. A deployment sets the limit with trainer.datasets.max-bytes. --name assigns the alias. If the alias already identifies another dataset, the command moves it. datasets name moves an alias without another upload.

A run descriptor, a workload’s evaluation settings, and an evaluation command accept a dataset name or hash. A command also accepts a unique hash prefix when its help says so.

Images

A user message can carry text and image parts, as in the following example:

{"messages":[
  {"role":"system","content":"Extract the invoice total as JSON."},
  {"role":"user","content":[
    {"type":"text","text":"Read this invoice."},
    {"type":"image","path":"images/invoice-1042.png"}
  ]},
  {"role":"assistant","content":"{\"total\":1290}"}
]}

path is relative to the JSONL file. datasets create reads each image and uploads images that the service doesn’t already have. It replaces each image path with the image’s content hash when it registers the examples. --no-upload reports missing images without uploading or registering anything. Registration rejects an example with a missing image.

Training on images uses a vision model’s processor and adapts its language backbone. The base model must therefore accept images. Graders receive text. Only sft trains on images. The trainer rejects a reinforcement run over a dataset with images.