0.4.0

This release adds H100 support, the full set of training parameters on the run plan, and more freedom in which base model a deployment trains.

Headlines

H100 support

A run can name the device it trains on and the host memory it needs. execution.deviceProfile names an accelerator the catalog declares, and execution.hostMemoryGiB states what the run’s loader needs from the node. If a plan specifies more than the device has, the trainer refuses it at submission and names the devices that qualify, rather than failing in the loader.

The following example names an H100 and 96 GiB of host memory:

{"execution": {"gpus": 1, "deviceProfile": "H100", "hostMemoryGiB": 96}}

A deployment with one kind of device leaves deviceProfile out and uses the catalog’s default.

The deployment manifests are one set that a region fills in. An H100 sandbox region ships with them, sized for a single H100 80 GiB per node. The H100 envelope in the catalog is an estimate until the first measured run on that hardware.

For more information, see Execution settings.

Training parameters

A run plan states every training parameter. hyperparams has four groups: lora for the adapter’s shape, optimizer for what performs the update, scheduler for how the learning rate moves, and training for the loop. rlHyperparams takes the same first three groups beside its own loop fields. model says how the base model is loaded: precision, gradient checkpointing, and the attention kernel.

Every field is part of the configuration hash, so two runs that trained the same way hash alike, and a run’s record says what it trained under. State as much or as little of a group as you want. Each field you leave out takes the default for the loop that runs, so stating one field of optimizer under a reinforcement run does not change the rest to supervised values.

If the runtime does not honor a field, the trainer refuses the plan at submission and names the field and the ceiling. A misspelled field is refused the same way rather than ignored.

The following plan states one or more fields in each group:

{
  "kind": "run",
  "workload": "triage",
  "baseModel": "google/gemma-4-E4B-it",
  "config": {
    "dataset": "triage-train",
    "method": "sft",
    "hyperparams": {
      "lora": {"rank": 16},
      "optimizer": {"learningRate": 0.0001, "weightDecay": 0.05},
      "scheduler": {"kind": "cosine", "warmupRatio": 0.05, "minLearningRateRatio": 0.1},
      "training": {"epochs": 4, "microBatchPerGpu": 2, "gradientAccumulation": 4}
    },
    "maxSeqLength": 2048,
    "model": {"precision": "bf16"}
  },
  "execution": {"gpus": 1, "deviceProfile": "H100"}
}

For every field, its default, and what it is refused for, see Configuration.

Base model flexibility

A deployment allows a catalog of base models. Selecting an allowed model registers it and warms its weights, and a registered model trains or serves without further setup.

The Gemma 4, Gemma 3, and Ministral families are supported. Gemma 4 loads as the multimodal model it is, and the adapter covers its language model’s layers. The run’s result names the modules it adapted.

base-models candidates walks a publisher’s repositories on the hub and reports the catalog’s verdict on each. With --catalog, it prints the entries to paste into the deployment’s catalog:

akka-optimize training base-models candidates --author google --match 'gemma-4-*'
akka-optimize training base-models candidates --author google --match 'gemma-4-E4B*' --catalog

base-models list shows what is selected and usable, and --all shows everything the deployment allows. select --wait and base-models wait wait for a selection to become ready and report whether its weights are cached:

akka-optimize training base-models list --all
akka-optimize training base-models select google/gemma-4-E4B-it --wait

A deployment with a shared weights cache fetches a base model’s weights once, when you select it, and every run and endpoint reads them from the shared volume. On such a deployment, a selection is not READY until the weights have arrived.

For more information, see Base models.

Other changes

Conversations and tool calls as training data

A training row is a whole conversation: system, user, assistant, and tool messages, ending with an assistant message. An assistant message can carry tool_calls, a tool message answers one, and a row can declare the tools it can call. A single-turn row registers as before, and a system message is optional.

The plan states which tokens the loss covers. loss holds assistantTurns (all or last), trainOnToolCalls, and trainOnToolResults. data names a validation dataset, whose loss and perplexity join the run’s metrics as it trains.

The following configuration trains on every assistant turn, including tool calls, and validates on a second dataset:

{
  "config": {
    "dataset": "triage-tools",
    "data": {"validationDataset": "triage-tools-val"},
    "loss": {"assistantTurns": "all", "trainOnToolCalls": true}
  }
}

For more information, see Conversations and tool calls and Data.

Endpoints reachable from your machine

An endpoint serves a selected base model or a trained candidate for inference. Creating one makes the model resident and ready to answer requests. A reinforcement run’s judge is served the same way, through its own endpoint.

An endpoint is reachable from outside its cluster. endpoints send sends one input and prints the answer. endpoints proxy opens a local OpenAI-compatible address for an application or a test script. endpoints create --wait and endpoints wait follow an endpoint until it answers, reporting its queue position and what the model is doing while it starts:

akka-optimize training endpoints create --model triage-candidate --wait
akka-optimize training endpoints send 970076e8 --input "Payment failed twice, card declined."
akka-optimize training endpoints proxy 970076e8 --port 8800

If a run’s adapter rank exceeds what the deployment’s serving engine can load, the trainer refuses the run at submission, so a run cannot spend hours producing an adapter no endpoint can serve.

For more information, see Endpoints.

Capacity and waiting

training capacity reports what the deployment can schedule and what holds it: each GPU pool’s declared and observed devices, every run, evaluation, or endpoint holding or waiting for devices, and each workload’s usage against its quota:

akka-optimize training capacity
akka-optimize training capacity --workload triage

Every --wait shows progress: a spinner and elapsed time when interactive, and one line per milestone otherwise. A queued run reports its position and why it waits, such as the workload’s quota or earlier queued work.

Failure reporting

A failed run’s record carries one sentence you can act on, such as the reason the runtime refused an example, or that the job ran out of memory and a smaller micro batch or sequence length uses less. Job output and exit detail are kept for operators. training info reports whether the service started completely and exits non-zero until it has.

For more information, see Why a run failed.

Parameter sweeps and run lineage

A run can branch from a snapshot of another run instead of starting from the base model, so exploring a change costs the additional steps trained, not a full run. training runs sweep submits one branch per point of a parameter grid from a single source snapshot, and --compare scores the source against the whole fan once every branch completes. training runs tree walks every generation below a run, and training runs ancestry walks back to the run a lineage started from.

The following sweeps a learning rate across three values and compares the results:

akka-optimize training runs sweep 970076e8 --snapshot step-4000 --steps 200 \
  --grid learning-rate.json --compare --dataset triage-val
akka-optimize training runs tree 970076e8

For more information about branching a run, see Branches.

Smaller changes

  • An endpoint or evaluation over a candidate takes the candidate’s workload when the request omits one. A base model, which belongs to no workload, still requires one.

  • Base models, endpoints, and comparisons accept an unambiguous ID prefix wherever a command names one. comparisons list recovers a comparison ID without repeating its submission.

  • Submitting a plan with no workload returns an error naming the missing field, not the reference grammar.

  • A completed run’s finalLoss is the last loss the training loop logged. The mean over the run is reported beside it as meanLoss.

  • An SFT smoke check stops only when reproduction over the whole training slice reaches its threshold, and reports UNSTABLE_READING when an early stop measures below it.

  • The CLI mints a development token for a trainer at http://localhost:9001 without reading a saved login, so a logged-in developer needs no extra steps to use a local trainer.

  • Training jobs on AWS exit after writing their result, and workload identity on EKS reaches S3.

Breaking changes

  • Plan shape. hyperparams.loraRank is hyperparams.lora.rank, hyperparams.learningRate is hyperparams.optimizer.learningRate, precision is model.precision, and attentionImplementation is model.attentionImplementation under config. rlHyperparams follows the same shape. Existing run and pipeline IDs change, and a stored run in the old shape does not replay.

  • Unknown fields are refused. A field no record of the plan declares is a 400 naming the field. Update any plan that still carries an old name.

  • Selecting a base model fetches its weights on a deployment with a shared weights cache. select returns when the weights are cached. In scripts, use --wait or base-models wait.

  • Static serving applications are gone. A model becomes resident when you select it and create an endpoint. A test target that reached a served model directly needs the address of a running endpoints proxy.

  • A judged reinforcement run needs quota for two claims, its training devices and its judge endpoint. Admission refuses a run whose workload quota holds only one.

  • runs get no longer prints a Logs row. Job output is an operator’s, behind separate commands.

Not in this release

The following are outside what a run plan accepts in this release:

  • Supervised fine-tuning runs on TRL. A plan cannot specify veRL’s SFT trainer or an FSDP2 engine for a supervised run, and the plan has no distributed group for sharding, offload, or sequence parallelism.

  • A completed run registers its adapter. A plan has no field for a second artifact with the adapter merged into the base model.

  • optimizer accepts kind, learningRate, weightDecay, and maxGradNorm. It has no betas or epsilon.

  • The chat template comes from the tokenizer. A plan cannot specify a different template.

  • A multi-node supervised run. A supervised run holds every device it claims on one node.

Unmeasured envelopes

The H100 and Gemma catalog envelopes are estimates. The Gemma 4 load path and the conversation masking are verified against the model configurations and the runtime’s source, not by a run. The first smoke run over Gemma 4 E4B on the H100 sandbox is what measures them.