Base models

A training run starts from a base model. Each deployment supports a fixed set of base models. Select a model to prepare it for training, evaluation, or serving.

Allowed and selected models

An allowed model is available for use in this deployment. The deployment’s hardware and training engines must support it. You can’t add models to this list.

A selected model is an allowed model that you have prepared for training, evaluation, or serving. Selecting a model registers it. If the deployment uses a shared cache, selection also downloads the weights in advance. Initially, no models are selected.

List the selected models:

akka-optimize training base-models list

To include every allowed model, add --all:

akka-optimize training base-models list --all

The repository identifies each model and is the value that other base model commands accept. The output also shows the pinned commit, parameter count, context length, and training methods that the hardware supports. STATUS shows the model’s selection state. A selected and prepared model is ready. A selected model whose weights haven’t arrived is preparing. An allowed model that nobody has selected is not-selected.

The allowed list includes only models that the training engines and hardware support. If the list doesn’t contain a model that you need, ask your operator to add it to the deployment.

To see what a run over the model can request, get the model. Under Runs this deployment admits, the output lists every admitted combination of training method, precision, and GPU count. It also shows the memory of each GPU, the host memory of the node that holds them, and the maximum sequence length:

akka-optimize training base-models get Qwen/Qwen3-8B

If a run requests an unsupported configuration, the trainer rejects it during submission before using a GPU. For the validation rules, see Base models.

Find compatible models

Before you ask an operator to allow a model, check whether the deployment can train it. The candidates command lists repositories from a Hugging Face account and evaluates each model against the deployment’s engines and hardware:

akka-optimize training base-models candidates --author google --match 'gemma-4-*'

--match filters repository names by using a glob pattern. In the pattern, * matches any sequence of characters. Omit --match to list all repositories from the account.

In the output, VERDICT is admissible if the deployment can train the model. Otherwise, the column explains why the model isn’t supported. Possible reasons include an unsupported architecture, a model that is too large for the hardware, or a model that the trainer can’t adapt.

The command reads data from Hugging Face but doesn’t modify the deployment. It doesn’t allow or select any models.

To generate catalog entries for the admissible models, an operator can add --catalog:

akka-optimize training base-models candidates --author google --match 'gemma-4-E4B*' --catalog

Select a model

akka-optimize training base-models select Qwen/Qwen3-8B

The command returns immediately while preparation continues in the background. Add --wait to wait until the model is ready. The wait ends after 30 minutes. To wait longer, use --timeout.

While the model is being prepared, you can start a run or evaluation, or create an endpoint. The trainer accepts the request and starts the work when the model is ready. In a script, use --wait to prepare the model before the script continues.

To check the preparation progress, run base-models get.

Model statuses

base-models get reports two statuses. Status shows whether the deployment configuration permits you to use the model.

Status Meaning

PENDING

The trainer is still reading the repository.

AVAILABLE

The training engines and hardware support the model.

UNSUPPORTED

The training engines or hardware do not support the model.

UNAVAILABLE

The trainer could not read the repository. Detail gives the reason.

WITHDRAWN

The model was removed from the allowed list.

Selection shows the model’s preparation progress. It appears only after you select the model.

Selection Meaning

PREPARING

The trainer is registering the model. This usually takes seconds.

FETCHING

The trainer is downloading the weights. This can take several minutes for a large model. Fetched counts the files downloaded so far.

READY

The model is ready for training, evaluation, and serving.

READY means that you can use the model. It doesn’t indicate whether the weights were downloaded, because the model is usable either way. Weights shows the cache state after the selection reaches READY:

Weights Meaning

cached

The weights are downloaded and ready for use.

not cached

This deployment does not download weights in advance. The first run that needs them downloads them before it starts, and takes longer to begin.

not cached: REASON

The download didn’t finish. The first run that needs the weights downloads them before it starts. REASON explains the failure.

base-models select --wait reports the same information when it returns, so a script doesn’t need to read the entry again.

Deselect a model

akka-optimize training base-models deselect Qwen/Qwen3-8B

The model remains allowed, and you can select it again. The trainer retains its registration because models trained from it still refer to it.

Deselecting does not stop an endpoint that is already serving the model. To stop one, see Endpoints.