Endpoints

An endpoint makes a model available for inference requests. Create one to test a selected base model or candidate model with your own requests.

Endpoints are for model testing rather than production traffic.

akka-optimize training endpoints create --model MODEL_ID --workload triage
akka-optimize training endpoints get ENDPOINT_ID
akka-optimize training endpoints list
akka-optimize training endpoints send ENDPOINT_ID --input TEXT
akka-optimize training endpoints proxy ENDPOINT_ID
akka-optimize training endpoints delete ENDPOINT_ID

For --model, specify a selected base model or candidate model. For --workload, specify the endpoint’s workload. The endpoint’s devices count against the workload’s GPU quota, and the trainer rejects a request that would exceed it. --workload defaults to the workload recorded by the model. A candidate needs this flag only to confirm its training workload, and the trainer rejects a different workload. A base model has no workload, so you must pass --workload to serve one.

For ENDPOINT_ID, specify the complete endpoint ID or a unique prefix. endpoints list shows the beginning of each ID.

endpoints create returns immediately, while the endpoint loads the model. Add --wait to return when it responds. The wait ends after 15 minutes. To wait longer, use --timeout.

endpoints delete stops the entire endpoint. The trainer rejects this command when the endpoint serves more than one model because other people might be using those models. To remove one of several models, see Sharing an endpoint.

A training run can start an endpoint to serve its judge. endpoints get shows the run in the Held by row. The trainer refuses to stop the endpoint while the run is calling it. The run releases the endpoint when it finishes.

An endpoint reports one of the following phases:

Phase Meaning

AWAITING_CAPACITY

The endpoint is queued, waiting for devices to free up.

STARTING

The endpoint has its devices and is loading the model.

READY

The model is loaded and the endpoint accepts requests.

FAILED

The endpoint couldn’t start. To see why, run endpoints get.

TERMINATED

The endpoint has stopped and released its devices.

Send one input

Send one input to a model:

akka-optimize training endpoints send ENDPOINT_ID --input "Payment failed twice, card declined."

The command prints the answer as the model produces it. The training service authenticates the request, so you don’t need separate credentials.

Use --file to read the input from a file, or --file - to read standard input.

Use --instructions to tell the model what to do with the input:

akka-optimize training endpoints send ENDPOINT_ID --file ticket.txt \
  --instructions "Reply with one label." --temperature 0

If the endpoint serves several models, select one with --model. Otherwise, the command uses the endpoint’s model.

With -o json, the command returns the model, output, and token usage instead of streaming text.

Send requests from an application

endpoints send sends one input and retains no history. To use a model from an application or test script, open a local address that forwards to the endpoint:

akka-optimize training endpoints proxy ENDPOINT_ID

An endpoint responds inside the deployment, where your machine can’t reach it directly. The proxy forwards requests through the training service, which authenticates them like other commands. The command prints its local address and continues forwarding until you stop it with Ctrl-C. Use --port to choose the port.

The address is OpenAI-compatible. Configure an existing client to use it, or send requests directly:

curl -s ADDRESS/chat/completions -H 'content-type: application/json' \
  -d '{"model":"MODEL_NAME","messages":[{"role":"user","content":"Hello"}]}'

Replace the following:

  • ADDRESS: the address the proxy prints, which ends in /v1

  • MODEL_NAME: the name endpoints get shows beside the model

The proxy forwards output as it arrives, so a request that sets "stream": true produces tokens as the model writes them.

Device usage

An endpoint reserves its devices from the time it starts until you delete it, even when it isn’t handling requests.

Endpoints, runs, and evaluations use the same devices. A running endpoint therefore leaves fewer devices for training and evaluation. Delete the endpoint when you finish testing.

Sharing an endpoint

Endpoints for the same base model at the same commit share one process and one copy of the base model’s weights. The process loads a candidate’s adapter when a request first uses that candidate. Eight candidates trained from one base model therefore share one endpoint and one set of devices.

endpoints create manages this sharing. If no endpoint serves the base model, the new endpoint receives its own devices and enters the STARTING phase. If an endpoint already serves the base model, the command adds your model to that endpoint without using more devices. The command does not reuse a deleted endpoint, even if you create an endpoint immediately after the deletion.

Each request specifies a model. A base model uses the repository as its name. A candidate uses the repository and adapter name. Run endpoints get to find the value to send.

The deployment determines how many adapters one process can load.

To stop serving one model and leave the rest of the endpoint running:

akka-optimize training endpoints detach ENDPOINT_ID MODEL_ID

Specify the model ID or a unique prefix. endpoints get shows the ID beside each served model. Use the model ID here, rather than the model name used in requests.

Detaching one model does not affect the others. Detaching the last one stops the endpoint and releases its devices.