Allocating hardware to models
An accelerator claims a slice of the device inventory described in Reviewing available hardware. A model deployment is placed on an accelerator, never on a device, so an accelerator has to exist before a model can run.
An accelerator states which device it claims and whether one model gets the card to itself.
Reviewing the accelerators you have
akka models accelerators list
NAME DEVICE TENANCY CARDS/REPLICA MEMORY/MODEL IN USE CAN PLACE READY
standard l4 Dedicated 1 23034Mi 0 1 True
CAN PLACE is the column that answers whether another model will fit. This accelerator holds the region’s one card and has not been deployed to yet, so it has a slot free. READY only says the accelerator itself is healthy.
To read the same information for one accelerator as a single line:
akka models accelerators get standard
|
Creating an accelerator does not check that the hardware behind it is free. An accelerator created against a fully allocated card reports |
Choosing a tenancy
Tenancy decides whether a model gets whole cards or shares one.
| Tenancy | Behavior |
|---|---|
|
Each model gets whole devices. This is the default. |
|
Several models split the memory of one device. |
Dedicated tenancy is the choice for a model that has to hold its latency under load, because nothing else competes for the card’s memory. Set devicesPerReplica above 1 to give one replica several cards. That requires a node holding that many cards, and the deployment’s parallelism factors have to multiply to exactly that number.
Shared tenancy fits several small models onto one card. Every model on a shared accelerator has to state engine.memoryPercent, because a model that does not declare its share is sized against the whole card and runs out of memory once the weights finish downloading. A shared accelerator splits a single device, so it cannot set devicesPerReplica above 1.
Creating an accelerator
Declare the accelerator in the model descriptor when you want it version controlled alongside the models that use it:
resource: Accelerator
resourceVersion: v1
metadata:
name: l4-dedicated
spec:
device: l4
tenancy: dedicated
Applying the descriptor creates it. See Accelerator for every field.
To create one directly instead:
akka models accelerators create fast --device l4
To split one card between two models:
akka models accelerators create small --device l4 --tenancy Shared --max-models 2
An accelerator is immutable once it exists. Re-applying a descriptor leaves an existing accelerator alone rather than reshaping it, so changing a tenancy means deleting the accelerator and creating it again.
Releasing capacity
akka models accelerators delete fast
Deleting an accelerator releases the capacity it held. Delete the deployments placed on it first, otherwise they have nowhere to run.
|
The features described in this section are an add-on to Akka Automated Operations. They are not included in the base product. |