← Journal · · Deep dive
Can you rename nvidia.com/gpu to the GPU model
Hey everyone! Been a while since I posted (honestly, I was on vacation after speaking at Tech Day)
Now I'm getting back to ML infrastructure research and want to share content with you again
I decided to check whether the Kubernetes resource nvidia.com/gpu (which represents GPU presence and capacity on nodes) can be renamed to a specific GPU, like nvidia.com/A100
I know that for MIG there's a config where you can define resource names for different partitions. Can you make a similar config for regular GPUs?
The resource registration logic lives in this file in the nvidia device plugin, and unfortunately, apart from plain nvidia.com/gpu or MIG-based transformations, you can't change the resource name. Only if you rewrite the code, like this
for i := 0; i < deviceCount; i++ { device, err := nvml.DeviceGetHandleByIndex(i) if err != nil { return fmt.Errorf("failed to get device handle: %v", err) }
model, err := nvml.DeviceGetName(device)
if err != nil {
return fmt.Errorf("failed to get device name: %v", err)
}
// Register resource with the model name
resourceName := "nvidia.com/" + strings.ToLower(model)
\_ = config.Resources.AddGPUResource("\*", resourceName)
}
But then how do you solve the following problem: allocating pods to a specific GPU?
You can use nodeSelector and the nvidia.com/gpu.product label, which sets the GPU name. It looks like this
apiVersion: v1 kind: Pod metadata: name: gpu-pod spec: containers:
- name: gpu-container image: nvidia/cuda:11.0-base resources: limits: nvidia.com/gpu: 1 nodeSelector: nvidia.com/gpu.product: "A100-SXM4-40GB" # Use the label to select the node
That's it! Drop a 🔥 if you're interested in more posts like this about my research. Also thinking about renaming myself to MLOps after all... 🕺 since the content is mostly about ML these days