← Journal · · Note

Scheduling AI workloads in k8s

In the previous post, the article Kubernetes goes AI-First: Unpacking the new AI conformance program explicitly mentions the problem of scheduling AI workloads. What's already out there in open source right now?

1️⃣ Reclaiming underutilized GPUs in Kubernetes using scheduler plugins

The native k8s scheduler looks at CPU and RAM, and knows nothing about GPU utilization. Scheduler plugins for k8s can help here: you can set rules that filter available nodes by Prometheus metrics, for example from the DCGM exporter.

Article: https://www.cncf.io/blog/2026/01/20/reclaiming-underutilized-gpus-in-kubernetes-using-scheduler-plugins Repo: https://github.com/kubernetes-sigs/scheduler-plugins

2️⃣ Running AI Workloads on Rack-Scale Supercomputers: From Hardware to Topology-Aware Scheduling NVIDIA recommends moving to topology-aware scheduling approaches, using DRA extensions (a thin API for working with vendor hardware), advanced meta-schedulers (the Run:ai/KAI scheduler), and automatic hardware discovery tools (Topograph).

Article: https://developer.nvidia.com/blog/running-ai-workloads-on-rack-scale-supercomputers-from-hardware-to-topology-aware-scheduling/ Repo: https://github.com/NVIDIA/topograph

Original on Telegram ↗

↑↓ select · Enter open · Esc close