← Journal · · Deep dive

KAI Scheduler: Nvidia's native K8s scheduler

Posts have been coming out often this week, but there's plenty to share! Finally got around to trying out KAI Scheduler. Nvidia open-sourced the scheduler used in their enterprise Run.AI platform. In short - it does a solid job solving the problem of scheduling workloads onto GPU resources for training tasks. How does it do it?

Let's go through the entities:

⌛️ Queues - an entity where you can define resource quotas. Very handy if you want to limit specific workloads to a limited amount of resources. Got a team? Add a label pointing to the queue name. If they exceed the GPU allocation quota, the pod won't start.

📈 Elastic Workload - a workload with built-in HPA. No further explanation really needed.

🚀 Workload Priority - four priority classes. They show which workloads can be preempted to allocate resources for more important processes.

🤝 GPU sharing - you can specify GPU fractions, e.g. 0.5. Yes, K8s won't let you run more than 2 pods per GPU in this case, but it's essentially time slicing, so there's no memory control at the K8s level. I ran 2 vLLM instances with 0.5 fractions each, and each one grabbed the full GPU and caused an OOM. You'll have to cap memory at the application level.

What I liked most is the quota mechanism in queues. The rest applies more to training, not so much to inference. There's also a project called Volcano from a Chinese team, but overall it's roughly the same thing.

Original on Telegram ↗

↑↓ select · Enter open · Esc close