← Journal · · Deep dive

GPU sharing like Yandex, but open source: Kueue

Yandex recently wrote about their Dev Cluster, and suspiciously many channels posted about it too (for example one and two): a homegrown system for sharing GPUs between ML engineers, grab a card when you need it, give it back when you don't, idle hardware goes to others. It's a cool solution, and the developers did a great job, but I wanted to try something like it myself, not just envy and admire it from afar.

We built a similar mechanism in Kubeflow on vanilla Kubernetes + Kueue. Here are the entities it's made of:

ResourceFlavor: hardware types. Each flavor is tied to nodes via nodeLabels/tolerations: standart (CPU), a100-80g, a100-40g, h100, hgx-h100, plus MIG slices (a100-80-10/20/40gb, h100-10/20/40gb), a single A100/H100 gets sliced into pieces, and a slice is also a quotable resource.

Cohort (hierarchical): the basis of sharing. A root cohort root, with a cohort for each team underneath it. Team guarantees (nominalQuota) are declared on the team's own cohort, and through the shared parent, teams can borrow each other's idle resources.

ClusterQueue: we generate two queues per team:

This is the equivalent of "give it back when the owner returns": a research job runs on someone else's idle GPUs, but the moment the owner's prod workload shows up, it gets preempted and sent back to the queue.

WorkloadPriorityClass (prod-workload, value 1000): preemption priority within the queue, withinClusterQueue: LowerPriority.

LocalQueue: the entry point for users. Every team/user namespace automatically gets a default and a prod queue (via the namespace-configuration-operator), and the user just specifies the queue name in their job.

Topology (TAS): topology-aware scheduling. Regular flavors get packed by zone/hostname, while for HGX H100 it's packed by InfiniBand scalable unit, so multi-node training doesn't get spread thin across the fabric.

Bottom line: guarantees for teams, plus utilization of idle capacity, plus fair preemption, all without a custom scheduler, all declarative, in GitOps, at 1000 MAU.

Original on Telegram ↗

↑↓ select · Enter open · Esc close