← CasesML platform
An ML platform on Kubeflow for 1000+ users
Dozens of data science and analytics teams, thousands of employees with access to ML. Instead of a DevBox per person and several team-level platforms, we built one platform where teams share GPUs.
Numbers
Before
- DevBox: every data scientist had their own GPU machine and SSH access. The best platform for one person and the worst for the company: GPUs sit idle and nobody can see who uses what.
- Team-level platforms: one team builds on Kubeflow, another on Airflow. People have to keep both tools in their heads, onboarding drags on, teams race each other to buy hardware.
- Pipelines fail silently: users cannot see why and go to the on-call engineers.
After
- One entry point for ~1000 monthly active users: JupyterLab, Pipelines, Model Registry and distributed training.
- GPUs are shared through Kueue: utilization went from about 15% to 60%.
- ~70k pipeline runs and up to 3k distributed training jobs per month.
- Kubeflow Pipelines patches cut the Kubernetes resources spent on pipelines by half, pod errors show up in the UI.
- A single Istio Sidecar setting saved about a terabyte of RAM across the nodes.
What we did
One entry point. Kubeflow Central Dashboard with JupyterLab, Pipelines, Model Registry, reports and distributed training. The platform is wired into the company context: LDAP/SSO, Vault, observability, S3, Trino, the data lake.
Teams and access through YAML and pull requests. A Go plugin for Kubeflow profiles: a team profile is a namespace with personal user namespaces inside, people come from LDAP groups.
GPU sharing on Kueue instead of a custom scheduler. Flavors per hardware type (CPU, A100, H100, HGX H100, MIG slices), a cohort with guarantees for each team, prod and regular queues. A regular job can borrow an idle GPU from another team, and when the owner needs it back, the job is preempted with a checkpoint. Topology-aware scheduling over InfiniBand for HGX H100. All declarative, in GitOps.
JupyterLab with remote kernels. Each cell starts a kernel in the cluster. After about an hour of idling, the resources and the GPU are released. Vault credentials and database connections are available right away.
Improved Kubeflow Pipelines and sent the changes upstream. Pod errors are visible in the UI, there are metrics and SLAs, fewer MySQL queries, and a central driver replaced the pair of driver pods per run.
Plumbing for thousands of namespaces. Istio Sidecar trims the Envoy config with 3000+ namespaces, secrets arrive through Vault Injection, distributed training runs through Kubeflow Trainer. Our own Model Registry on S3 and MySQL and Model Delivery: press "Deliver" and the model is swapped in the running inference service without a restart.
Stack
KubeflowKubeflow PipelinesArgo WorkflowsKueueKubeflow TrainerJupyterLabIstioVaultLDAP/SSOMySQLS3Grafana
My role
MLOps engineer on the ML platform team. My area: GPU utilization, Kueue for batch jobs, Vault integration.