Production ML platform roadmap

The path from the first GPU in a cluster to a platform where teams train and deploy models on their own. For each step: what problem shows up, how to solve it, and where I wrote or talked about it.

What level is your team at? Check the MLOps maturity levels →

DAY 0 · FOUNDATION

01

GPU in Kubernetes

Problem: The cluster has GPUs, but pods can't see them or crash on an incompatible CUDA version.

  • the same driver and GPU Operator version on every node
  • a driver / CUDA / framework compatibility matrix
  • nodes and the cluster described as code

DAY 1 · LAUNCH

DAY 2 · PRODUCTION

07

GPU sharing

Problem: A small model got a whole A100, and the card sits idle.

  • MIG where you need isolation
  • time-slicing and MPS for dev and light workloads
  • HAMi as an alternative without MIG
09

The whole platform

Problem: Every team builds its own little MLOps for training and inference, and the solutions drift apart.

  • a shared horizontal layer for training and inference, while models stay with the product teams
  • the platform as an internal product: a golden path, deployment templates and standard pipelines
  • SLO-based observability, not just "the pod is alive", driver and model updates without downtime
  • don't build it "like big tech" until you have that scale
10

Safe model rollout

Problem: A new model goes to everyone at once, and when the metrics drop it gets rolled back by hand.

  • canary or blue-green: a small share of traffic first
  • A/B on a share of users with business metrics, not just accuracy
  • automatic rollback when the metrics drop

MLOps maturity levels

Day 0/1/2 is the order in which you build infrastructure. Maturity is about how the team works with models. Check what you already have, and you'll see your level and what to close next. The criteria are my own take on the public maturity models from Microsoft Azure and Google Cloud. Without an honest assessment it's easy to optimize the wrong thing: at levels 1-2 it's too early to think about GPU sharing.

1Manualstarting point

Unstructured notebooks, experiments aren't recorded anywhere, data is prepared by hand, the model is deployed by hand or not at all. No monitoring. Almost everyone starts here.

2Repeatable

Code is in Git, training runs from a script, there is CI with tests. Results can be repeated, but a lot depends on people.

3Reproducible

A shared tracker and registry, automated training, versioned data. Any model in prod can be rebuilt.

4Automated

Rollout to a share of traffic with automatic rollback, quality monitoring with alerts, CI/CD checks the model before prod.

5Self-improving

The cycle from data to monitoring runs without manual steps: drift triggers retraining by itself, teams have templates.

A level counts only when all its criteria and all the ones below are checked: a hole at level 2 won't let you claim level 4. Checkmarks are stored only in your browser.

↑↓ select · Enter open · Esc close