Production ML platform roadmap The path from the first GPU in a cluster to a platform where teams train and deploy models on their own. For each step: what problem shows up, how to solve it, and where I wrote or talked about it.
What level is your team at? Check the MLOps maturity levels →
DAY 0 foundation DAY 1 launch DAY 2 production MATURITY 1 → 5 DAY 0 · FOUNDATION Starter kit without a GPU
No GPUs or LLMs yet? You still need a platform. For a small team five things are enough:
Git with reviews and basic CI one experiment tracker: MLflow or ClearML an orchestrator: Airflow or Prefect a model service in a container with one way to roll out and roll back basic monitoring with room for model quality The goal is simple: experiments are reproducible, deployment is repeatable, prod is observable. Feature store, a registry as a product and templates come later.
01 GPU in Kubernetes Problem: The cluster has GPUs, but pods can't see them or crash on an incompatible CUDA version.
the same driver and GPU Operator version on every node a driver / CUDA / framework compatibility matrix nodes and the cluster described as code 02 Choosing hardware for the model Problem: The model didn't fit on the card, or it did, but there was no room left for the KV cache.
estimate memory for weights and KV cache before you buy pick the precision for your hardware: bf16, FP8, int4 look at memory bandwidth, not just VRAM DAY 1 · LAUNCH 03 First serving Problem: The model lives in a data scientist's notebook and gets rolled out to prod by hand.
one inference server: Triton or vLLM the model as an artifact in a registry, not a file on disk deploy through CI, not over ssh 04 Self-service for ML teams Problem: Every notebook and pipeline starts with a ticket to DevOps.
JupyterHub with GPU profiles pipelines and training in Kubeflow per-team quotas instead of handing out cards manually DAY 2 · PRODUCTION 05 Cold start Problem: A new replica takes minutes to come up: images and weights weigh tens of gigabytes.
lazy image loading: eStargz, SOCI, Nydus weights separate from the image: S3, OCI volumes fast weight loading into the GPU 06 Inference autoscaling Problem: CPU-based HPA doesn't see GPU load, and traffic spikes arrive faster than replicas start.
scale on queue length and concurrency KEDA or Knative instead of plain HPA keep a minimum of warm replicas 07 GPU sharing Problem: A small model got a whole A100, and the card sits idle.
MIG where you need isolation time-slicing and MPS for dev and light workloads HAMi as an alternative without MIG 08 Utilization and scheduling Problem: GPUs are allocated but not busy, while teams wait in a queue.
measure real utilization, not allocation quotas and priorities in Kueue preemption and fair share between teams 09 The whole platform Problem: Every team builds its own little MLOps for training and inference, and the solutions drift apart.
a shared horizontal layer for training and inference, while models stay with the product teams the platform as an internal product: a golden path, deployment templates and standard pipelines SLO-based observability, not just "the pod is alive", driver and model updates without downtime don't build it "like big tech" until you have that scale 10 Safe model rollout Problem: A new model goes to everyone at once, and when the metrics drop it gets rolled back by hand.
canary or blue-green: a small share of traffic first A/B on a share of users with business metrics, not just accuracy automatic rollback when the metrics drop 11 Monitor quality, not just uptime Problem: The service is alive and answers 200 OK, while the model quietly degrades.
input data drift and predictions versus reality business metrics and alerts next to RPS and latency logs and traces to find the cause, not just the symptom MLOps maturity levels Day 0/1/2 is the order in which you build infrastructure. Maturity is about how the team works with models. Check what you already have, and you'll see your level and what to close next. The criteria are my own take on the public maturity models from Microsoft Azure and Google Cloud . Without an honest assessment it's easy to optimize the wrong thing: at levels 1-2 it's too early to think about GPU sharing.
1 Manualstarting point
Unstructured notebooks, experiments aren't recorded anywhere, data is prepared by hand, the model is deployed by hand or not at all. No monitoring. Almost everyone starts here.
2 RepeatableCode is in Git, training runs from a script, there is CI with tests. Results can be repeated, but a lot depends on people.
3 ReproducibleA shared tracker and registry, automated training, versioned data. Any model in prod can be rebuilt.
4 AutomatedRollout to a share of traffic with automatic rollback, quality monitoring with alerts, CI/CD checks the model before prod.
5 Self-improvingThe cycle from data to monitoring runs without manual steps: drift triggers retraining by itself, teams have templates.
Your team
Level 1 · Manual Level 2 · Repeatable Level 3 · Reproducible Level 4 · Automated Level 5 · Self-improving What to close next to reach level 2 · Repeatable
to reach level 3 · Reproducible
to reach level 4 · Automated
to reach level 5 · Self-improving
Versioning All code lives in Git, changes go through reviewGit Experiments Training runs from a script with parameters, not from notebook cellsExperiment tracking Envs Orchestration The steps from data to model are described in code and run with one commandML pipelines Data Data preparation is a script, not a manual exportData eng Deployment The model is packaged in a container and rolled out by a written runbookstep 03 · First serving docker Model serving Monitoring There are basic alerts: the model service is alive and respondingMonitoring CI/CD CI runs a linter and tests on every merge requestCI/CD Testing Versioning Data and models are versioned: for every model you can see which data version it was trained onVersioning Experiments All runs are in a shared tracker, the best models are in a registry with a link to their runstep 04 · Self-service for ML teams Experiment tracking Orchestration Training runs as an automated pipeline on shared infrastructure, not on one data scientist's machinestep 04 · Self-service for ML teams ML pipelines DS workspace Data There is a data catalog and shared feature tables, data is checked before trainingData eng Data validation Deployment Inference takes the model only from the registry, never from a file on diskstep 03 · First serving Model serving Experiment tracking Monitoring Besides uptime you can see the model's own metrics: quality and driftstep 11 · Monitor quality, not just uptime ML monitoring CI/CD The pipeline builds the image and deploys it with no manual stepsstep 03 · First serving CI/CD docker Versioning From a model in prod you can trace back to the code, data and experiment it came fromVersioning Experiment tracking Experiments A new model is automatically compared with the current one on the same test setsModel metrics Orchestration Retraining starts on a schedule or on an event, without a personML pipelines Data Data validation stops the pipeline and sends an alertData validation Deployment A new version goes to a share of traffic (canary, blue-green, A/B) and rolls back automaticallystep 10 · Safe model rollout step 06 · Inference autoscaling Rollout & A/B Monitoring Business metrics and drift have alerts next to RPS and latencystep 11 · Monitor quality, not just uptime ML monitoring Monitoring CI/CD CI/CD works with the registry and doesn't let a model that failed the checks into prodstep 10 · Safe model rollout CI/CD Model metrics Versioning The platform collects versions, lineage and metadata by itself, the team doesn't have to rememberstep 09 · The whole platform Versioning ML platform Experiments A new task starts from a ready template, not from a blank pagestep 09 · The whole platform Experiment tracking ML platform Orchestration Retraining is triggered by monitoring when it sees drift, not by a person with a calendarstep 11 · Monitor quality, not just uptime step 08 · Utilization and scheduling ML pipelines ML monitoring Data The path from raw data to features has no manual steps and a check at every stageData validation Data eng Deployment Rollout and rollback are the same for all teams (golden path), with in-house tools or patched open source where neededstep 09 · The whole platform Rollout & A/B ML platform Open/InnerSource Monitoring Drift gets an automatic reaction (retraining or rollback), degradation is traced to its root causestep 11 · Monitor quality, not just uptime ML monitoring CI/CD Teams take standard pipeline templates, nobody writes CI from scratchstep 09 · The whole platform CI/CD ML platform Every criterion is checked. Now it's about keeping it that way :)
Reset checkmarks A level counts only when all its criteria and all the ones below are checked: a hole at level 2 won't let you claim level 4. Checkmarks are stored only in your browser.