MLOps is made of three tracks: ML, development and operations. You don't need calculus or the transformer architecture. But you do need to understand how a data scientist works: they are your customer, and they need the right tools. The platform for them is assembled from ready-made blocks, like LEGO.
The circles are what I deal with every day, not a mandatory list: everyone's set will be different. Go through all of them, even if you came from a neighboring field, and rate yourself honestly. Your own vertical (backend, DevOps or data) is usually already in "know it", and the rest you bring up to enough: every circle has a hint on that level, go deeper only when a task calls for it. At the end, download the plan: it will include your gaps, tasks and sources.
Where to start
From scratch: networks (you can refresh them in an evening) → Fit/Predict on Titanic → experiment tracking → inference on GPU and LLMs.
From DevOps or sysadmin work: the typical way in is a DS team that can't get a model to prod, and the task lands on you. Close the DevOps part first, that's already a profession. Do Fit/Predict in parallel: don't put ML off, better find out early whether you actually like it. Then DVC, MLflow and the difference between online and batch inference. Leave distributed training for later.
From data engineering: you already understand data, start with Kubernetes and the "Full cycle on MLflow" project.
From backend: you can skip the hardware, start with packaging a model: the "First model without a GPU" project. For you a model is a tricky component with its own demands on latency, memory, updates and A/B. Learn ONNX and feature caching.
From data science: you already have the math and frameworks, but models break on the way to prod. Learn OOP, tests, logging, Docker, basic CI/CD and an API on FastAPI: the docker, API, Testing and CI/CD circles and the "First model without a GPU" project. The goal is to package research into a reproducible artifact.
Where to practice if there are no tasks at work
On a laptop without a GPU: Docker, minikube, MLflow, ONNX Runtime, Triton on CPU. You can try Kubernetes in the browser on Killercoda.
Cheap VMs and the cloud: a couple of VMs from an inexpensive hosting provider, or managed Kubernetes on the bonus clouds give to new accounts.
GPU: rent a VM with a card for an hour or two. For LLMs you can't tell which config is better without a real GPU.
What they ask in interviews
Big tech usually has no separate MLOps track. Infra engineers go through the SRE track: Linux, Kubernetes, System Design. GPUs come up at the final round at most, and there it matters more how you approach problems. ML specifics are asked at mid-size and small companies. System Design is common for senior positions: mock interviews help here.
Specializations: a data platform makes working with data easier, an ML platform speeds up testing hypotheses, an inference platform is responsible for resources and reliability in prod. These are verticals. There is a horizontal too: just as DevOps sits next to developers, MLOps sits next to data scientists and builds the pipeline from data to prod.
In an interview, concepts matter more than specific tools: if the company uses a different tool, they'll retrain you.
A common task: "roll out an experiment tracker". They check your reasoning, not your MLflow knowledge: first find out what problem we're solving, gather the criteria, and only then pick a tool. "Everyone uses it" and "the team next door has it" are not arguments.
What MLOps doesn't do
MLOps by itself doesn't make a model more accurate and doesn't replace A/B tests. It makes testing hypotheses faster and cheaper.
Deep database tuning (indexes, query plans) is a job for a DBA or backend: you spot the problem in monitoring and hand it over. Why a model didn't bring business value is for the product manager, analyst or data scientist to figure out. Your part is prediction logs and metrics that make it visible.
not rateddon't knowhave gapsknow itI have articles or posts on it
End-to-end projects
drag the map with your finger, tap a circle
Skill map
ML
Fit/Predict
Be a data scientist at least a couple of times. You don't need calculus: take ready-made code, fit trains, predict makes predictions. That's how you start to understand your customers.
Enough for MLOps: you've gone through the basic DS process: data, training, hyperparameter tuning, the model as an artifact.
Try it hands-on
Titanic on Kaggle: from data to a model tutorial ↗
download a model from Hugging Face and run inference tutorial ↗
What a neural network is made of, what weights are and why they take up gigabytes.
Enough for MLOps: you understand weights, epochs, batches and the difference between training and inference. You don't need to know the transformer architecture.
Try it hands-on
train a small network in PyTorch and save the weights tutorial ↗
get to know the formats: safetensors, ONNX, GGUF tutorial ↗
The model works, and then the metrics show it's losing quality. That's drift. If the model is retrained weekly, it often doesn't have time to degrade. And before a rollout the new version is compared with the current one, otherwise "it got better" means nothing.
Enough for MLOps: you know the main quality metrics and what drift is. You compare models under the same conditions: the same test sets, several slices, unusual scenarios like a demand spike or weird input.
Try it hands-on
compute precision, recall and ROC-AUC for the model from Fit/Predict tutorial ↗
add a gate to CI: a new model moves on only if it's no worse than the current one on the same test sets tutorial ↗
LLMs are a separate track, not like classic fit/predict. The key questions: will the model fit on the card, which quantization to pick, how to check quality.
Enough for MLOps: you can estimate VRAM: parameters × bytes (70B in bf16 ≈ 140 GB) plus KV cache and activations. You know quantization, fine-tuning, the OpenAI protocol and streaming over SSE.
Try it hands-on
pick a GPU for Qwen3 32B and compare with DeepSeek tutorial ↗
turn daily traffic into RPS: 200,000 requests a day is ≈ 2.5 RPS tutorial ↗
A data scientist keeps running the code with slightly different parameters. It used to be notebooks v1, v2 in a folder. A tracker records the code, parameters, data version and environment next to the metrics, so any run can be repeated. The best model goes to the registry with a link to its run, and inference takes it from there. MLOps deploys all of this for a hundred people with guaranteed resources.
Enough for MLOps: you've set up MLflow or ClearML, logged a couple of experiments yourself and registered a model with its data version, parameters and metrics.
Try it hands-on
set up MLflow and log the training from Fit/Predict tutorial ↗
Model Registry: from a model version to deployment tutorial ↗
In ML you version not just code but data and models too, and link them together. Otherwise you can't say why accuracy dropped from 90% to 85% on the new dataset version. DVC keeps large files in external storage and puts only references to them in Git.
Enough for MLOps: you've versioned a dataset with DVC and can tell which data version the production model was trained on. Don't dive deep until you have a task for it.
Go deeper if needed: lineage: where the data came from and which steps it went through, plus a data catalog. That way you can walk from a production model back to the source table.
ML + DEV
ML pipelines
CI/CD for models. An orchestrator wires the steps into a DAG (cleaning, training, testing) and runs it on a schedule or on an event. Experiments with different hyperparameters run in parallel, and the best model goes to prod. In Kubernetes, Kubeflow Pipelines is the natural choice: every step is a separate container.
Enough for MLOps: you've built a pipeline: several training runs, picking the best model, deployment.
Try it hands-on
a pipeline in ClearML Pipelines or Kubeflow Pipelines tutorial ↗
Go deeper if needed: continuous training: retraining is triggered by monitoring that has spotted drift, not by a person.
ML + DEV
Data eng
ETL, DAGs in orchestrators and Feature Store. Example: a scoring service receives only a client ID, and the Feature Store fetches their history and income by it. It also makes sure a feature is computed the same way in training and in inference. Many "unstable model" issues actually grow out of the data.
Enough for MLOps: you understand ETL, DAGs and why you need a Feature Store (online data in Redis, cold data in S3). It's a hard topic, no need to go deep right away.
Try it hands-on
figure out how a data lake differs from a warehouse tutorial ↗
run Kafka in Docker and push a stream of events through it tutorial ↗
Go deeper if needed: how a data platform is built: sources (databases, logs, Kafka) → a lake on S3 in Parquet → a DWH like ClickHouse or Greenplum → Spark for transformations, with a catalog, lineage and access control on top.
ML + DEV
Data validation
A model learns from whatever it's given. A source breaks: a column goes empty, prices arrive in cents instead of dollars, and the model faithfully trains on garbage. So data is checked before training: schema, ranges, gaps, anomalies. If the checks fail, the pipeline stops and sends an alert.
Enough for MLOps: you've added a data check step to a pipeline that fails the run and sends an alert when the schema or distribution drifts.
Try it hands-on
describe a dataframe schema in Pandera and make the pipeline fail on broken data tutorial ↗
compare two datasets with the Data Drift preset in Evidently tutorial ↗
Where a data scientist lives: Jupyter, JupyterHub, DevBox. In Jupyter every cell keeps state, so DS code doesn't look like a developer's: it's a tool for experiments.
Enough for MLOps: you understand how Jupyter works and have deployed JupyterHub with profiles for different environments.
ML is basically matrix math, and a GPU has thousands of times more cores for it than a CPU. The main limit is video memory. The card's architecture decides what's available: MIG and bf16 since Ampere, FP8 since Hopper, FP4 since Blackwell.
Enough for MLOps: you know the generations, VRAM and teraflops, data center vs gaming cards. You install the latest driver, but check the minimum version for the architecture and compatibility with the OS kernel. CUDA lives in the image.
Go deeper if needed: the CUDA version in nvidia-smi is the compatible one, not the installed one. If you deploy at a customer's site, reproduce their environment down to the driver.
ML + OPS
GPU in K8s
GPU Operator installs drivers and the device plugin, and pods request the nvidia.com/gpu resource. Preparing a GPU node can take up to 5 minutes, and that hurts autoscaling.
Enough for MLOps: you've installed GPU Operator and run a pod with a GPU.
In vanilla Kubernetes a container gets a whole card: a 500 MB service eats an H100. A common ask from companies: several data scientists on one GPU in JupyterHub.
Enough for MLOps: you know MIG (up to 7 isolated slices), MPS, time-slicing and HAMi, and pick one based on the card's architecture and the workload.
Try it hands-on
slice an A100 with MIG and run two pods tutorial ↗
compare MIG, MPS, time-slicing and HAMi tutorial ↗
Go deeper if needed: fair share between teams, gang scheduling for distributed training.
ML + OPS
Distributed training
How GPUs are wired together: PCIe (through CPU and RAM), NVLink (a bridge between cards), NVSwitch (a switch for 8 cards), InfiniBand between nodes. Sometimes a cluster is put together by people who don't know what these are.
Enough for MLOps: you know it exists and when you need it. NVLink matters for training and for models that don't fit on one card. For inference of small models PCIe is enough.
Try it hands-on
NVLink, NVSwitch, PCIe, InfiniBand: what connects to what tutorial ↗
Ollama is for home. vLLM currently wins on effort vs result: TensorRT-LLM is faster on paper, but building it is not a 5-minute job. For fault tolerance and routing, look at vLLM Production Stack.
Enough for MLOps: you've run vLLM in Docker and tuned the config: max-num-seqs, context, tensor parallel. You understand the trade-off between minimum latency and maximum throughput.
Try it hands-on
run the same model in Ollama, vLLM, SGLang and Triton tutorial ↗
Go deeper if needed: KV cache shared across replicas (LMCache), CUDA graphs, offloading weights to CPU, config autotuning.
ML + OPS
Optimization
Just as DevOps slims down a developer's Dockerfile, MLOps turns a suboptimal model into an artifact that runs fast and weighs little. The format chain: PyTorch or TensorFlow → ONNX → TensorRT.
Enough for MLOps: you understand formats and precisions (bf16, fp8, int4). You check quantization on your own set of reference answers, not on benchmarks: int4 takes about 30% of the memory.
Try it hands-on
export a model to ONNX and run it in ONNX Runtime tutorial ↗
An inference server takes a model as an artifact and exposes an endpoint. At the start a model in a container with FastAPI is enough, but once there are many models that update often, a homemade service will hit framework versions, formats and zero-downtime updates. Inference can be online (HTTP, gRPC) or offline (a scheduled job).
Enough for MLOps: you've set up Triton or KServe and understand the difference between a model format and a server. You know the two approaches: a service per model, or a shared server that loads a model by name. For a couple of simple models ONNX Runtime or BentoML in Docker is enough, Triton would be overkill there.
Go deeper if needed: Triton is hard to get into and has a lot of docs: start with the Conceptual Guide. Next up is NVIDIA Dynamo for distributed inference.
ML + OPS + DEV
Autoscaling
Spikes arrive faster than replicas start: a GPU node takes minutes to get ready, and images and weights weigh tens of gigabytes. CPU-based HPA doesn't see GPU load.
Enough for MLOps: you scale on queue length or concurrency, not CPU, and understand what a cold start is made of.
Try it hands-on
KServe with concurrency-based autoscaling tutorial ↗
Go deeper if needed: pre-provisioned nodes, a weight cache on NFS, Tensorizer, scale to zero.
ML + OPS + DEV
Rollout & A/B
A new model doesn't go to everyone at once. Canary: a small share of traffic first, you watch errors and business metrics, then add more. Blue-green keeps the old version next to the new one so you can roll back in seconds. A/B shows whether the new model is better on real users, not on a test set.
Enough for MLOps: you understand canary, blue-green and A/B, have rolled a model out to a share of traffic and rolled it back when the metrics dropped.
Try it hands-on
canary in KServe: 10% of traffic to the new model version, then promote or roll back tutorial ↗
Argo Rollouts: canary with Prometheus metric analysis and automatic rollback tutorial ↗
Go deeper if needed: automatic rollback on SLOs and business metrics with Argo Rollouts or Flagger. A/B needs more than a rollout: it needs an honest measurement of how much traffic and time it takes for the difference not to be noise.
ML + OPS + DEV
ML platform
A platform shows up when every team builds its own little MLOps and the solutions drift apart. The platform team moves the shared routine into a horizontal layer, while models and business logic stay with the product teams. It's assembled from blocks, like LEGO: IDE, experiment tracking, S3, pipelines, inference. In Kubernetes almost all of it installs with Helm charts.
Enough for MLOps: you can draw a platform out of blocks across three layers (data, ML, inference), pick one tool per layer for a small company and set up a couple of them.
Go deeper if needed: the platform as an internal product. Risks: vendor lock-in with managed clouds, long integrations, product speed versus platform standards, and over-engineering "like big tech" without big tech scale.
ML + OPS + DEV
ML monitoring
The metrics are the same as in the Ops part: inference is a regular service. But a service can answer 200 OK while the model overprices every order. So on top you watch the model's own signals: input drift, the gap between prediction and reality, and business metrics.
Enough for MLOps: inference metrics in Prometheus (RPS, latency, queue, GPU), business metrics and an alert on drift.
Try it hands-on
attach Evidently to a model and shift the data tutorial ↗
dig into observability in vLLM Production Stack tutorial ↗
Go deeper if needed: prediction logs, so the product manager and analyst can study the model's effect themselves. An automatic reaction to drift: retraining or rollback. Arize and WhyLabs are managed alternatives to Evidently: faster to start, but expensive to move away from later.
ML + OPS + DEV
Cost
Choose hardware by metrics, not by gut feeling: GPU or CPU, gaming or data center card, with NVLink or without. You tell the client: at this price you get these metrics, at that price those.
Enough for MLOps: you compare configurations by price and metrics (TTFT, throughput) and know that small models and recommenders are often cheaper on CPU.
The easiest place to start the track: you can refresh the basics in an evening. In practice: a data scientist wants to open a service in the browser, and you set up ingress and HTTPS with automatic certificate renewal.
Enough for MLOps: OSI, IP, DNS, TLS, how VMs are connected, ingress in Kubernetes.
Try it hands-on
explain what happens when you type google.com tutorial ↗
Read and watch
bookOlifer: Computer Networks (no need to read it all)
Open the product list of any cloud and you'll see every System Design component: VMs, managed databases and Kubernetes, S3, registry.
Enough for MLOps: you don't need to learn every cloud: pick one and build a mini project on GitHub where the infrastructure is deployed with Terraform.
Try it hands-on
deploy managed Kubernetes in Yandex Cloud (new accounts get a bonus) tutorial ↗
Almost every ML platform is built on Kubernetes: orchestration, isolation, zero-downtime updates. All MLOps tools install with Helm charts, and a Helm chart is almost Docker Compose for Kubernetes.
Enough for MLOps: you install Helm charts and know the objects from pod to ingress. If you don't need autoscaling, zero-downtime updates and environment isolation, a VM with Docker Compose is enough for a small project.
Try it hands-on
minikube on a laptop: the same objects without a cloud tutorial ↗
Go deeper if needed: writing Helm charts, Kustomize, security policies.
OPS
Monitoring
So you don't have to watch services 24/7: an alert comes in, you go fix it. Exporters on services, Prometheus scrapes metrics, Grafana draws them, alerts go to Telegram. Concepts matter more than a specific tool.
Enough for MLOps: you've set up Prometheus, Grafana and an alert yourself.
Go deeper if needed: VictoriaMetrics is faster than Prometheus. Next to metrics live logs (Loki, or the heavy ELK for full-text search) and traces with OpenTelemetry. The top level is a single system where metrics, logs and traces are linked.
OPS
Security
A regular Ops task, nothing MLOps-specific: no credentials in the repo, code gets checked, inference sits behind authorization. A typical hole in LLM services: requests aren't logged, access isn't restricted at all, and nobody knows where the service gets its data.
Enough for MLOps: tokens in Vault or masked CI variables, linters and SAST in the pipeline, authorization in front of inference.
bookDesigning Data-Intensive Applications, Martin Kleppmann
OPS
IaC
A snowflake server: someone SSHed in, did things by hand, left no docs, and the next person gets a quest. A phoenix server is rebuilt from code. Clicking through 40 VMs in a cloud takes ages, Terraform does it in a minute.
Enough for MLOps: Terraform brings up a VM, Ansible configures it. In the cloud and in small companies this is the basics, in big tech the infrastructure is usually already in place.
Try it hands-on
Terraform brings up a VM, Ansible configures it tutorial ↗
MLOps builds infrastructure from the same components: S3, databases, load balancers. For an infra engineer an ML service on the diagram is just a box. A GPU is one more resource, the place where the weights live.
Enough for MLOps: you can pass a mock interview: requirements, load estimates, API, databases, components. You understand that fault tolerance comes from replicas in different zones behind a load balancer, not from weights spread across several GPUs.
Try it hands-on
mock interview: design a service step by step tutorial ↗
The easiest way to run an LLM is in a container. By default a container has no GPU access, you need a runtime.
Enough for MLOps: a Dockerfile with the right layer order (things that change often go last), a container with a GPU, Docker Compose and a private registry.
Try it hands-on
run a container with a GPU through Container Toolkit tutorial ↗
slim down an image with a multi-stage build tutorial ↗
Go deeper if needed: large ML images: CUDA alone weighs ~6 GB. Lazy loading, zstd, P2P distribution. Scanning images for vulnerabilities, for example with Trivy.
OPS + DEV
CI/CD
Same as in DevOps: from the repo through a pipeline to prod. But in ML the pipeline runs not only on new code but on new data too, and it carries artifacts: datasets and models. CI/CD delivers the code and the image, while serving handles the requests: these are different processes.
Enough for MLOps: you've written a couple of pipelines in GitLab CI or GitHub Actions.
Try it hands-on
build a pipeline: linter, tests, image, deploy tutorial ↗
Without tests on real GPUs there's no answer to which configuration is optimal. You reproduce the client's scenario (50 chats, a request every 30 seconds) and compare TTFT and latency. One request from a Python script is not load.
Enough for MLOps: you can generate parallel load and measure TTFT, latency, throughput. You understand the difference between request rate and concurrency.
Go deeper if needed: if a T4 and an A100 give the same result, the batch size in the config is most likely one.
DEV
Python/Go
Tools change so fast that the docs lag behind. It happens: you try a feature from the docs, and in the issues the maintainer says it's no longer supported. You only find out if you dig into the code.
Enough for MLOps: you read Python (the language of ML) or Go (the language of services) with confidence.
Try it hands-on
write a script that calls a model API in parallel tutorial ↗
Data scientists' setups fall apart when the PyTorch or CUDA version changes. And in CI dependencies are installed on every run, so if packages take half an hour to download, that's slow: uv helps here.
Enough for MLOps: you understand conda, venv, poetry, uv and can build a reproducible environment.
Try it hands-on
anaconda, venv, uv: what's the difference tutorial ↗
Inference is a web service with the model inside as a black box. You collect the same metrics from it: RPS, latency.
Enough for MLOps: HTTP and REST, the OpenAI protocol and V2 (KServe, Triton). gRPC can wait for now. If you've never built a service with FastAPI, make a mini prototype.
Try it hands-on
a FastAPI service with a model and Swagger tutorial ↗
Go deeper if needed: SSE for token streaming, WebSocket.
DEV
Testing
You've deployed the platform: now check that the services are reachable and working. Unit tests are more for developers.
Enough for MLOps: a linter and mypy in the pipeline, unit and integration tests, an e2e check of the platform. Code review happens regularly, not when someone remembers.
Everything MLOps uses lives on GitHub. In big tech MLOps engineers take open source and patch it for their platform. Stars, CNCF status and who builds the project help you evaluate it.
Enough for MLOps: you read open source code, search through issues, have opened a PR.
Try it hands-on
take apart an open source inference platform tutorial ↗