← Journal · · Talk sources
MLechny Put 2026: talk sources
I collected links by topic from my talk, so you can not only watch the presentation, but also go deeper into the details.
And also the recording of the talk.
1. Three approaches to building an MLOps platform When we were designing the MLOps course for Yandex Practicum, we broke down three variants of MLOps platform architecture, using examples from well-known Western companies. If you want to look at different platform approaches not in the abstract, but through real architectural patterns, this is a good place to start. https://habr.com/ru/companies/yandex_praktikum/articles/1014322/
2. Kubeflow: profiles and enterprise extensions If you're interested in multi-tenancy in Kubeflow, I recommend looking at how profiles are built in Kubeflow, and then checking out an example of an enterprise plugin for automatically adding profiles and integrating with a corporate access model. Profiles: https://www.kubeflow.org/docs/components/central-dash/profiles/ IAM plugin example: https://github.com/kubeflow/kubeflow/blob/v1.10-branch/components/profile-controller/controllers/plugin_iam.go
3. Kueue We use Kueue as the main scheduler for AI workloads. If you're interested in job queues, quotas, fair sharing, and the basic principles of platform-level scheduling for ML workloads in Kubernetes, here's where you can see the core concepts. https://github.com/kubernetes-sigs/kueue https://kueue.sigs.k8s.io/docs/concepts/
4. Istio Sidecar If you use Istio as a service mesh, I recommend taking a closer look at Sidecar settings. It helps cut extra load on Envoy and noticeably save on memory. https://istio.io/latest/docs/reference/config/networking/sidecar/
5. Distributed training in Kubeflow If you're doing distributed training in Kubeflow, I recommend paying attention to the trainer-operator. It's one of the basic entry points into the platform layer for distributed training. https://github.com/kubeflow/trainer
6. Secrets and access This article does a good job covering options for integrating a Kubernetes cluster with a corporate Vault, including the practical side of working with secrets and access. https://habr.com/ru/companies/oleg-bunin/articles/919234/
7. KServe and LLM inference If you want to see what LLM inference looks like in KServe today, here's a good entry point into LLMInferenceService and the current generative inference model in the platform. https://kserve.github.io/website/docs/model-serving/generative-inference/llmisvc/llmisvc-overview
8. A timeline for MLOps through 2035 And separately, I'll leave a talk with thoughts on where MLOps platforms are heading, and why the platform layer will eventually show up practically everywhere. https://www.youtube.com/watch?v=kWBpQZIGmik
Thanks for listening to the talk! You can ask questions in the comments :)