The text version of my infra.conf 2026 talk about an ML platform on Kubeflow is out. With links to our PRs in the community and manifests, so you can reproduce the platform yourself.
How the ML platform at Avito grew from DevBox setups and per-team unit platforms into a single Kubeflow-based platform, and what problems we hit along the way. I also cover the inference platform and what agentic platforms might look like. Useful both for people building platforms at large companies and for those working with DevBoxes or small unit platforms.
Cloud native is becoming AI native: DRA, Workload API and Inference Gateway, vLLM in Foundry Local, reliable LLM inference at Databricks, and serverless GPUs from Modal.
We built GPU sharing between ML engineers, similar to Yandex's Dev Cluster, on vanilla Kubernetes + Kueue in Kubeflow. Breaking down which entities it's made of.
As promised in the talk, links to our proposals, PRs and issues in Kubeflow Pipelines: central driver, recurring runs optimization, metrics, and pod error diagnostics.
On June 4 at infra.conf I'm talking about how we build an ML platform on Kubeflow at Avito: queues on Kueue, pipeline patches, observability, inference on KServe, and our own model registry.
Links by talk topic, and the recording of the talk: approaches to MLOps platforms, Kubeflow profiles, Kueue, Istio Sidecar, distributed training, and LLM inference in KServe.
The «MLOps for Model Development and Monitoring» course at Yandex Practicum, with a focus on infrastructure: I put together the program core and designed the course infrastructure.
Launching infrastructure digests: shipping heavy model weights, AI conformance in Kubernetes, multi-node inference on Dynamo, and llm-d in the CNCF Sandbox.
On April 22 I'm speaking at MLechny Put for the third time, this time about platformization: the path from a DevBox with JupyterLab to a centralized platform, Kueue, Kubeflow, and agentic platforms.
Recorded a podcast with Yura Klassen from Kupper: how to move into MLOps, what a typical workday looks like, what pain points you'll hit, and what skills the market needs.
Starting a series of posts about inference platforms from the basics: what inference is, how an inference service differs from a web service, and why you should separate business logic from the model.
Tomorrow I'll be at the career meetup by the self and Zvuk community: we'll talk about careers in ML and MLOps, big tech interviews, and career tracks.
Posts have gotten rarer, but there's content coming: researching inference platforms for ML and LLM, hitting meetups in Moscow, and writing part two of the article on picking infrastructure for LLMs.
Thanks for listening to the talk! I open-sourced a lab for picking infrastructure for LLMs, and I'm sharing the pictures that didn't make it into the presentation.
This Saturday I'm talking at "Ya pro backend" about how I choose infrastructure for LLMs with GenAI Perf. New in the talk: automating the selection with Terraform, the vLLM production stack, and Argo Workflows.
After a two-month pause, sharing my plans: writing an MLOps course, prepping a DevOops talk on GPU workload allocation in K8s, and releasing part one of the article on picking infrastructure for LLMs.
The first part of a series on choosing infrastructure for LLM inference when a client asks "deploy Qwen for me". We break down what makes up the required VRAM (model parameters, activations, KV cache, buffers), how to pick a GPU and run basic inference. Useful if you are sizing hardware for an LLM for the first time.
I cache HuggingFace weights in S3 for vLLM. Why you can't just dump the HF cache as-is, and the options: rclone, downloading to a local folder, tensorizing.
On May 29 I'm hosting ODS DataFest at Selectel and talking about inference autoscaling bottlenecks in K8s, GPU sharing, and schedulers for ML workloads.
Looking back at Selectel's ML meetup through a co-organizer's article. And if you want a text version of my talk on picking infrastructure, drop some reactions.
I got hands-on with KAI Scheduler, the open source scheduler from Run:ai. Breaking down the entities: queues with quotas, elastic workloads, priorities, and GPU sharing.
A guide article on a repo of Triton Inference Server tutorials is out: a demo platform, deploying different model formats and LLMs, auto-configuration, and simple UIs.
A guide to a repository of recipes for NVIDIA Triton Inference Server: an inference platform demo with autoscaling and canary deployments, running models in different formats and popular LLMs, and tuning the configuration with Model Navigator and Model Analyzer. Handy if you are building your own Triton-based inference in Kubernetes.
That's the title of my talk at the Selectel ML meetup on April 23. And from the picture you can guess which franchise the talk will reference, and what I think about TensorRT-LLM.
On April 23 I'm hosting the Selectel ML meetup in St. Petersburg and giving a talk on picking infrastructure for LLMs: GPUs for inference, vLLM and SGLang configuration, and load testing.
NVIDIA released open source Dynamo for multi-GPU LLM inference with vLLM, SGLang, TensorRT-LLM, and mistral.rs backends. Haven't tested it myself yet, but benchmarks are coming.
A guideline from a consultation: how to deploy inference on a GPU VM, from NVIDIA architectures, drivers, and CUDA to Docker images and picking an inference server.
A new episode of the "Segodnya na retro" podcast is out with me as a guest: we talked about personal brand, how speaking helps at work, and whether companies should invest in developer speakers.
The market itself doesn't quite know what MLOps is. I explain it with a diagram: three entities (Data, ML, and Inference), specialist verticals, and a horizontal value-delivery pipeline.
Found my niche (ML infrastructure), so I'm renaming the channel again. Sharing my research plans for the year and what kinds of questions you can bring to me.
Wrapping up the year with a second part on Habr: how we implemented autoscaling, canary deploy, the inference graph, and automatic Triton configuration tuning.
The second part about the Selectel inference platform: five features on top of standard Triton in Kubernetes. Canary deployments, autoscaling with faster image pulling, an inference graph, Triton optimization and a UI that needs no developers. For those who need more than just deploying a Triton Helm chart.
Found a calculator that computes how much GPU memory an LLM needs for training or inference, both for off-the-shelf models and for your own parameters.
The first part, based on the HighLoad talk: platform requirements, why we started with Seldon and moved away from it, and what the platform looks like now.
The first part of the story of the Selectel inference platform: what requirements we had, why we started with Seldon Core and ended up with NVIDIA Triton, and how the platform infrastructure and its delivery to clients work. Useful if you are choosing what to build your inference on.
Talking at HighLoad++ about building an Inference platform on Triton: why we moved away from Seldon, canary deploy, GPU node autoscaling, model chains on Ray, and a UI without frontend engineers.
At KubeCon they gave a detailed talk about DRA for GPUs: native sharing via TimeSlicing and MPS is already available, and dynamic MIG is promised for 1.33.
Needed an image-selection form when creating an instance in JupyterHub. Instead of jinja2 templates, found a simpler solution: profileList with unlisted_choice in kubespawner.
Figuring out whether you can register GPUs in Kubernetes under the model name, like MIG partitions. Spoiler: only by patching the device plugin, and nodeSelector helps you pick the card.
At Tech Day we announced the Inference platform based on NVIDIA Triton, which I'm developing. Sharing the recording of the product talk and an announcement of the technical one at HighLoad++.
This Thursday I'm talking at TechDay about our new product, an inference platform built on NVIDIA Triton: model lifecycle, autoscaling, canary deploys, and an inference graph on Ray.
Based on the webinar: I walk step by step through how resource autoscaling with GPUs works in the cloud, and show a practical setup with vLLM and GPT-2.
How GPU inference autoscaling works in Kubernetes: HPA, the node autoscaler, GPU Operator, and why image pulling slows scaling down. In practice we deploy GPT-2 on vLLM, build custom metrics with Prometheus Adapter and load-test it. Useful if you are preparing ML services for peak traffic.
Inviting you to a Kubernetes meetup where this time I'm in the audience: talks from Selectel, Flant, Magnit Tech, and Hilbert Team. Let's think about how to reuse this in ML infrastructure.
Why setting up a GPU on a server turns into a puzzle of framework, driver and OS kernel compatibility, and how NVIDIA GPU Operator automates it in Kubernetes. We look at running multiple driver versions, precompiled driver containers and the Mandela effect in nvidia-smi, plus how we use the operator for autoscaling and GPU sharing.
In Kubernetes 1.31 you'll be able to mount another container into a pod as a volume. For ML, that's a way to version model weights and configs in a container registry.
Express42 published the State of DevOps 2024 results. Wrote down what caught my eye: the cloud market, AI assistants, managed services, and DevOps engineer skills.
Ahead of a PyCon talk about the GPU Operator, I collected links on where to check the dependencies between GPU architecture, driver, CUDA, OS kernel, and framework versions.
A written version of the talk from the ML meetup: an ML system by analogy with manufacturing, picking a configuration for ML workloads, finding bottlenecks, and using GPUs efficiently.
A text version of my talk at the Selectel ML meetup: what ML systems can borrow from factory assembly lines, how to pick the minimal configuration for an ML workload, and how to find bottlenecks with Goldratt's theory of constraints and a profiler. And then how to use GPUs efficiently in training and inference.
Heading off on vacation and leaving a collection for DevOps: awesome-lists, roadmaps, guides. Sources were gathered with colleagues, and ChatGPT helped structure it.
Continuing to dig into pulling and deploying images in containerd: trying nerdctl, eStargz, Nydus, SOCI, and zstd on my own images. Spoiler: my favorite is zstd.
ML images weigh 10-15 GB, so a new replica during autoscaling takes ages to come up. I try a caching registry and find out the bottleneck is layer extraction.
Ahead of Selectel's Admin meetup, which I'm hosting, I share my approach to talks: story first, slides second. Plus the books and exercises that helped.
I'm speaking at MLechny Put (Selectel's ML meetup) today. Sharing the stream and the sources I used to prepare: ML system diagrams, MLOps principles, books on lean manufacturing and theory of constraints.
Starting a series 'from manufacturing automation to DevOps': I go through lean manufacturing principles using a software product and a coffee plant as examples.
Starting my own channel: I'll share research on DevOps and MLOps, thoughts on tools, and extra material for Habr articles and talks. All in plain language.
How our DataML products team at Selectel moved away from ClickOps and automated the deployment of a cloud ML platform with GitLab and Terraform: from downstream pipelines to a monolithic Terraform job, then splitting it up with remote state, plus automated tests. Useful if you want to stop being afraid of deploying to prod.
Can you repartition MIG on the fly while the GPU is already under load? We look at solutions from run.ai and Nebuly, test dynamic MIG in Kubernetes with experiments, and compare time-slicing, MPS and MIG on the same benchmark. For those who already share GPUs and want to do it more flexibly.
The follow-up on GPU sharing, this time in Kubernetes: how GPU Operator splits GPUs with MIG and time-slicing, and how to set up load balancing, monitoring and autoscaling with Prometheus Adapter and HPA. By the end of the evening we have a prototype of an autoscaling inference platform.
A hands-on guide to GPU sharing: why "one GPU per container" is inefficient and how CUDA streams, time-slicing, MIG, MPS and vGPU differ. Then we set up MIG in practice, run inference servers on the partitions and measure throughput. For those who want to share GPUs between workloads and colleagues.
GPU sharing
This site uses Yandex Metrica cookies to count visits. Details