Journal

Posts from the "AI & ML Infrastructure" channel and articles on Habr.

Habr · AvitoTech

ML for large companies: from DevBox to a platform for a thousand users ↗

How the ML platform at Avito grew from DevBox setups and per-team unit platforms into a single Kubeflow-based platform, and what problems we hit along the way. I also cover the inference platform and what agentic platforms might look like. Useful both for people building platforms at large companies and for those working with DevBoxes or small unit platforms.

ML platforms · Inference

Digest

mlinfra digest · 23.06.2026

Speculative decoding from Modal, prefix caching in GKE Inference Gateway, the Databricks AI platform, and Ray on AKS.

Inference · ML platforms

Digest

mlinfra digest · 17.06.2026

Cloud native is becoming AI native: DRA, Workload API and Inference Gateway, vLLM in Foundry Local, reliable LLM inference at Databricks, and serverless GPUs from Modal.

Inference · LLMs and hardware

Deep dive

GPU sharing like Yandex, but open source: Kueue

We built GPU sharing between ML engineers, similar to Yandex's Dev Cluster, on vanilla Kubernetes + Kueue in Kubeflow. Breaking down which entities it's made of.

Scheduling and utilization · GPU sharing

Announcement

A talk at infra.conf 2026!

On June 4 at infra.conf I'm talking about how we build an ML platform on Kubeflow at Avito: queues on Kueue, pipeline patches, observability, inference on KServe, and our own model registry.

ML platforms

Talk sources

MLechny Put 2026: talk sources

Links by talk topic, and the recording of the talk: approaches to MLOps platforms, Kubeflow profiles, Kueue, Istio Sidecar, distributed training, and LLM inference in KServe.

ML platforms · Scheduling and utilization

Digest

mlinfra digest · 04/21/2026

A digest ahead of MLechny Put: updates on open source serving projects, async inference via GKE Inference Gateway, and DRANET on AKS.

Inference

Digest

mlinfra digest · 15.04.2026

The hard part isn't spinning up a model anymore, it's managing it in production: DRANET, Inference Gateway, and an inference operator from AWS.

Inference

Announcement

We wrote a course on MLOps!

The «MLOps for Model Development and Monitoring» course at Yandex Practicum, with a focus on infrastructure: I put together the program core and designed the course infrastructure.

Career and community

Note

Scheduling AI workloads in k8s

What's already out there in open source for scheduling AI workloads: scheduler plugins with GPU metrics, and topology-aware scheduling from NVIDIA.

Scheduling and utilization

Digest

mlinfra digest · 04/08/2026

Launching infrastructure digests: shipping heavy model weights, AI conformance in Kubernetes, multi-node inference on Dynamo, and llm-d in the CNCF Sandbox.

Inference · ML platforms

Announcement

MLechny Put 2026

On April 22 I'm speaking at MLechny Put for the third time, this time about platformization: the path from a DevBox with JupyterLab to a centralized platform, Kueue, Kubeflow, and agentic platforms.

ML platforms

Personal

2025 wrap-up

Wrapping up the year by track: talks and articles, mentoring and the MLOps course, channel life, and a bit of personal stuff.

Career and community

Announcement

Podcast about a career in MLOps

Recorded a podcast with Yura Klassen from Kupper: how to move into MLOps, what a typical workday looks like, what pain points you'll hit, and what skills the market needs.

Career and community

Deep dive

Inference serving platform: in the beginning was the Word

Starting a series of posts about inference platforms from the basics: what inference is, how an inference service differs from a web service, and why you should separate business logic from the model.

Inference

Announcement

Career meetup by self and Zvuk

Tomorrow I'll be at the career meetup by the self and Zvuk community: we'll talk about careers in ML and MLOps, big tech interviews, and career tracks.

Career and community

Talk sources

Following up on choosing infrastructure for LLMs

Thanks for listening to the talk! I open-sourced a lab for picking infrastructure for LLMs, and I'm sharing the pictures that didn't make it into the presentation.

LLMs and hardware

Announcement

Announcement: talk at "Ya pro backend"

This Saturday I'm talking at "Ya pro backend" about how I choose infrastructure for LLMs with GenAI Perf. New in the talk: automating the selection with Terraform, the vLLM production stack, and Argo Workflows.

LLMs and hardware

Personal

I'm back: an MLOps course, DevOops, and an article on LLMs

After a two-month pause, sharing my plans: writing an MLOps course, prepping a DevOops talk on GPU workload allocation in K8s, and releasing part one of the article on picking infrastructure for LLMs.

Career and community · LLMs and hardware

Habr · Selectel

Taming LLMs: choosing inference infrastructure. Part 1 ↗

The first part of a series on choosing infrastructure for LLM inference when a client asks "deploy Qwen for me". We break down what makes up the required VRAM (model parameters, activations, KV cache, buffers), how to pick a GPU and run basic inference. Useful if you are sizing hardware for an LLM for the first time.

LLMs and hardware · Inference

Deep dive

How to store HuggingFace model weights in S3

I cache HuggingFace weights in S3 for vLLM. Why you can't just dump the HF cache as-is, and the options: rclone, downloading to a local folder, tensorizing.

Inference

Personal

We're 500 people now!

We hit 500! To mark it, a video with a DGX B200 cluster and some channel stats: posts, talks, articles, consultations, and podcasts.

Career and community

Talk sources

Inference autoscaling in k8s: talk sources

Sources for my talk at ODS Data Fest, all in one message: inference on Triton, autoscaling and speeding it up, GPU allocation and sharing.

Inference · GPU sharing

Announcement

ODS visiting Selectel!

On May 29 I'm hosting ODS DataFest at Selectel and talking about inference autoscaling bottlenecks in K8s, GPU sharing, and schedulers for ML workloads.

Inference · GPU sharing

Note

How MLechny Put 2025 went

Looking back at Selectel's ML meetup through a co-organizer's article. And if you want a text version of my talk on picking infrastructure, drop some reactions.

Career and community

Talk sources

List of useful sources: "Taming the LLM"

Sources for the talk on picking infrastructure for LLMs: VRAM estimation, inference frameworks, vLLM configuration, and load testing tools.

LLMs and hardware · Inference

Deep dive

KAI Scheduler: Nvidia's native K8s scheduler

I got hands-on with KAI Scheduler, the open source scheduler from Run:ai. Breaking down the entities: queues with quotas, elastic workloads, priorities, and GPU sharing.

Scheduling and utilization · GPU sharing

Habr · Selectel

Cooking with Triton: recipes for your own inference platform ↗

A guide to a repository of recipes for NVIDIA Triton Inference Server: an inference platform demo with autoscaling and canary deployments, running models in different formats and popular LLMs, and tuning the configuration with Model Navigator and Model Analyzer. Handy if you are building your own Triton-based inference in Kubernetes.

Inference

Deep dive

Alternative GPU sharing

MIG, MPS, and TimeSlicing aren't the only options. Looking at alternatives: HAMi with vGPU and dynamic-mig, DRA, and sharing in KAI Scheduler.

GPU sharing

Announcement

ML meetup at Selectel 2025!

On April 23 I'm hosting the Selectel ML meetup in St. Petersburg and giving a talk on picking infrastructure for LLMs: GPUs for inference, vLLM and SGLang configuration, and load testing.

LLMs and hardware · Inference

Note

NVIDIA Dynamo AI

NVIDIA released open source Dynamo for multi-GPU LLM inference with vLLM, SGLang, TensorRT-LLM, and mistral.rs backends. Haven't tested it myself yet, but benchmarks are coming.

Inference

Deep dive

A short guide to deploying your inference on GPU

A guideline from a consultation: how to deploy inference on a GPU VM, from NVIDIA architectures, drivers, and CUDA to Docker images and picking an inference server.

Inference · LLMs and hardware

Announcement

Do developers need public speaking?

A new episode of the "Segodnya na retro" podcast is out with me as a guest: we talked about personal brand, how speaking helps at work, and whether companies should invest in developer speakers.

Career and community

Note

Who exactly is an MLOps engineer?

The market itself doesn't quite know what MLOps is. I explain it with a diagram: three entities (Data, ML, and Inference), specialist verticals, and a horizontal value-delivery pipeline.

Career and community · ML platforms

Personal

Rebrand and plans for the year

Found my niche (ML infrastructure), so I'm renaming the channel again. Sharing my research plans for the year and what kinds of questions you can bring to me.

Career and community

Personal

2024 wrap-up

Looking back at 2024's talks, from DevOps Conf and the ML meetup to PyCon, TechDay, and HighLoad++, and sharing plans for next year.

Career and community

Announcement

Recording of the HighLoad++ 24 talk

The HighLoad++ recordings are out: my talk about the Inference platform is now available, along with the slides and Q&A.

Inference

Announcement

Part two of the Inference platform article!

Wrapping up the year with a second part on Habr: how we implemented autoscaling, canary deploy, the inference graph, and automatic Triton configuration tuning.

Inference

Habr · Selectel

Five elements of the Selectel inference platform: how we built our own Avatar ↗

The second part about the Selectel inference platform: five features on top of standard Triton in Kubernetes. Canary deployments, autoscaling with faster image pulling, an inference graph, Triton optimization and a UI that needs no developers. For those who need more than just deploying a Triton Helm chart.

Inference

Note

A calculator for GPU memory for LLMs

Found a calculator that computes how much GPU memory an LLM needs for training or inference, both for off-the-shelf models and for your own parameters.

LLMs and hardware

Habr · Selectel

NVIDIA Triton Inference Server: building production ML without developers ↗

The first part of the story of the Selectel inference platform: what requirements we had, why we started with Seldon Core and ended up with NVIDIA Triton, and how the platform infrastructure and its delivery to clients work. Useful if you are choosing what to build your inference on.

Inference · ML platforms

Announcement

Talk at HighLoad++ 2024: an Inference platform on Triton

Talking at HighLoad++ about building an Inference platform on Triton: why we moved away from Seldon, canary deploy, GPU node autoscaling, model chains on Ray, and a UI without frontend engineers.

Inference

Note

Native GPU resource support in K8s

At KubeCon they gave a detailed talk about DRA for GPUs: native sharing via TimeSlicing and MPS is already available, and dynamic MIG is promised for 1.33.

GPU sharing

Deep dive

Custom forms in JupyterHub

Needed an image-selection form when creating an instance in JupyterHub. Instead of jinja2 templates, found a simpler solution: profileList with unlisted_choice in kubespawner.

ML platforms

Deep dive

Can you rename nvidia.com/gpu to the GPU model

Figuring out whether you can register GPUs in Kubernetes under the model name, like MIG partitions. Spoiler: only by patching the device plugin, and nodeSelector helps you pick the card.

GPU sharing

Announcement

We launched Selectel's Inference platform

At Tech Day we announced the Inference platform based on NVIDIA Triton, which I'm developing. Sharing the recording of the product talk and an announcement of the technical one at HighLoad++.

Inference

Announcement

Triton inference platform: talk at TechDay

This Thursday I'm talking at TechDay about our new product, an inference platform built on NVIDIA Triton: model lifecycle, autoscaling, canary deploys, and an inference graph on Ray.

Inference

Note

AI Conf 2024: talks worth remembering

Attended AI Conf by Ontico and wrote down the talks I remember: testing LLM applications, crowdsourcing data labeling, GPT in model training.

Career and community

Habr · Selectel

How to survive Black Friday traffic? GPU inference autoscaling in Kubernetes ↗

How GPU inference autoscaling works in Kubernetes: HPA, the node autoscaler, GPU Operator, and why image pulling slows scaling down. In practice we deploy GPT-2 on vLLM, build custom metrics with Prometheus Adapter and load-test it. Useful if you are preparing ML services for peak traffic.

Inference

Announcement

Kubernetes Meetup at Selectel

Inviting you to a Kubernetes meetup where this time I'm in the audience: talks from Selectel, Flant, Magnit Tech, and Hilbert Team. Let's think about how to reuse this in ML infrastructure.

Career and community

Announcement

New article: boosting GPU utilization

A written version of the talk from the ML meetup: an ML system by analogy with manufacturing, picking a configuration for ML workloads, finding bottlenecks, and using GPUs efficiently.

Scheduling and utilization

Habr · Selectel

The unbearable lightness of raising GPU utilization ↗

A text version of my talk at the Selectel ML meetup: what ML systems can borrow from factory assembly lines, how to pick the minimal configuration for an ML workload, and how to find bottlenecks with Goldratt's theory of constraints and a profiler. And then how to use GPUs efficiently in training and inference.

Scheduling and utilization

Personal

Back from vacation plus a podcast about DevOps

Back from a Rammstein concert. While I prep new content, you can listen to a podcast where I talk about my DevOps experience and the tasks I run into.

Career and community

Note

Useful resources for Ops folks

Heading off on vacation and leaving a collection for DevOps: awesome-lists, roadmaps, guides. Sources were gathered with colleagues, and ChatGPT helped structure it.

Career and community

Personal

My GPU sharing article made the Technotext finals

My article on GPU sharing in Kubernetes made the Technotext 2023 shortlist. A quick recap: several pods on one GPU and autoscaling inference on MIG.

GPU sharing · Career and community

Deep dive

Why ML images take so long to deploy

ML images weigh 10-15 GB, so a new replica during autoscaling takes ages to come up. I try a caching registry and find out the bottleneck is layer extraction.

Inference

Note

How I prep for talks

Ahead of Selectel's Admin meetup, which I'm hosting, I share my approach to talks: story first, slides second. Plus the books and exercises that helped.

Career and community

Talk sources

MLechny Put 2024: talk sources

I'm speaking at MLechny Put (Selectel's ML meetup) today. Sharing the stream and the sources I used to prepare: ML system diagrams, MLOps principles, books on lean manufacturing and theory of constraints.

Scheduling and utilization

Note

Lean manufacturing for a DevOps engineer

Starting a series 'from manufacturing automation to DevOps': I go through lean manufacturing principles using a software product and a coffee plant as examples.

Note

GPU sharing: TimeSlicing, MPS and MIG

A quick rundown of three ways to split an NVIDIA GPU between workloads: how TimeSlicing, MPS and MIG differ in GPU support and memory isolation.

GPU sharing

Personal

Introduction: what this channel is about

Starting my own channel: I'll share research on DevOps and MLOps, thoughts on tools, and extra material for Habr articles and talks. All in plain language.

Habr · Selectel

How we made cloud platform deploys 20x faster and got rid of panic attacks ↗

How our DataML products team at Selectel moved away from ClickOps and automated the deployment of a cloud ML platform with GitLab and Terraform: from downstream pipelines to a monolithic Terraform job, then splitting it up with remote state, plus automated tests. Useful if you want to stop being afraid of deploying to prod.

ML platforms

Habr · Selectel

Dividing the indivisible in Kubernetes: GPU sharing with MIG and time-slicing ↗

The follow-up on GPU sharing, this time in Kubernetes: how GPU Operator splits GPUs with MIG and time-slicing, and how to set up load balancing, monitoring and autoscaling with Prometheus Adapter and HPA. By the end of the evening we have a prototype of an autoscaling inference platform.

GPU sharing · Inference

↑↓ select · Enter open · Esc close