<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"><channel><title>ml-infra.pro · journal</title><description>Anton Alekseev, MLOps engineer: GPUs in Kubernetes, inference, ML platforms. Roadmaps, talks and a journal on ML infrastructure.</description><link>https://ml-infra.pro/en/</link><language>en</language><item><title>ML for big companies: an article on an ML platform built on Kubeflow</title><link>https://ml-infra.pro/en/journal/2026-06-29-122/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2026-06-29-122/</guid><description>The text version of my infra.conf 2026 talk about an ML platform on Kubeflow is out. With links to our PRs in the community and manifests, so you can reproduce the platform yourself.</description><pubDate>Mon, 29 Jun 2026 08:21:20 GMT</pubDate></item><item><title>ML for large companies: from DevBox to a platform for a thousand users</title><link>https://habr.com/ru/companies/avito/articles/1050838/</link><guid isPermaLink="true">https://habr.com/ru/companies/avito/articles/1050838/</guid><description>How the ML platform at Avito grew from DevBox setups and per-team unit platforms into a single Kubeflow-based platform, and what problems we hit along the way. I also cover the inference platform and what agentic platforms might look like. Useful both for people building platforms at large companies and for those working with DevBoxes or small unit platforms.</description><pubDate>Fri, 26 Jun 2026 00:00:00 GMT</pubDate></item><item><title>mlinfra digest · 23.06.2026</title><link>https://ml-infra.pro/en/journal/2026-06-23-121/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2026-06-23-121/</guid><description>Speculative decoding from Modal, prefix caching in GKE Inference Gateway, the Databricks AI platform, and Ray on AKS.</description><pubDate>Tue, 23 Jun 2026 08:09:45 GMT</pubDate></item><item><title>mlinfra digest · 17.06.2026</title><link>https://ml-infra.pro/en/journal/2026-06-17-120/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2026-06-17-120/</guid><description>Cloud native is becoming AI native: DRA, Workload API and Inference Gateway, vLLM in Foundry Local, reliable LLM inference at Databricks, and serverless GPUs from Modal.</description><pubDate>Wed, 17 Jun 2026 07:49:48 GMT</pubDate></item><item><title>GPU sharing like Yandex, but open source: Kueue</title><link>https://ml-infra.pro/en/journal/2026-06-11-119/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2026-06-11-119/</guid><description>We built GPU sharing between ML engineers, similar to Yandex&apos;s Dev Cluster, on vanilla Kubernetes + Kueue in Kubeflow. Breaking down which entities it&apos;s made of.</description><pubDate>Thu, 11 Jun 2026 10:12:42 GMT</pubDate></item><item><title>Contributions to Kubeflow Pipelines: links from the talk</title><link>https://ml-infra.pro/en/journal/2026-06-04-116/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2026-06-04-116/</guid><description>As promised in the talk, links to our proposals, PRs and issues in Kubeflow Pipelines: central driver, recurring runs optimization, metrics, and pod error diagnostics.</description><pubDate>Thu, 04 Jun 2026 15:41:01 GMT</pubDate></item><item><title>A talk at infra.conf 2026!</title><link>https://ml-infra.pro/en/journal/2026-05-25-115/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2026-05-25-115/</guid><description>On June 4 at infra.conf I&apos;m talking about how we build an ML platform on Kubeflow at Avito: queues on Kueue, pipeline patches, observability, inference on KServe, and our own model registry.</description><pubDate>Mon, 25 May 2026 13:12:38 GMT</pubDate></item><item><title>MLechny Put 2026: talk sources</title><link>https://ml-infra.pro/en/journal/2026-04-22-114/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2026-04-22-114/</guid><description>Links by talk topic, and the recording of the talk: approaches to MLOps platforms, Kubeflow profiles, Kueue, Istio Sidecar, distributed training, and LLM inference in KServe.</description><pubDate>Wed, 22 Apr 2026 13:03:01 GMT</pubDate></item><item><title>mlinfra digest · 04/21/2026</title><link>https://ml-infra.pro/en/journal/2026-04-21-113/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2026-04-21-113/</guid><description>A digest ahead of MLechny Put: updates on open source serving projects, async inference via GKE Inference Gateway, and DRANET on AKS.</description><pubDate>Tue, 21 Apr 2026 13:41:55 GMT</pubDate></item><item><title>mlinfra digest · 15.04.2026</title><link>https://ml-infra.pro/en/journal/2026-04-15-112/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2026-04-15-112/</guid><description>The hard part isn&apos;t spinning up a model anymore, it&apos;s managing it in production: DRANET, Inference Gateway, and an inference operator from AWS.</description><pubDate>Wed, 15 Apr 2026 08:05:52 GMT</pubDate></item><item><title>We wrote a course on MLOps!</title><link>https://ml-infra.pro/en/journal/2026-04-14-111/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2026-04-14-111/</guid><description>The «MLOps for Model Development and Monitoring» course at Yandex Practicum, with a focus on infrastructure: I put together the program core and designed the course infrastructure.</description><pubDate>Tue, 14 Apr 2026 08:02:31 GMT</pubDate></item><item><title>Scheduling AI workloads in k8s</title><link>https://ml-infra.pro/en/journal/2026-04-09-110/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2026-04-09-110/</guid><description>What&apos;s already out there in open source for scheduling AI workloads: scheduler plugins with GPU metrics, and topology-aware scheduling from NVIDIA.</description><pubDate>Thu, 09 Apr 2026 10:37:06 GMT</pubDate></item><item><title>mlinfra digest · 04/08/2026</title><link>https://ml-infra.pro/en/journal/2026-04-08-109/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2026-04-08-109/</guid><description>Launching infrastructure digests: shipping heavy model weights, AI conformance in Kubernetes, multi-node inference on Dynamo, and llm-d in the CNCF Sandbox.</description><pubDate>Wed, 08 Apr 2026 07:45:13 GMT</pubDate></item><item><title>MLechny Put 2026</title><link>https://ml-infra.pro/en/journal/2026-04-06-108/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2026-04-06-108/</guid><description>On April 22 I&apos;m speaking at MLechny Put for the third time, this time about platformization: the path from a DevBox with JupyterLab to a centralized platform, Kueue, Kubeflow, and agentic platforms.</description><pubDate>Mon, 06 Apr 2026 13:48:19 GMT</pubDate></item><item><title>Platformization: why it exists and whether you need it</title><link>https://ml-infra.pro/en/journal/2026-02-16-105/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2026-02-16-105/</guid><description>Why companies go into platformization, how I split ML platforms into Data, ML, and Inference layers. And an open question: do you need it right now?</description><pubDate>Mon, 16 Feb 2026 14:18:55 GMT</pubDate></item><item><title>2025 wrap-up</title><link>https://ml-infra.pro/en/journal/2025-12-30-104/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2025-12-30-104/</guid><description>Wrapping up the year by track: talks and articles, mentoring and the MLOps course, channel life, and a bit of personal stuff.</description><pubDate>Tue, 30 Dec 2025 16:58:23 GMT</pubDate></item><item><title>Podcast about a career in MLOps</title><link>https://ml-infra.pro/en/journal/2025-12-21-101/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2025-12-21-101/</guid><description>Recorded a podcast with Yura Klassen from Kupper: how to move into MLOps, what a typical workday looks like, what pain points you&apos;ll hit, and what skills the market needs.</description><pubDate>Sun, 21 Dec 2025 14:10:43 GMT</pubDate></item><item><title>Inference serving platform: in the beginning was the Word</title><link>https://ml-infra.pro/en/journal/2025-12-01-98/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2025-12-01-98/</guid><description>Starting a series of posts about inference platforms from the basics: what inference is, how an inference service differs from a web service, and why you should separate business logic from the model.</description><pubDate>Mon, 01 Dec 2025 15:27:44 GMT</pubDate></item><item><title>Career meetup by self and Zvuk</title><link>https://ml-infra.pro/en/journal/2025-11-29-97/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2025-11-29-97/</guid><description>Tomorrow I&apos;ll be at the career meetup by the self and Zvuk community: we&apos;ll talk about careers in ML and MLOps, big tech interviews, and career tracks.</description><pubDate>Sat, 29 Nov 2025 14:50:50 GMT</pubDate></item><item><title>Plans: inference platforms, meetups, and part two of the article</title><link>https://ml-infra.pro/en/journal/2025-11-29-96/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2025-11-29-96/</guid><description>Posts have gotten rarer, but there&apos;s content coming: researching inference platforms for ML and LLM, hitting meetups in Moscow, and writing part two of the article on picking infrastructure for LLMs.</description><pubDate>Sat, 29 Nov 2025 12:55:19 GMT</pubDate></item><item><title>Following up on choosing infrastructure for LLMs</title><link>https://ml-infra.pro/en/journal/2025-10-04-91/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2025-10-04-91/</guid><description>Thanks for listening to the talk! I open-sourced a lab for picking infrastructure for LLMs, and I&apos;m sharing the pictures that didn&apos;t make it into the presentation.</description><pubDate>Sat, 04 Oct 2025 10:46:36 GMT</pubDate></item><item><title>Announcement: talk at &quot;Ya pro backend&quot;</title><link>https://ml-infra.pro/en/journal/2025-09-30-90/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2025-09-30-90/</guid><description>This Saturday I&apos;m talking at &quot;Ya pro backend&quot; about how I choose infrastructure for LLMs with GenAI Perf. New in the talk: automating the selection with Terraform, the vLLM production stack, and Argo Workflows.</description><pubDate>Tue, 30 Sep 2025 17:07:04 GMT</pubDate></item><item><title>DevOops talk sources: GPU inference in K8s: acceleration, sharing, and scaling without pain</title><link>https://ml-infra.pro/en/journal/2025-09-17-89/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2025-09-17-89/</guid><description>All sources for the DevOops talk in one message: inference as a web service, faster autoscaling, GPU schedulers, and GPU sharing.</description><pubDate>Wed, 17 Sep 2025 11:06:24 GMT</pubDate></item><item><title>I&apos;m back: an MLOps course, DevOops, and an article on LLMs</title><link>https://ml-infra.pro/en/journal/2025-08-29-88/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2025-08-29-88/</guid><description>After a two-month pause, sharing my plans: writing an MLOps course, prepping a DevOops talk on GPU workload allocation in K8s, and releasing part one of the article on picking infrastructure for LLMs.</description><pubDate>Fri, 29 Aug 2025 09:20:15 GMT</pubDate></item><item><title>Taming LLMs: choosing inference infrastructure. Part 1</title><link>https://habr.com/ru/companies/selectel/articles/941784/</link><guid isPermaLink="true">https://habr.com/ru/companies/selectel/articles/941784/</guid><description>The first part of a series on choosing infrastructure for LLM inference when a client asks &quot;deploy Qwen for me&quot;. We break down what makes up the required VRAM (model parameters, activations, KV cache, buffers), how to pick a GPU and run basic inference. Useful if you are sizing hardware for an LLM for the first time.</description><pubDate>Fri, 29 Aug 2025 00:00:00 GMT</pubDate></item><item><title>Time to remember how to cook up a Triton Inference Server</title><link>https://ml-infra.pro/en/journal/2025-06-24-87/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2025-06-24-87/</guid><description>The Evrone channel published the recording of my talk on a private inference platform install. A good reason to remember how the platform is built.</description><pubDate>Tue, 24 Jun 2025 11:16:26 GMT</pubDate></item><item><title>Tensorizing, or fast loading of model weights into GPU</title><link>https://ml-infra.pro/en/journal/2025-06-06-84/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2025-06-06-84/</guid><description>How tensorizing weights speeds up loading a model into the GPU at vLLM startup. Inside: measurements on Qwen3-8B and Qwen3-32B and config examples.</description><pubDate>Fri, 06 Jun 2025 10:01:36 GMT</pubDate></item><item><title>How to store HuggingFace model weights in S3</title><link>https://ml-infra.pro/en/journal/2025-06-05-83/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2025-06-05-83/</guid><description>I cache HuggingFace weights in S3 for vLLM. Why you can&apos;t just dump the HF cache as-is, and the options: rclone, downloading to a local folder, tensorizing.</description><pubDate>Thu, 05 Jun 2025 07:34:42 GMT</pubDate></item><item><title>We&apos;re 500 people now!</title><link>https://ml-infra.pro/en/journal/2025-05-30-79/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2025-05-30-79/</guid><description>We hit 500! To mark it, a video with a DGX B200 cluster and some channel stats: posts, talks, articles, consultations, and podcasts.</description><pubDate>Fri, 30 May 2025 07:36:12 GMT</pubDate></item><item><title>Inference autoscaling in k8s: talk sources</title><link>https://ml-infra.pro/en/journal/2025-05-29-78/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2025-05-29-78/</guid><description>Sources for my talk at ODS Data Fest, all in one message: inference on Triton, autoscaling and speeding it up, GPU allocation and sharing.</description><pubDate>Thu, 29 May 2025 15:28:01 GMT</pubDate></item><item><title>ODS visiting Selectel!</title><link>https://ml-infra.pro/en/journal/2025-05-27-76/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2025-05-27-76/</guid><description>On May 29 I&apos;m hosting ODS DataFest at Selectel and talking about inference autoscaling bottlenecks in K8s, GPU sharing, and schedulers for ML workloads.</description><pubDate>Tue, 27 May 2025 15:14:20 GMT</pubDate></item><item><title>How MLechny Put 2025 went</title><link>https://ml-infra.pro/en/journal/2025-05-14-75/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2025-05-14-75/</guid><description>Looking back at Selectel&apos;s ML meetup through a co-organizer&apos;s article. And if you want a text version of my talk on picking infrastructure, drop some reactions.</description><pubDate>Wed, 14 May 2025 12:55:29 GMT</pubDate></item><item><title>List of useful sources: &quot;Taming the LLM&quot;</title><link>https://ml-infra.pro/en/journal/2025-04-23-74/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2025-04-23-74/</guid><description>Sources for the talk on picking infrastructure for LLMs: VRAM estimation, inference frameworks, vLLM configuration, and load testing tools.</description><pubDate>Wed, 23 Apr 2025 15:32:41 GMT</pubDate></item><item><title>KAI Scheduler: Nvidia&apos;s native K8s scheduler</title><link>https://ml-infra.pro/en/journal/2025-04-18-72/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2025-04-18-72/</guid><description>I got hands-on with KAI Scheduler, the open source scheduler from Run:ai. Breaking down the entities: queues with quotas, elastic workloads, priorities, and GPU sharing.</description><pubDate>Fri, 18 Apr 2025 14:20:43 GMT</pubDate></item><item><title>How to prep Triton: recipes for your inference platform</title><link>https://ml-infra.pro/en/journal/2025-04-17-71/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2025-04-17-71/</guid><description>A guide article on a repo of Triton Inference Server tutorials is out: a demo platform, deploying different model formats and LLMs, auto-configuration, and simple UIs.</description><pubDate>Thu, 17 Apr 2025 13:32:16 GMT</pubDate></item><item><title>Cooking with Triton: recipes for your own inference platform</title><link>https://habr.com/ru/companies/selectel/articles/901358/</link><guid isPermaLink="true">https://habr.com/ru/companies/selectel/articles/901358/</guid><description>A guide to a repository of recipes for NVIDIA Triton Inference Server: an inference platform demo with autoscaling and canary deployments, running models in different formats and popular LLMs, and tuning the configuration with Model Navigator and Model Analyzer. Handy if you are building your own Triton-based inference in Kubernetes.</description><pubDate>Thu, 17 Apr 2025 00:00:00 GMT</pubDate></item><item><title>Taming the LLM: picking inference infrastructure without the headache</title><link>https://ml-infra.pro/en/journal/2025-04-16-70/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2025-04-16-70/</guid><description>That&apos;s the title of my talk at the Selectel ML meetup on April 23. And from the picture you can guess which franchise the talk will reference, and what I think about TensorRT-LLM.</description><pubDate>Wed, 16 Apr 2025 09:01:19 GMT</pubDate></item><item><title>Alternative GPU sharing</title><link>https://ml-infra.pro/en/journal/2025-04-10-66/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2025-04-10-66/</guid><description>MIG, MPS, and TimeSlicing aren&apos;t the only options. Looking at alternatives: HAMi with vGPU and dynamic-mig, DRA, and sharing in KAI Scheduler.</description><pubDate>Thu, 10 Apr 2025 07:29:02 GMT</pubDate></item><item><title>ML meetup at Selectel 2025!</title><link>https://ml-infra.pro/en/journal/2025-04-02-65/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2025-04-02-65/</guid><description>On April 23 I&apos;m hosting the Selectel ML meetup in St. Petersburg and giving a talk on picking infrastructure for LLMs: GPUs for inference, vLLM and SGLang configuration, and load testing.</description><pubDate>Wed, 02 Apr 2025 11:43:48 GMT</pubDate></item><item><title>Announcement: talk at the Ostrovok DevOps meetup!</title><link>https://ml-infra.pro/en/journal/2025-03-24-64/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2025-03-24-64/</guid><description>On March 27 in Moscow I&apos;m talking about inference autoscaling in K8s: scale to zero and scale from zero, image caching and pull acceleration.</description><pubDate>Mon, 24 Mar 2025 09:57:00 GMT</pubDate></item><item><title>NVIDIA Dynamo AI</title><link>https://ml-infra.pro/en/journal/2025-03-20-63/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2025-03-20-63/</guid><description>NVIDIA released open source Dynamo for multi-GPU LLM inference with vLLM, SGLang, TensorRT-LLM, and mistral.rs backends. Haven&apos;t tested it myself yet, but benchmarks are coming.</description><pubDate>Thu, 20 Mar 2025 18:06:14 GMT</pubDate></item><item><title>A short guide to deploying your inference on GPU</title><link>https://ml-infra.pro/en/journal/2025-03-17-62/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2025-03-17-62/</guid><description>A guideline from a consultation: how to deploy inference on a GPU VM, from NVIDIA architectures, drivers, and CUDA to Docker images and picking an inference server.</description><pubDate>Mon, 17 Mar 2025 08:52:15 GMT</pubDate></item><item><title>Do developers need public speaking?</title><link>https://ml-infra.pro/en/journal/2025-02-26-61/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2025-02-26-61/</guid><description>A new episode of the &quot;Segodnya na retro&quot; podcast is out with me as a guest: we talked about personal brand, how speaking helps at work, and whether companies should invest in developer speakers.</description><pubDate>Wed, 26 Feb 2025 16:42:04 GMT</pubDate></item><item><title>Who exactly is an MLOps engineer?</title><link>https://ml-infra.pro/en/journal/2025-02-19-60/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2025-02-19-60/</guid><description>The market itself doesn&apos;t quite know what MLOps is. I explain it with a diagram: three entities (Data, ML, and Inference), specialist verticals, and a horizontal value-delivery pipeline.</description><pubDate>Wed, 19 Feb 2025 10:34:29 GMT</pubDate></item><item><title>Rebrand and plans for the year</title><link>https://ml-infra.pro/en/journal/2025-02-12-58/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2025-02-12-58/</guid><description>Found my niche (ML infrastructure), so I&apos;m renaming the channel again. Sharing my research plans for the year and what kinds of questions you can bring to me.</description><pubDate>Wed, 12 Feb 2025 08:03:02 GMT</pubDate></item><item><title>2024 wrap-up</title><link>https://ml-infra.pro/en/journal/2024-12-31-47/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2024-12-31-47/</guid><description>Looking back at 2024&apos;s talks, from DevOps Conf and the ML meetup to PyCon, TechDay, and HighLoad++, and sharing plans for next year.</description><pubDate>Tue, 31 Dec 2024 14:22:10 GMT</pubDate></item><item><title>Recording of the HighLoad++ 24 talk</title><link>https://ml-infra.pro/en/journal/2024-12-28-45/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2024-12-28-45/</guid><description>The HighLoad++ recordings are out: my talk about the Inference platform is now available, along with the slides and Q&amp;A.</description><pubDate>Sat, 28 Dec 2024 11:02:11 GMT</pubDate></item><item><title>Part two of the Inference platform article!</title><link>https://ml-infra.pro/en/journal/2024-12-27-44/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2024-12-27-44/</guid><description>Wrapping up the year with a second part on Habr: how we implemented autoscaling, canary deploy, the inference graph, and automatic Triton configuration tuning.</description><pubDate>Fri, 27 Dec 2024 08:09:26 GMT</pubDate></item><item><title>Five elements of the Selectel inference platform: how we built our own Avatar</title><link>https://habr.com/ru/companies/selectel/articles/867972/</link><guid isPermaLink="true">https://habr.com/ru/companies/selectel/articles/867972/</guid><description>The second part about the Selectel inference platform: five features on top of standard Triton in Kubernetes. Canary deployments, autoscaling with faster image pulling, an inference graph, Triton optimization and a UI that needs no developers. For those who need more than just deploying a Triton Helm chart.</description><pubDate>Fri, 27 Dec 2024 00:00:00 GMT</pubDate></item><item><title>A calculator for GPU memory for LLMs</title><link>https://ml-infra.pro/en/journal/2024-12-18-43/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2024-12-18-43/</guid><description>Found a calculator that computes how much GPU memory an LLM needs for training or inference, both for off-the-shelf models and for your own parameters.</description><pubDate>Wed, 18 Dec 2024 12:51:04 GMT</pubDate></item><item><title>Part one of the article about the inference platform is out!</title><link>https://ml-infra.pro/en/journal/2024-12-16-42/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2024-12-16-42/</guid><description>The first part, based on the HighLoad talk: platform requirements, why we started with Seldon and moved away from it, and what the platform looks like now.</description><pubDate>Mon, 16 Dec 2024 08:42:47 GMT</pubDate></item><item><title>NVIDIA Triton Inference Server: building production ML without developers</title><link>https://habr.com/ru/companies/selectel/articles/866256/</link><guid isPermaLink="true">https://habr.com/ru/companies/selectel/articles/866256/</guid><description>The first part of the story of the Selectel inference platform: what requirements we had, why we started with Seldon Core and ended up with NVIDIA Triton, and how the platform infrastructure and its delivery to clients work. Useful if you are choosing what to build your inference on.</description><pubDate>Mon, 16 Dec 2024 00:00:00 GMT</pubDate></item><item><title>How we built the Inference platform: sources</title><link>https://ml-infra.pro/en/journal/2024-12-02-41/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2024-12-02-41/</guid><description>Sources for my HighLoad talk in one message: Istio and canary deploy, autoscaling, faster image pulling, GPU sharing, and more.</description><pubDate>Mon, 02 Dec 2024 07:40:27 GMT</pubDate></item><item><title>Talk at HighLoad++ 2024: an Inference platform on Triton</title><link>https://ml-infra.pro/en/journal/2024-11-27-39/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2024-11-27-39/</guid><description>Talking at HighLoad++ about building an Inference platform on Triton: why we moved away from Seldon, canary deploy, GPU node autoscaling, model chains on Ray, and a UI without frontend engineers.</description><pubDate>Wed, 27 Nov 2024 14:41:09 GMT</pubDate></item><item><title>Native GPU resource support in K8s</title><link>https://ml-infra.pro/en/journal/2024-11-15-36/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2024-11-15-36/</guid><description>At KubeCon they gave a detailed talk about DRA for GPUs: native sharing via TimeSlicing and MPS is already available, and dynamic MIG is promised for 1.33.</description><pubDate>Fri, 15 Nov 2024 08:46:46 GMT</pubDate></item><item><title>Custom forms in JupyterHub</title><link>https://ml-infra.pro/en/journal/2024-11-06-35/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2024-11-06-35/</guid><description>Needed an image-selection form when creating an instance in JupyterHub. Instead of jinja2 templates, found a simpler solution: profileList with unlisted_choice in kubespawner.</description><pubDate>Wed, 06 Nov 2024 10:00:37 GMT</pubDate></item><item><title>Can you rename nvidia.com/gpu to the GPU model</title><link>https://ml-infra.pro/en/journal/2024-10-30-33/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2024-10-30-33/</guid><description>Figuring out whether you can register GPUs in Kubernetes under the model name, like MIG partitions. Spoiler: only by patching the device plugin, and nodeSelector helps you pick the card.</description><pubDate>Wed, 30 Oct 2024 13:53:53 GMT</pubDate></item><item><title>We launched Selectel&apos;s Inference platform</title><link>https://ml-infra.pro/en/journal/2024-10-11-32/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2024-10-11-32/</guid><description>At Tech Day we announced the Inference platform based on NVIDIA Triton, which I&apos;m developing. Sharing the recording of the product talk and an announcement of the technical one at HighLoad++.</description><pubDate>Fri, 11 Oct 2024 14:30:43 GMT</pubDate></item><item><title>Triton inference platform: talk at TechDay</title><link>https://ml-infra.pro/en/journal/2024-10-08-31/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2024-10-08-31/</guid><description>This Thursday I&apos;m talking at TechDay about our new product, an inference platform built on NVIDIA Triton: model lifecycle, autoscaling, canary deploys, and an inference graph on Ray.</description><pubDate>Tue, 08 Oct 2024 14:45:07 GMT</pubDate></item><item><title>AI Conf 2024: talks worth remembering</title><link>https://ml-infra.pro/en/journal/2024-09-28-30/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2024-09-28-30/</guid><description>Attended AI Conf by Ontico and wrote down the talks I remember: testing LLM applications, crowdsourcing data labeling, GPT in model training.</description><pubDate>Sat, 28 Sep 2024 11:25:33 GMT</pubDate></item><item><title>New article on inference autoscaling in Kubernetes is out</title><link>https://ml-infra.pro/en/journal/2024-09-18-29/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2024-09-18-29/</guid><description>Based on the webinar: I walk step by step through how resource autoscaling with GPUs works in the cloud, and show a practical setup with vLLM and GPT-2.</description><pubDate>Wed, 18 Sep 2024 16:58:11 GMT</pubDate></item><item><title>How to survive Black Friday traffic? GPU inference autoscaling in Kubernetes</title><link>https://habr.com/ru/companies/selectel/articles/844026/</link><guid isPermaLink="true">https://habr.com/ru/companies/selectel/articles/844026/</guid><description>How GPU inference autoscaling works in Kubernetes: HPA, the node autoscaler, GPU Operator, and why image pulling slows scaling down. In practice we deploy GPT-2 on vLLM, build custom metrics with Prometheus Adapter and load-test it. Useful if you are preparing ML services for peak traffic.</description><pubDate>Wed, 18 Sep 2024 00:00:00 GMT</pubDate></item><item><title>Kubernetes Meetup at Selectel</title><link>https://ml-infra.pro/en/journal/2024-09-10-28/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2024-09-10-28/</guid><description>Inviting you to a Kubernetes meetup where this time I&apos;m in the audience: talks from Selectel, Flant, Magnit Tech, and Hilbert Team. Let&apos;s think about how to reuse this in ML infrastructure.</description><pubDate>Tue, 10 Sep 2024 09:50:06 GMT</pubDate></item><item><title>How kittens set up GPUs in Kubernetes with their paws, and what the Mandela effect has to do with it</title><link>https://habr.com/ru/companies/selectel/articles/839528/</link><guid isPermaLink="true">https://habr.com/ru/companies/selectel/articles/839528/</guid><description>Why setting up a GPU on a server turns into a puzzle of framework, driver and OS kernel compatibility, and how NVIDIA GPU Operator automates it in Kubernetes. We look at running multiple driver versions, precompiled driver containers and the Mandela effect in nvidia-smi, plus how we use the operator for autoscaling and GPU sharing.</description><pubDate>Fri, 30 Aug 2024 00:00:00 GMT</pubDate></item><item><title>Webinar: setting up GPU in Kubernetes with GPU Operator</title><link>https://ml-infra.pro/en/journal/2024-08-29-26/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2024-08-29-26/</guid><description>Webinar in an hour: setting up a k8s cluster to work with GPUs, deploying GPT-2 inference, node autoscaling by traffic, and splitting GPUs with MIG.</description><pubDate>Thu, 29 Aug 2024 11:56:56 GMT</pubDate></item><item><title>Read Only Volumes from OCI artifacts in Kubernetes 1.31</title><link>https://ml-infra.pro/en/journal/2024-08-27-25/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2024-08-27-25/</guid><description>In Kubernetes 1.31 you&apos;ll be able to mount another container into a pod as a volume. For ML, that&apos;s a way to version model weights and configs in a container registry.</description><pubDate>Tue, 27 Aug 2024 09:20:37 GMT</pubDate></item><item><title>State of DevOps 2024: what&apos;s interesting about the cloud</title><link>https://ml-infra.pro/en/journal/2024-08-17-24/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2024-08-17-24/</guid><description>Express42 published the State of DevOps 2024 results. Wrote down what caught my eye: the cloud market, AI assistants, managed services, and DevOps engineer skills.</description><pubDate>Sat, 17 Aug 2024 13:03:50 GMT</pubDate></item><item><title>Driver, CUDA, and framework compatibility: links for the PyCon talk</title><link>https://ml-infra.pro/en/journal/2024-07-24-23/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2024-07-24-23/</guid><description>Ahead of a PyCon talk about the GPU Operator, I collected links on where to check the dependencies between GPU architecture, driver, CUDA, OS kernel, and framework versions.</description><pubDate>Wed, 24 Jul 2024 10:04:39 GMT</pubDate></item><item><title>New article: boosting GPU utilization</title><link>https://ml-infra.pro/en/journal/2024-06-27-21/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2024-06-27-21/</guid><description>A written version of the talk from the ML meetup: an ML system by analogy with manufacturing, picking a configuration for ML workloads, finding bottlenecks, and using GPUs efficiently.</description><pubDate>Thu, 27 Jun 2024 13:02:35 GMT</pubDate></item><item><title>The unbearable lightness of raising GPU utilization</title><link>https://habr.com/ru/companies/selectel/articles/822651/</link><guid isPermaLink="true">https://habr.com/ru/companies/selectel/articles/822651/</guid><description>A text version of my talk at the Selectel ML meetup: what ML systems can borrow from factory assembly lines, how to pick the minimal configuration for an ML workload, and how to find bottlenecks with Goldratt&apos;s theory of constraints and a profiler. And then how to use GPUs efficiently in training and inference.</description><pubDate>Thu, 27 Jun 2024 00:00:00 GMT</pubDate></item><item><title>Back from vacation plus a podcast about DevOps</title><link>https://ml-infra.pro/en/journal/2024-06-21-20/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2024-06-21-20/</guid><description>Back from a Rammstein concert. While I prep new content, you can listen to a podcast where I talk about my DevOps experience and the tasks I run into.</description><pubDate>Fri, 21 Jun 2024 14:04:27 GMT</pubDate></item><item><title>Useful resources for Ops folks</title><link>https://ml-infra.pro/en/journal/2024-06-09-19/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2024-06-09-19/</guid><description>Heading off on vacation and leaving a collection for DevOps: awesome-lists, roadmaps, guides. Sources were gathered with colleagues, and ChatGPT helped structure it.</description><pubDate>Sun, 09 Jun 2024 04:38:53 GMT</pubDate></item><item><title>Speeding up image pulling: eStargz, Nydus, SOCI, and zstd</title><link>https://ml-infra.pro/en/journal/2024-05-26-17/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2024-05-26-17/</guid><description>Continuing to dig into pulling and deploying images in containerd: trying nerdctl, eStargz, Nydus, SOCI, and zstd on my own images. Spoiler: my favorite is zstd.</description><pubDate>Sun, 26 May 2024 10:37:17 GMT</pubDate></item><item><title>My GPU sharing article made the Technotext finals</title><link>https://ml-infra.pro/en/journal/2024-05-15-16/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2024-05-15-16/</guid><description>My article on GPU sharing in Kubernetes made the Technotext 2023 shortlist. A quick recap: several pods on one GPU and autoscaling inference on MIG.</description><pubDate>Wed, 15 May 2024 15:08:34 GMT</pubDate></item><item><title>Lazy loading images with Stargz Snapshotter</title><link>https://ml-infra.pro/en/journal/2024-05-08-15/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2024-05-08-15/</guid><description>Digging deeper: how image pulling and extraction actually work, and how lazy layer loading via Stargz Snapshotter speeds up service startup.</description><pubDate>Wed, 08 May 2024 10:55:41 GMT</pubDate></item><item><title>Why ML images take so long to deploy</title><link>https://ml-infra.pro/en/journal/2024-05-08-14/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2024-05-08-14/</guid><description>ML images weigh 10-15 GB, so a new replica during autoscaling takes ages to come up. I try a caching registry and find out the bottleneck is layer extraction.</description><pubDate>Wed, 08 May 2024 10:55:12 GMT</pubDate></item><item><title>How I prep for talks</title><link>https://ml-infra.pro/en/journal/2024-04-24-13/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2024-04-24-13/</guid><description>Ahead of Selectel&apos;s Admin meetup, which I&apos;m hosting, I share my approach to talks: story first, slides second. Plus the books and exercises that helped.</description><pubDate>Wed, 24 Apr 2024 15:09:42 GMT</pubDate></item><item><title>MLechny Put 2024: talk sources</title><link>https://ml-infra.pro/en/journal/2024-04-18-9/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2024-04-18-9/</guid><description>I&apos;m speaking at MLechny Put (Selectel&apos;s ML meetup) today. Sharing the stream and the sources I used to prepare: ML system diagrams, MLOps principles, books on lean manufacturing and theory of constraints.</description><pubDate>Thu, 18 Apr 2024 15:05:08 GMT</pubDate></item><item><title>Lean manufacturing for a DevOps engineer</title><link>https://ml-infra.pro/en/journal/2024-04-17-8/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2024-04-17-8/</guid><description>Starting a series &apos;from manufacturing automation to DevOps&apos;: I go through lean manufacturing principles using a software product and a coffee plant as examples.</description><pubDate>Wed, 17 Apr 2024 19:21:12 GMT</pubDate></item><item><title>GPU sharing: TimeSlicing, MPS and MIG</title><link>https://ml-infra.pro/en/journal/2024-04-04-6/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2024-04-04-6/</guid><description>A quick rundown of three ways to split an NVIDIA GPU between workloads: how TimeSlicing, MPS and MIG differ in GPU support and memory isolation.</description><pubDate>Thu, 04 Apr 2024 11:36:24 GMT</pubDate></item><item><title>Introduction: what this channel is about</title><link>https://ml-infra.pro/en/journal/2024-04-04-4/</link><guid isPermaLink="true">https://ml-infra.pro/en/journal/2024-04-04-4/</guid><description>Starting my own channel: I&apos;ll share research on DevOps and MLOps, thoughts on tools, and extra material for Habr articles and talks. All in plain language.</description><pubDate>Thu, 04 Apr 2024 08:39:38 GMT</pubDate></item><item><title>How we made cloud platform deploys 20x faster and got rid of panic attacks</title><link>https://habr.com/ru/companies/selectel/articles/803883/</link><guid isPermaLink="true">https://habr.com/ru/companies/selectel/articles/803883/</guid><description>How our DataML products team at Selectel moved away from ClickOps and automated the deployment of a cloud ML platform with GitLab and Terraform: from downstream pipelines to a monolithic Terraform job, then splitting it up with remote state, plus automated tests. Useful if you want to stop being afraid of deploying to prod.</description><pubDate>Thu, 04 Apr 2024 00:00:00 GMT</pubDate></item><item><title>How to split a GPU and share it with colleagues? Dynamic GPU sharing in Kubernetes with MIG, MPS and time-slicing</title><link>https://habr.com/ru/companies/selectel/articles/776132/</link><guid isPermaLink="true">https://habr.com/ru/companies/selectel/articles/776132/</guid><description>Can you repartition MIG on the fly while the GPU is already under load? We look at solutions from run.ai and Nebuly, test dynamic MIG in Kubernetes with experiments, and compare time-slicing, MPS and MIG on the same benchmark. For those who already share GPUs and want to do it more flexibly.</description><pubDate>Fri, 24 Nov 2023 00:00:00 GMT</pubDate></item><item><title>Dividing the indivisible in Kubernetes: GPU sharing with MIG and time-slicing</title><link>https://habr.com/ru/companies/selectel/articles/756934/</link><guid isPermaLink="true">https://habr.com/ru/companies/selectel/articles/756934/</guid><description>The follow-up on GPU sharing, this time in Kubernetes: how GPU Operator splits GPUs with MIG and time-slicing, and how to set up load balancing, monitoring and autoscaling with Prometheus Adapter and HPA. By the end of the evening we have a prototype of an autoscaling inference platform.</description><pubDate>Wed, 30 Aug 2023 00:00:00 GMT</pubDate></item><item><title>How to split a GPU into parts and share it with colleagues: a hands-on guide to MIG</title><link>https://habr.com/ru/companies/selectel/articles/748544/</link><guid isPermaLink="true">https://habr.com/ru/companies/selectel/articles/748544/</guid><description>A hands-on guide to GPU sharing: why &quot;one GPU per container&quot; is inefficient and how CUDA streams, time-slicing, MIG, MPS and vGPU differ. Then we set up MIG in practice, run inference servers on the partitions and measure throughput. For those who want to share GPUs between workloads and colleagues.</description><pubDate>Tue, 18 Jul 2023 00:00:00 GMT</pubDate></item></channel></rss>