large classifieds marketplaceML platform
An ML platform on Kubeflow for 1000+ users
15%→~60%
GPU utilization: from a DevBox per person to one platform
~1000 MAU~70k runs/monthKubeflow · Kueue · Istio
Anton Alekseev · ML infrastructure
ML and inference platforms on Kubernetes: GPU sharing, scheduling, inference autoscaling.
You can get a model running with docker run in an evening. Then it has to survive real load, share GPUs with other teams, get updates without downtime and not burn through the budget. That part is what I do.
DAY 0 · DESIGN
DAY 1 · LAUNCH
DAY 2 · OPERATIONS
5%Average GPU utilization in production Kubernetes clusters, according to the Cast AI 2026 report. An idle card costs as much as a busy one.cast.ai ↗
What I built and what came out of it. Before and after numbers from real projects.
large classifieds marketplaceML platform
15%→~60%
GPU utilization: from a DevBox per person to one platform
~1000 MAU~70k runs/monthKubeflow · Kueue · Istio
large classifieds marketplaceinference
days→minutes
from a model in the registry to production, instead of a web service in the PaaS mixed with business logic
>10k RPSgeo-distributed clusterprefill/decode
cloud providerML platform
8 h→24 min
to deploy the platform for a client
bankNDA
Helm → prod
a platform on the bank's servers and a churn model with automated training and rollout
Yandex Practicumeducation
10 modules
program expert for every module: responsible for the curriculum
getMentormentoring
Anton does not just know the subject - he is a practicing professional, and his answers come from hands-on experience.
George
Eleven steps from the first GPU in a cluster to a platform where teams train and deploy models on their own.
Open →career trackA skill map at the intersection of ML, development and operations. Rate yourself and download a plan of what to learn.
Open →The text version of my infra.conf 2026 talk about an ML platform on Kubeflow is out. With links to our PRs in the community and manifests, so you can reproduce the platform yourself.
ML platforms
How the ML platform at Avito grew from DevBox setups and per-team unit platforms into a single Kubeflow-based platform, and what problems we hit along the way. I also cover the inference platform and what agentic platforms might look like. Useful both for people building platforms at large companies and for those working with DevBoxes or small unit platforms.
ML platforms · Inference
Speculative decoding from Modal, prefix caching in GKE Inference Gateway, the Databricks AI platform, and Ray on AKS.
Inference · ML platforms
infra.conf 2026 (Yandex Infrastructure), Moscow
MLechny Put 2026 (Selectel ML meetup), Moscow
ITMO x IT Infrastructure Club (podcast)
self x Zvuk career meetup, Moscow
The easiest way is to message me on Telegram.