Anton Alekseev · ML infrastructure

I make sure no GPU sits idle

ML and inference platforms on Kubernetes: GPU sharing, scheduling, inference autoscaling.

Message me on Telegram

What I solve

You can get a model running with docker run in an evening. Then it has to survive real load, share GPUs with other teams, get updates without downtime and not burn through the budget. That part is what I do.

DAY 0 · DESIGN

Which hardware and how to share it

  • Sizing infrastructure for LLMs: memory, KV cache, precision, GPU model.
  • GPU sharing by workload type: MIG, MPS, time-slicing, HAMi.
  • Quotas and priorities for teams: who can preempt whom.
Steps in the roadmap →

DAY 1 · LAUNCH

A platform, not a pile of scripts

  • ML platform: teams launch notebooks, pipelines and training on their own.
  • Inference platform on KServe, Triton and vLLM: a model gets deployed without a ticket.
  • IaC and GitOps: the cluster is recreated with a single command.
Steps in the roadmap →

DAY 2 · OPERATIONS

What breaks after launch

  • Autoscaling on queue depth and concurrency, not CPU: with vLLM the CPU is idle while the GPU is at 95%.
  • Cold start: weights and images take tens of gigabytes, a replica takes minutes to come up.
  • Utilization and fair share: priorities, preemption and quotas so research does not eat prod.
  • Observability per MIG slice and SLOs on TTFT, not just "the pod is alive".
  • Updating drivers, CUDA and models with zero downtime.
Steps in the roadmap →

5%Average GPU utilization in production Kubernetes clusters, according to the Cast AI 2026 report. An idle card costs as much as a busy one.cast.ai ↗

Cases

What I built and what came out of it. Before and after numbers from real projects.

large classifieds marketplaceinference

An inference platform on KServe for ML and LLMs

days→minutes

from a model in the registry to production, instead of a web service in the PaaS mixed with business logic

>10k RPSgeo-distributed clusterprefill/decode

getMentormentoring

20+ one-on-one consultations

Anton does not just know the subject - he is a practicing professional, and his answers come from hands-on experience.

George
Book on getMentor ↗
All cases →

Roadmaps

Journal

Habr · AvitoTech

ML for large companies: from DevBox to a platform for a thousand users ↗

How the ML platform at Avito grew from DevBox setups and per-team unit platforms into a single Kubeflow-based platform, and what problems we hit along the way. I also cover the inference platform and what agentic platforms might look like. Useful both for people building platforms at large companies and for those working with DevBoxes or small unit platforms.

ML platforms · Inference

Digest

mlinfra digest · 23.06.2026

Speculative decoding from Modal, prefix caching in GKE Inference Gateway, the Databricks AI platform, and Ray on AKS.

Inference · ML platforms

Full journal →

Talks

All talks →

Get in touch

The easiest way is to message me on Telegram.

TelegramLinkedIngetMentor

↑↓ select · Enter open · Esc close