← Journal · · Digest

mlinfra digest · 04/08/2026

Hey everyone!

#mlinfra_digest

I set up a research feed for myself on infrastructure for AI/ML. I want to start sharing the most interesting articles and posts on the channel from time to time.

There won't be news here about new SOTA models, how agents learned to embroider, or which Hollywood star started a GitHub repo (hi, Milla Jovovich), just an infrastructure digest about cooking up AI platforms (including for your agents).

1️⃣ The weight of AI models: Why infrastructure always arrives slowly A really good piece about how the bottleneck in AI infra is often not the model itself, but delivering, storing, and rolling out large weights. I've done posts before about DragonFly p2p, but for optimizing image pulls. Let's see how it can help with model weights: weights get split into layers like a Docker image and downloaded from different nodes. Models via torrent :)) https://www.cncf.io/blog/2026/03/27/the-weight-of-ai-models-why-infrastructure-always-arrives-slowly/

2️⃣ Kubernetes goes AI-First: Unpacking the new AI conformance program One of the strongest signals about which capabilities will soon become the baseline for AI-ready Kubernetes: DRA, AI-aware autoscaling, observability, scheduling. DRA is becoming the standard in the West, every KubeCon has a talk about it (and it's already in beta, with dynamic MIG added too). Scaling by GPU utilization looks like a necessary part of such platforms. https://opensource.googleblog.com/2026/04/kubernetes-goes-ai-first-unpacking-the-new-ai-conformance-program.html

3️⃣ The platform under the model: How cloud native powers AI engineering in production A good overview piece that pulls the cloud-native substrate under AI workloads into one picture. I especially like seeing Kueue and Kubeflow there, well, they're CNCF projects, so no surprise there. Handy if you want to quickly check your own roadmap against it and understand what primitives a production AI platform is built from these days. https://www.cncf.io/blog/2026/03/26/the-platform-under-the-model-how-cloud-native-powers-ai-engineering-in-production/

4️⃣ How NVIDIA Dynamo 1.0 Powers Multi-Node Inference at Production Scale Here we've got two articles about distributed inference. Let's split into two camps: if you're on Team Dynamo, write in the comments how much it helps you with disaggregated inference, and whether your TPOT is stable. https://developer.nvidia.com/blog/nvidia-dynamo-1-production-ready/

5️⃣ llm-d officially a CNCF Sandbox project Personally I'm keeping a closer eye on llm-d: CNCF Sandbox is a good sign, but what matters even more is that it's supported in KServe :) https://cloud.google.com/blog/products/containers-kubernetes/llm-d-officially-a-cncf-sandbox-project

Original on Telegram ↗

↑↓ select · Enter open · Esc close