← Journal · · Digest
mlinfra digest · 04/21/2026
Hello!
#mlinfra_digest
Today's digest is coming out a day early. And you know why? Of course you do: tomorrow I'll be speaking again, as usual, at MLechny Put 2026, make sure to come or join online :)
While we're waiting for that, let's take a look at some interesting articles from recently.
1️⃣ State of the Model Serving Communities A good roundup of the latest updates on open source serving/inference projects. Handy when you want to quickly check what's happening around KServe, the Gateway API Inference Extension, llm-d, and neighboring initiatives, without reading through dozens of separate announcements. https://inferenceops.substack.com/p/state-of-the-model-serving-communities-b93
2️⃣ Unifying real-time and async inference with GKE Inference Gateway An interesting practical breakdown of how to do async inference without maintaining separate serving setups for each type of load. And what's especially nice is that async inference here isn't just an idea to read about anymore, you can actually get your hands on it, for example through llm-d async and the Async Processor.
Article: https://cloud.google.com/blog/products/containers-kubernetes/unifying-real-time-and-async-inference-with-gke-inference-gateway Repo: https://github.com/llm-d-incubation/llm-d-async/blob/main/README.md Docs: https://llm-d.ai/docs/guide/Installation/asynchronous-processing
3️⃣ Optimizing RDMA performance for AI workloads on AKS with DRANET This one's a more low-level infra story, and about DRA again! It's about RDMA, and the topology of workload placement (so they end up close to each other and connected by something like NVLink). Things that really start to affect latency and throughput once you go beyond simple single-node serving. DRA is treated as a solution not just for GPUs, it's a general standard: network topology will get implemented through ResourceClaims too. https://blog.aks.azure.com/2026/04/01/dranet-rdma-optimization-for-ai-on-aks