← Journal · · Digest
mlinfra digest · 17.06.2026
#mlinfra_digest
Hi everyone!
Openclaw brought me a fresh batch of interesting articles on platform infrastructure, I picked out the most relevant ones.
1️⃣ Cloud native is now AI-native: Engineering production-ready AI
Cloud native is transforming into AI native.
Specific primitives are moving to the front: DRA, Pod Groups / Workload API, Inference Gateway, plus there's already talk about AI Conformance and security for agentic flows.
Are you already fully testing the move to DRA? Or is your prod cluster 10 versions behind upstream? :)
2️⃣ Scale On-Prem AI with Foundry Local on Azure Local: Multi-Node Inference and vLLM Support
Microsoft added multi-node inference, a vLLM runtime, and a planner that automatically picks a memory-safe config for the model and hardware in Foundry Local on Azure Local.
What caught my interest here was the planner itself: how Azure picks vLLM parameters. There's also a great table there of which parameters to tune. Attached it to the post.
3️⃣ Reliable LLM Inference at Scale
What I liked at Databricks is the focus on the fact that in production the problem isn't abstract benchmarks, it's how the system behaves under spiky demand.
They write about:
• model-aware capacity planning,
• cost-aware routing,
• autoscaling based on the real cost of the load,
• recovery from silent failures and latency degradation.
4️⃣ How we achieved truly serverless GPUs
Modal has a strong engineering breakdown of why truly serverless GPUs are hard.
The key problem is startup latency, and they solve it with warm GPU buffers, lazy filesystem/image loading, checkpoint/restore, and even CUDA checkpoint/restore.
I'd also add model tensorizing, to get models onto the GPU faster, and RDMA.