← Journal · · Digest
mlinfra digest · 23.06.2026
#mlinfra_digest
Another roundup from my handyman.
1️⃣ Speculation Is All You Need
🔎 What it's about: Modal shows that speculative decoding gives a 2-3x boost to interactive inference, and releases new DFlash draft models for Qwen.
2️⃣ GKE Inference Gateway prefix caching accelerates AI inference
🔎 What it's about: Google describes prefix-cache-aware routing in GKE Inference Gateway and gives benchmarks on throughput, wait time, and inter-token latency.
3️⃣ What's New in the AI Platform: Agents for ML Engineering, Our Deep Learning Platform, and New Capabilities for Real-Time ML
🔎 What it's about: Databricks announced a serverless GPU AI Runtime, strengthened real-time serving, and an ML agent built into the feature/training/serving/monitoring loop.
4️⃣ Anyscale on Azure: Powering Enterprise AI at Massive Scale on Azure Kubernetes Service
🔎 What it's about: Microsoft and Anyscale are pushing Ray as a unified runtime for data prep, training, fine-tuning, inference, and agentic execution on top of AKS.
5️⃣ Portable vLLM Model Inference Kernels in Helion
🔎 What it's about: PyTorch/Red Hat integrated Helion kernels into vLLM for FP8 serving on Qwen3 and showed a throughput boost in the fusion-heavy inference path.