← Journal · · Digest

mlinfra digest · 23.06.2026

#mlinfra_digest

Another roundup from my handyman.

1️⃣ Speculation Is All You Need
🔎 What it's about: Modal shows that speculative decoding gives a 2-3x boost to interactive inference, and releases new DFlash draft models for Qwen.

2️⃣ GKE Inference Gateway prefix caching accelerates AI inference
🔎 What it's about: Google describes prefix-cache-aware routing in GKE Inference Gateway and gives benchmarks on throughput, wait time, and inter-token latency.

3️⃣ What's New in the AI Platform: Agents for ML Engineering, Our Deep Learning Platform, and New Capabilities for Real-Time ML
🔎 What it's about: Databricks announced a serverless GPU AI Runtime, strengthened real-time serving, and an ML agent built into the feature/training/serving/monitoring loop.

4️⃣ Anyscale on Azure: Powering Enterprise AI at Massive Scale on Azure Kubernetes Service
🔎 What it's about: Microsoft and Anyscale are pushing Ray as a unified runtime for data prep, training, fine-tuning, inference, and agentic execution on top of AKS.

5️⃣ Portable vLLM Model Inference Kernels in Helion
🔎 What it's about: PyTorch/Red Hat integrated Helion kernels into vLLM for FP8 serving on Qwen3 and showed a throughput boost in the fusion-heavy inference path.

Original on Telegram ↗

↑↓ select · Enter open · Esc close