← Journal · · Talk sources

Inference autoscaling in k8s: talk sources

Materials from my talk at ODS Data Fest, all sources in one message.

1. Inference on Triton

How we built a platform on top of Triton Triton metrics for scaling Triton dashboard for scaling info

2. Autoscaling

What is it? Autoscaling in k8s with GPU nodes Node autoscaler in k8s How the GPU operator works Webinar on vLLM autoscaling

Speeding up autoscaling Moving the model cache to PVC S3 Harbor and Kuik for caching images on worker nodes Lazy loading images: how it helps speed up container pulls eStargz, SOCI, Nydus: snapshotters for lazy image loading ZSTD image format, faster than gzip Great talk on speeding up image uploads

3. GPU allocation

Sharing GPUs MIG, MPS, Timeslicing: GPU sharing technologies and their comparison A benchmark for measuring GPU performance based on YOLOv5 Hami-Project and an article on setting it up Autoscaling with MIG Dynamic MIG using InstaSlice and Nos

Scheduling

DRA driver KAI-scheduler Kueue Volcano

Thanks a lot for listening to the talk, leave reactions and comments under this post :)

Original on Telegram ↗

↑↓ select · Enter open · Esc close