← Journal · · Talk sources
Inference autoscaling in k8s: talk sources
Materials from my talk at ODS Data Fest, all sources in one message.
1. Inference on Triton
How we built a platform on top of Triton Triton metrics for scaling Triton dashboard for scaling info
2. Autoscaling
What is it? Autoscaling in k8s with GPU nodes Node autoscaler in k8s How the GPU operator works Webinar on vLLM autoscaling
Speeding up autoscaling Moving the model cache to PVC S3 Harbor and Kuik for caching images on worker nodes Lazy loading images: how it helps speed up container pulls eStargz, SOCI, Nydus: snapshotters for lazy image loading ZSTD image format, faster than gzip Great talk on speeding up image uploads
3. GPU allocation
Sharing GPUs MIG, MPS, Timeslicing: GPU sharing technologies and their comparison A benchmark for measuring GPU performance based on YOLOv5 Hami-Project and an article on setting it up Autoscaling with MIG Dynamic MIG using InstaSlice and Nos
Scheduling
DRA driver KAI-scheduler Kueue Volcano
Thanks a lot for listening to the talk, leave reactions and comments under this post :)