← Journal · · Talk sources
DevOops talk sources: GPU inference in K8s: acceleration, sharing, and scaling without pain
DevOops talk sources: GPU inference in K8s: acceleration, sharing, and scaling without pain
Hi to everyone who just joined! This message has all the sources in a convenient format, dig in :)
1. Inference as a web service
Article from Nvidia - the main challenges when prepping inference
2. Speeding up autoscaling
Offloading the model cache to S3 via PVC Harbor and Kuik for caching images on workloads Lazy loading of images - how it helps speed up container pulls eStargz, SOCI, Nydus - snapshotters for lazy image loading ZSTD image format - faster than gzip An interesting talk about speeding up image delivery
3. GPU schedulers
KAI Scheduler Kueue Volcano DRA driver Interesting KubeCon talks about DRA and resource sharing - one and two
4. GPU sharing MIG, MPS, Timeslicing - GPU sharing technologies and their comparison Benchmark for measuring GPU performance based on YOLOv5 Hami-Project and an article on setting it up Autoscaling with MIG Dynamic MIG using InstaSlice, Nos, and Hami
Thanks for listening to the talk! There's a lot more coming in the channel about AI/ML infrastructure 🔥