← Journal · · Talk sources

DevOops talk sources: GPU inference in K8s: acceleration, sharing, and scaling without pain

DevOops talk sources: GPU inference in K8s: acceleration, sharing, and scaling without pain

Hi to everyone who just joined! This message has all the sources in a convenient format, dig in :)

1. Inference as a web service

Article from Nvidia - the main challenges when prepping inference

2. Speeding up autoscaling

Offloading the model cache to S3 via PVC Harbor and Kuik for caching images on workloads Lazy loading of images - how it helps speed up container pulls eStargz, SOCI, Nydus - snapshotters for lazy image loading ZSTD image format - faster than gzip An interesting talk about speeding up image delivery

3. GPU schedulers

KAI Scheduler Kueue Volcano DRA driver Interesting KubeCon talks about DRA and resource sharing - one and two

4. GPU sharing MIG, MPS, Timeslicing - GPU sharing technologies and their comparison Benchmark for measuring GPU performance based on YOLOv5 Hami-Project and an article on setting it up Autoscaling with MIG Dynamic MIG using InstaSlice, Nos, and Hami

Thanks for listening to the talk! There's a lot more coming in the channel about AI/ML infrastructure 🔥

Original on Telegram ↗

↑↓ select · Enter open · Esc close