← Journal · · Note
GPU sharing: TimeSlicing, MPS and MIG
Imagine you've got a hefty infrastructure with a GPU for ML training and inferencing (using trained models in production) tasks. But some models don't fully utilize the GPU, and you can't fit several of them at once, because, say, Kubernetes by default allocates a whole GPU to one Pod.
It's like having a 300 square meter house where you only use 40. Sounds fancy, but you're overpaying quite a bit too)
You can split your GPU (from NVIDIA) into parts, basically renting out the square meters of your house by the day)
Timeslicing: splits the GPU between tasks using time slices. Similar to "multithreading" on a single-core processor.
- available on all GPUs👍
- easy to set up😎
- no memory isolation⚠️
MPS: a special server at the CUDA level launches separate processes
- Volta architecture and above✌️
- maximum of 48 threads⚠️
- GPU memory isolation at the CUDA level🎞
MIG: splits the GPU at the hardware level
- Ampere architecture only💲
- maximum of 7 partitions🎰
- GPU memory isolation at the hardware level👌
Read more in my blog on Habr. Also, a comparison of these three technologies 😉