← Journal · · Talk sources

How we built the Inference platform: sources

Following up on my talk at HighLoad, here are the sources that came up during the talk. All in one place, in one message 🐹

1. Istio + Canary deploy

How Istio works How to implement basic auth Canary deploy in Istio

2. Autoscaling

What is it? Autoscaling in k8s with GPU nodes Node autoscaler in k8s How the GPU operator works

Speeding up autoscaling Moving the model cache to PVC S3 Lazy loading of images: how it helps speed up container pulling eStargz, SOCI, Nydus: snapshotters for lazy image loading ZSTD image format: faster than gzip

Sharing GPUs MIG, MPS, Timeslicing: GPU sharing technologies and their comparison A GPU performance benchmark based on YOLOv5

3. Inference Graph

Ray Serve, Ray Cluster: a graph across multiple nodes Triton model ensembles: a graph on a single node

4. Optimizing Triton

How Triton optimization works Perf_client: runs load tests for Triton Model Analyzer: picks a Triton config for you Model Navigator: picks a model format for you

5. UI without developers

Istio dashboard DCGM dashboard Triton dashboard Kiali for traffic visualization

That wraps up the list of sources, hope you'll find something useful here! For those who couldn't make it to my talk, I'll soon publish articles on Habr retelling it, with a few extra clarifications added, so stay tuned :)

If you still have questions after my talk, you can ask them in the comments under this post or catch me at the conference :) Thanks for listening, you're awesome! 😊

Original on Telegram ↗

↑↓ select · Enter open · Esc close