← Journal · · Announcement
Talk at HighLoad++ 2024: an Inference platform on Triton
Next week I'll be speaking at HighLoad++ 2024: I'll talk about our experience building an Inference platform on top of Nvidia Triton Server. I'll share the following with you
- Why we moved away from Seldon
- How we implemented canary deploy for models
- How autoscaling of GPU nodes works in k8s
- How to build a model chain using Ray
- What Triton optimization on GPU gives you, and how to automatically pick the best configuration
- How we built a UI without frontend engineers
Come to the talk on December 2 at 10:00 in Hall 11! I'll answer your questions and just share the experience :)
Also, a bit of self-promo: we're now entering the second stage of developing our platform and looking for a Go developer to join the team! If you're into ML and have Go development experience (you'll be writing the backend for our platform, and maybe even a k8s operator!), message me directly :) The vacancy itself is here