← Journal · · Announcement
New article on inference autoscaling in Kubernetes is out
Hi!
Today my new article came out, based on the webinar I recently gave, about inference autoscaling in Kubernetes 👩💻
In it I walked step by step through how resource autoscaling works in the cloud when you're using GPUs, and also showed in practice how you can implement such a setup yourself with real inference (I used the vLLM framework and the GPT-2 chat model as a base).
So dive in and read the article 😎😎😎