← Journal · · Personal

Plans: inference platforms, meetups, and part two of the article

Hey everyone!

Want to make a couple of small announcements for the near future.

1️⃣ Even though posts have gotten a bit rarer, that doesn't mean the content has run out for you. Right now I'm actively researching the problem of building inference platforms, both for ML workloads and for LLMs. Expect posts with an overview of open source solutions, along with thoughts on whether to run inference on the company's existing PaaS or go with a separate solution.

2️⃣ Over the next two weeks I'll be attending meetups in Moscow, where we can network. There will be separate posts about those!

3️⃣ Starting to write part two of the article on picking infrastructure for LLMs, based on my talk. I'll go deeper into the genai-perf load testing tool, which parameters affect performance, and how LLM quantization helps.

So stay tuned! Let's keep exploring AI/ML infrastructure together!

Original on Telegram ↗

↑↓ select · Enter open · Esc close