← Casesinference
An inference platform on KServe for ML and LLMs
Classic ML models and LLMs run in production side by side. Instead of shipping each model as a web service in the shared PaaS, we built a dedicated inference platform on KServe in a geo-distributed cluster.
Numbers
Before
- A model ran as a regular web service in the company PaaS, mixed with business logic: its own code, its own build, its own rollout.
- Getting a model to production took days.
- ML models and LLMs need different things: an LLM needs heavy traffic on one model, ML models need fast updates.
After
- A model from the registry goes live on the platform in minutes, with no separate service full of business logic.
- ML and LLMs in one place, more than 10k RPS on the platform.
- The platform runs in a geo-distributed cluster across several data centers.
- A new model version is swapped into the running service without a restart.
What we did
A dedicated platform instead of the shared PaaS. We chose open source KServe: it serves both ML and LLMs in one place, and the community keeps adding features.
ML inference in two KServe modes: Knative (serverless) and Standard.
Disaggregated LLM inference. LLMs live in separate KServe resources (LLMInferenceService): prefill and decode are split and scale independently, the KV cache is distributed.
Model Registry and Model Delivery are tied to inference. The registry keeps ML models, Triton repositories and LLMs from Hugging Face in S3, all versioned. Delivery swaps the model in a running service without a restart.
Geo-distributed cluster: the platform runs in several data centers at once.
Stack
KServeLLMInferenceServiceKnativeIstioS3MySQLTritonHugging Face
My role
I build the inference platform: Knative and Standard modes for ML and LLMs and a service for distributed disaggregated LLM inference.