← Journal · · Deep dive
Inference serving platform: in the beginning was the Word

Let's start a series of posts about inference platforms. I know you've been waiting for this!
And we'll start with a prologue, so: "In the beginning was the Word"
What is inference?
Let's turn to the Cambridge Dictionary and define the word Inference.
a guess that you make or an opinion that you form based on the information that you have
A model's prediction is a guess, based on the input data and the information it already has (its weights).
result = model.predict(data)
So from the definition's point of view, inference is the result of the predict operation working on the data passed to it.
An inference service is an API endpoint in front of the model.predict(data) method.
So where do you keep the business logic? In this article, NVIDIA breaks down exactly this: cases of migrating from a monolith to a separate inference service.
Why separate the business logic?
Here it's important to understand the difference from a web service with business logic. We get a new entity: the model. A kind of black box with its own lifecycle. Separating the model into its own entity lets you solve the scaling problem, especially when you need expensive resources like GPUs.
Pre/post-processing can live next to inference to reduce latency. But you need a convenient mechanism for scaling replicas of each entity separately (since they might not need a GPU, while the model does). Using pre/post-processing to implement business logic doesn't spin up a separate service, but it does put limits on maintainability and scaling.
To get rid of the headache of orchestrating and scaling resources, people use an inference serving platform.
Engine, service, platform: what's the difference?
Inference service: the final endpoint in front of the model, with autoscaling, a service mesh, the whole package. This is what the user actually interacts with. A combination of the engine and the platform.
Inference engine (runtime): a server wrapper around your model. Exposes an endpoint to your model (onnx, torch, llm). For example: vllm, triton, ml-server, onnx-runtime
Inference platform: the orchestrator of your inference services. Includes the service mesh, resource scaling, and orchestration of the inference services you create. I categorize examples of such platforms into two groups:
K8s-oriented platforms: provide CRDs for creating inference services from manifests. Orchestrate at the k8s level.
Dev-oriented platforms: a kind of Kubernetes inside Kubernetes. To create inference services, an "inference" cluster is created that essentially mirrors the work of k8s. But the interface is more geared toward developers.
I also recommend checking out this article as an overview of inference platforms. It's still relevant.
In upcoming posts I'd like to talk about each platform in more detail. For now, please vote in the poll, curious to see what everyone's using.