← Journal · · Deep dive
A short guide to deploying your inference on GPU
Hey! I recently walked through a guideline on how to properly deploy inference on a GPU VM during a consultation. Decided to share it with you!
🚀 In the beginning, there was the Word, and the Word was Hardware.
Nvidia's datacenter GPU architecture evolves like this:
Volta → Ampere → Hopper → Blackwell
Key architecture features (each new generation includes the previous one's perks):
Volta
- MPS with 48 threads available.
Ampere
- MIG: the ability to split a GPU at the hardware level into up to 7 parts.
- Flash Attention: significantly boosts performance.
- bfloat16: a high-performance format for generative models
An interesting article about GPUs
🚀 Drivers and CUDA
By default, unless you have specific requirements, install the latest driver version depending on your OS kernel.
I also recommend checking out my driver, CUDA, and framework compatibility matrix.
🚀 Docker images
To get Docker → GPU working, you need the Container Toolkit.
The CUDA version is baked into the Docker image; on the host you only install the driver and the Toolkit.
🚀 Inference frameworks
Ollama: good for GGUF models, great for home use. Runs with a single command.
vLLM: the standard choice for serving LLMs.
Triton + vLLM: Triton as a wrapper over vLLM, lets you use DCGM exporter metrics.
Triton + TensorRT: the optimal choice if you're short on performance (RPS).
That's it for now. Drop a 🔥 reaction if you'd like each point broken out into its own, more detailed post! Every step has its own quirks and pitfalls that I've already stepped on, and I'd love for you to avoid them :)