← Journal · · Deep dive

A short guide to deploying your inference on GPU

Hey! I recently walked through a guideline on how to properly deploy inference on a GPU VM during a consultation. Decided to share it with you!

🚀 In the beginning, there was the Word, and the Word was Hardware.

Nvidia's datacenter GPU architecture evolves like this:

Volta → Ampere → Hopper → Blackwell

Key architecture features (each new generation includes the previous one's perks):

Volta

Ampere

An interesting article about GPUs

🚀 Drivers and CUDA

By default, unless you have specific requirements, install the latest driver version depending on your OS kernel.

I also recommend checking out my driver, CUDA, and framework compatibility matrix.

🚀 Docker images

To get Docker → GPU working, you need the Container Toolkit.

The CUDA version is baked into the Docker image; on the host you only install the driver and the Toolkit.

🚀 Inference frameworks

Ollama: good for GGUF models, great for home use. Runs with a single command.

vLLM: the standard choice for serving LLMs.

Triton + vLLM: Triton as a wrapper over vLLM, lets you use DCGM exporter metrics.

Triton + TensorRT: the optimal choice if you're short on performance (RPS).

That's it for now. Drop a 🔥 reaction if you'd like each point broken out into its own, more detailed post! Every step has its own quirks and pitfalls that I've already stepped on, and I'd love for you to avoid them :)

Original on Telegram ↗

↑↓ select · Enter open · Esc close