← Journal · · Talk sources
List of useful sources: "Taming the LLM"
Hey everyone, sharing the useful sources behind my talk on picking infrastructure for LLMs
- Qwen model card from HuggingFace
- Article on estimating VRAM when picking a GPU
- Online VRAM calculator
- Inference frameworks: Ollama, SGLang, vLLM
- vLLM configuration
- Load testing tools: locust, k6, gatling, apache jmeter, Yandex Tank
- Load testing tool for inference from Nvidia - perf analyzer, genai-perf
- genai-perf modes - Analyze, Sessions
- Quantizations: GGUF, AWQ, GPTQ
- Triton backends for LLM - vLLM, TensorRT-LLM
- Utility for importing Triton config for LLM
- Benchmarks for TensorRT-LLM
- Article comparing LLM inference backends from BentoML Cloud
How did you like the talk? Leave your reactions and comments!