← Journal · · Note
NVIDIA Dynamo AI

NVIDIA released a new open source solution for LLM inference: Dynamo AI
As backends it supports popular frameworks: mistral.rs, SGLang, vLLM, and TensorRT-LLM.
The focus is on LLMs with a large number of billions of parameters, mostly ones running on multiple GPUs.
And indeed there are some questions about Triton on that front.
I haven't tested it myself yet, once I get around to it I'll try to share benchmarks. For now, the screenshot shows benchmarks from the developers.
You can read more in this article.