Simplifying Model Serving Across Multiple GPUs with NVIDIA TensorRT Multi-Device Integration in NVIDIA Dynamo-Triton
NVIDIA TensorRT multi-device inference lets a single network execute across multiple GPUs via NCCL collectives, shipping in TensorRT 11.0.
“NVIDIA TensorRT multi-device inference is a new capability that enables a single TensorRT network to execute across multiple GPUs using NCCL-backed distributed collectives while retaining TensorRT inference optimizations.”
NVIDIA introduced TensorRT multi-device inference, fully supported in TensorRT 11.0, allowing a single TensorRT network to run across multiple GPUs using NCCL-backed distributed collectives while keeping inference optimizations. It addresses generative AI compute and memory demands that exceed a single GPU, but is an incremental infrastructure feature rather than a broad industry-shifting signal.