The Hallway Track
Product Launches

Simplifying Model Serving Across Multiple GPUs with NVIDIA TensorRT Multi-Device Integration in NVIDIA Dynamo-Triton

NVIDIA Developer Blog · Sep 21, 2026 · Product Launches

NVIDIA TensorRT multi-device inference lets a single network execute across multiple GPUs via NCCL collectives, shipping in TensorRT 11.0.

“NVIDIA TensorRT multi-device inference is a new capability that enables a single TensorRT network to execute across multiple GPUs using NCCL-backed distributed collectives while retaining TensorRT inference optimizations.”

NVIDIA introduced TensorRT multi-device inference, fully supported in TensorRT 11.0, allowing a single TensorRT network to run across multiple GPUs using NCCL-backed distributed collectives while keeping inference optimizations. It addresses generative AI compute and memory demands that exceed a single GPU, but is an incremental infrastructure feature rather than a broad industry-shifting signal.

nvidia tensorrt gpu-inference model-serving dynamo-triton

Watch / read the original source →