Setting a World Record for MoE Pre-Training on NVIDIA GB300 NVL72
NVIDIA GB300 NVL72 sets world record pre-training DeepSeek-V3 671B at 1,648 TFLOPs per GPU
“As compute per token falls, communication increasingly determines how efficiently models scale across thousands of GPUs.”
NVIDIA's GB300 NVL72 achieved a world record for MoE pre-training by running DeepSeek-V3 671B at 1,648 TFLOPs per GPU, signaling a shift in the frontier model training bottleneck from compute to inter-GPU communication. This matters because it validates MoE as the dominant architecture for frontier models and highlights that future hardware gains must prioritize interconnect bandwidth over raw compute throughput.