Can LLMs Write Fast Multi-GPU Kernels? — Simran Arora, Together AI
Together AI built a benchmark to test whether frontier LLMs can generate efficient multi-GPU kernels
“we've sort of shifted the bottleneck to multi-GPU communication”
Simran Arora of Together AI presented research on whether frontier LLMs understand the principles behind multi-GPU kernel design, introducing a new evaluation called 'parallel kernel bench.' The work is motivated by a fundamental shift in AI hardware bottlenecks: as single-GPU kernel efficiency has improved (flash attention, DeepSeek architectures), the limiting factor has moved to multi-GPU communication. Evaluating LLMs on this task tests whether models can reason about a domain that directly gates the performance of large-scale AI training and inference.