The Hallway Track
Research Findings

Can LLMs Write Fast Multi-GPU Kernels? — Simran Arora, Together AI

AI Engineer · Aug 27, 2026 · Research Findings

Together AI built a benchmark to test whether frontier LLMs can generate efficient multi-GPU kernels

“we've sort of shifted the bottleneck to multi-GPU communication”

Simran Arora of Together AI presented research on whether frontier LLMs understand the principles behind multi-GPU kernel design, introducing a new evaluation called 'parallel kernel bench.' The work is motivated by a fundamental shift in AI hardware bottlenecks: as single-GPU kernel efficiency has improved (flash attention, DeepSeek architectures), the limiting factor has moved to multi-GPU communication. Evaluating LLMs on this task tests whether models can reason about a domain that directly gates the performance of large-scale AI training and inference.

multi-GPU kernel generation benchmarks Together AI GPU performance LLM capabilities

Watch / read the original source →