Weight Folding, CUDA Streams, and the Bug That Made My Model Speak Backwards — Filip Makraduli
Two lines of algebra make transformer RMS norm cheaper and faster via weight folding and deferred division.
“the GPUs are not slow or bad at math, but they're bad at everything else around the actual math”
A co-authored arXiv paper proposes algebraic tricks to fuse normalization into matrix multiplication, apply weight folding, and defer division, reducing the wall-clock overhead of RMS norm layers that fire many times per decode step. It matters as a low-cost efficiency improvement to the transformer architecture, though it is an incremental optimization rather than a major industry signal.