GEM Training: How Meta Doubled the Efficiency of Its LLM-Scale Ads Foundation Model
Meta doubled GEM ads model training efficiency to 20-25% MFU while scaling FLOPs 4x in 12 months
Meta's GEM, the foundation model powering ads recommendations on Facebook and Instagram, now trains at LLM scale on thousands of GPUs after achieving a 2x efficiency improvement to 20-25% MFU alongside a 4x FLOPs increase over 12 months. The gains required purpose-built innovations — custom attention kernels, MXFP8 mixed precision, and topology-aware 5D parallelism — because standard LLM training infrastructure does not transfer to GEM's hybrid sparse/dense recommendation architecture. This matters as a signal that frontier-scale ML engineering increasingly requires domain-specific hardware/software co-design rather than off-the-shelf LLM tooling.