🔬“We have foundation models for language, not for physics” — Anima Anandkumar, Bren Professor of Computing
Standard transformer scaling fails for physics: trillion-token context requirements make it impossible.
“If each dimension is even a few hundred grid points, which is where industrial scale starts... we're talking hundreds of billions to even a trillion context length. So forget ever having a transformer for anything of this scale, all of the world's compute will not be enough.”
Caltech's Anima Anandkumar argues that transformer-based foundation models hit a fundamental wall for physical systems like weather and fusion, where required context lengths reach trillions of tokens. Her alternative — Neural Operators — combines data with physical laws to model continuous, multi-scale systems without brute-force scaling. This signals a divergence from the 'scale solves everything' paradigm dominating LLM research, with significant implications for scientific AI investment and architecture choices.