Anima Anandkumar argues that alternative AI weather model architectures blow up over long-term rollouts because they treat the Earth as a rectangle rather than a sphere.
“This is where the Allen AI Institute has now built climate models based on our neural operator architecture. And that's the only one that works As an AI emulator, right? None of the other architectures work for climate because climate requires us to assume the world is a globe and that if you are repeatedly rolling out, you kind of can keep that information.”
quote is from the automated transcript, cleaned for reading:
filler sounds and stutters are removed, nothing is rephrased. names can be misheard
(the analysis reads context, assessments check outside sources). how →
More from Anima Anandkumar
PredictionNot checkable as stated
Transformers will never scale to high-resolution 4D physics simulations
“So forget ever having a transformer for anything of this scale. All of the world's compute will not be enough. And first of all, they all have to be co-located to be able to ever do this. So that's why we need other architectures.”
Anima AnandkumarAug 26, 2026▶ 29:41🔬 Why Transformers Hit a Wall the Moment Physics Shows Up — Anima Anandkumar, Caltech
AssertionSupported
AI models predict fusion reactor plasma disruption one million times faster
“You know, I talk about plasma and fusion reactor. You know, we barely have a few thousand samples, but we are able to accurately predict events like disruption very well. And we are able to do that a million times faster than what traditional simulations were …”
Anima AnandkumarAug 26, 2026▶ 43:50🔬 Why Transformers Hit a Wall the Moment Physics Shows Up — Anima Anandkumar, Caltech
AssertionOpen · timeframe Sep 2029
Anandkumar: Multi-physics models outperform single-physics models of equivalent parameter size
“And in fact, I was going to add that it turns out that having the model of the same size with multiple areas of physics does better than giving all of those parameters to each single physics. So if you had separate models and made them big enough as the origin…”
Anandkumar: Standard Transformers cannot scale to 5-trillion context lengths for physics
“On the other hand, if you think about using transformer architectures that have worked so well for language, that just wouldn't be able to support a five trillion context length. No matter all the compute in the world is thrown at it. So that kind of quadratic…”
Anandkumar: Dense physics feedback enables better AI self-improvement than sparse LLMs
“And the difference there is compared to language where self-improvement needs something like human feedback or other reward signals that are very sparse. They just tell you yes or no, thumbs up or down. We have dense feedback because the physics laws, there's …”
Physics-Informed Neural Networks fail on chaotic, time-dependent differential equations
“Optimization ends up being usually very difficult, especially for problems that are time-dependent, meaning it's not just stationary, you also have time, and the time component in many cases could be turbulent, like in the case of fluid dynamics, you know, you…”
Anima AnandkumarAug 26, 2026▶ 12:24🔬 Why Transformers Hit a Wall the Moment Physics Shows Up — Anima Anandkumar, Caltech
Made with StarZero
Turn any episode into a week of clips.
This entire site, over 200 episodes transcribed, diarized, checked and made playable,
runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the
moments worth sharing, cuts them, captions them, and reframes them for every feed.
We use essential cookies to make the site work. With your permission we
also use analytics cookies (Google Analytics and Mixpanel) to understand
usage and improve StarZero. See our Cookie Policy.