Aug 26, 2026 · 1h 23m · latent-space

🔬 Why Transformers Hit a Wall the Moment Physics Shows Up — Anima Anandkumar, Caltech

Anima Anandkumar · 59m spoken RJ Haneke · 8m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Caltech Professor Anima Anandkumar explores how neural operators and Fourier Neural Operators (FNOs) surpass traditional numerical solvers and standard transformers in simulating complex physical systems, enabling breakthroughs in weather forecasting, fusion energy, and inverse design.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 5.7 Guest teaching 4.9 Guest disagreement 1.6 The hosts pushing back 2.1
05100:0020:0040:001:00:001:20:003:12–6:22 · The hosts as informed peer 5/10 Bridging AI and Science Through Symbolic Verification and TorchLean Brandon asks a solid introductory question linking differential equations, machine learning, and Anima's work on TorchLean. Anima explains her thesis on AI for science, symbolic verification, and neural network safety.6:22–10:57 · The hosts as informed peer 6/10 Formal Proofs, Certified Robustness, and Bounds in TorchLean RJ and Brandon probe how formal verification and bounds work on unconstrained neural networks like transformers. Anima explains certified robustness algorithms like CROWN and acknowledges Lean's current CPU/scalability limitations.10:57–15:50 · The hosts as informed peer 6/10 Physics-Informed Neural Networks vs. Data-Driven Neural Operators Brandon compares physics-informed neural networks (PINNs) to classical numerical analysis discretization errors. Anima educates the hosts on why PINNs fail on chaotic or time-dependent problems and why data-driven neural operators overcome optimization landscape bottlenecks.15:51–19:13 · The hosts as informed peer 5/10 Continuous Function Learning and Resolution Independence RJ asks about function fitting and regularization in under-constrained settings. Anima explains that neural operators represent continuous mappings between infinite-dimensional function spaces, enabling zero-shot super-resolution.19:13–24:34 · The hosts as informed peer 6/10 Fourier Neural Operators and Non-Local Dependency Modeling RJ asks why operating in the Fourier dual domain is advantageous and how architecture differs from standard models. Anima details the balance between quasi-linear computational efficiency and capturing non-local spatial-temporal dependencies.24:34–26:48 · The hosts as informed peer 6/10 Non-Linear Latent Spaces and Frequency Bounds in FNOs Brandon asks about frequency bounds and Nyquist limits in non-linear settings. Anima clarifies that lifting input signals into higher-dimensional non-linear latent spaces allows the network to bypass linear basis truncation constraints.26:48–31:52 · The hosts as informed peer 7/10 Why Transformers Fail on High-Resolution Multi-Dimensional Physics Anima argues transformers are completely intractable for multi-dimensional physics due to context length scaling. RJ pushes back citing vision/video autoregressive codebooks, but Anima firmly rebuts that video models operate at perceptual low-res rather than high-fidelity physics grid requirements.31:53–38:38 · The hosts as informed peer 5/10 FourCastNet: Revolutionizing Weather and Climate Modeling RJ asks about the genesis of FourCastNet and the shift in community thinking. Anima recounts defying skepticism from traditional meteorologists and achieving order-of-magnitude speedups while open-sourcing the model.38:39–44:44 · The hosts as informed peer 6/10 Climate Rollouts, Probabilistic Ensembles, and Rare Physical Events Brandon and RJ question how deterministic rollouts handle long-term climate chaos and rare events. Anima highlights probabilistic ensemble forecasting and notes physical structure enables AI to generalize well on extreme events.44:44–50:16 · The hosts as informed peer 5/10 Multi-Scale Physical Visualizations and Atmospheric Rivers Brandon connects the conversation to AlphaFold and structural biology while RJ asks to review visualizations. Anima presents multi-scale modeling visuals, explaining how soft physical constraints are incorporated into the loss function.50:16–55:57 · The hosts as informed peer 6/10 Training on Reanalysis Data, Operational Deployment at ECMWF, and FourCastNet Iterations Brandon and RJ ask how training data is constructed from reanalysis data assimilation and how versions 1 through 3 of FourCastNet evolved. Anima notes early landfall predictions during Hurricane Lee at ECMWF.55:57–1:04:03 · The hosts as informed peer 7/10 Long-Term Climate Rollouts and Spherical Geometry Stability RJ and Brandon closely question how ensemble rollouts maintain physical conservation laws and spherical harmonics without blowing up. Brandon spots polar singularities on the visual rollout, leading Anima to explain why spherical basis geometry stabilizes multi-month runs.1:04:04–1:07:14 · The hosts as informed peer 7/10 Simulating Tokamak Plasma and Digital Twins for Fusion Energy Brandon demonstrates specific physics knowledge by bringing up magnetohydrodynamic (MHD) equations and plasma disruptions in tokamaks. Anima explains how AI digital twins can evaluate both tokamak and stellarator designs.1:07:14–1:10:23 · The hosts as informed peer 5/10 Anima Anandkumar's Career Journey: From Tensor Methods to Principled AI Brandon asks about Anima's career evolution from theoretical tensor methods to applied deep learning. Anima explains why empirical AI is coming full circle back to principled mathematical inductive biases for extrapolation.1:10:24–1:14:20 · The hosts as informed peer 6/10 Physics Foundation Models: Carbon Sequestration, Aerodynamics, and Inverse Design Brandon lists candidate physical PDE systems (electromagnetism, diffusion, heat dissipation). Anima showcases carbon sequestration modeling and topological deformation for car aerodynamics, introducing the concept of general physics foundation models.1:14:21–1:17:23 · The hosts as informed peer 5/10 Multi-Physics Coupling and Inverse Design in Nanotechnology RJ asks if models can transfer to completely unseen physical phenomena. Anima clarifies that unseen physics cannot be invented from zero data, but coupled multi-physics (e.g. thermal-mechanical coupling and inverse lithography) can be learned via modular curricula.1:17:23–1:21:05 · The hosts as informed peer 4/10 Open Source Ecosystem, UN Advisory Board, and AI Policy Brandon and RJ wrap up by discussing the open-source neuraloperator library, Anima's appointment to the UN Scientific Advisory Board, and her call for distinguishing scientific AI from language models in regulation.3:12–6:22 · Guest teaching 4/10 Bridging AI and Science Through Symbolic Verification and TorchLean Brandon asks a solid introductory question linking differential equations, machine learning, and Anima's work on TorchLean. Anima explains her thesis on AI for science, symbolic verification, and neural network safety.6:22–10:57 · Guest teaching 5/10 Formal Proofs, Certified Robustness, and Bounds in TorchLean RJ and Brandon probe how formal verification and bounds work on unconstrained neural networks like transformers. Anima explains certified robustness algorithms like CROWN and acknowledges Lean's current CPU/scalability limitations.10:57–15:50 · Guest teaching 6/10 Physics-Informed Neural Networks vs. Data-Driven Neural Operators Brandon compares physics-informed neural networks (PINNs) to classical numerical analysis discretization errors. Anima educates the hosts on why PINNs fail on chaotic or time-dependent problems and why data-driven neural operators overcome optimization landscape bottlenecks.15:51–19:13 · Guest teaching 6/10 Continuous Function Learning and Resolution Independence RJ asks about function fitting and regularization in under-constrained settings. Anima explains that neural operators represent continuous mappings between infinite-dimensional function spaces, enabling zero-shot super-resolution.19:13–24:34 · Guest teaching 5/10 Fourier Neural Operators and Non-Local Dependency Modeling RJ asks why operating in the Fourier dual domain is advantageous and how architecture differs from standard models. Anima details the balance between quasi-linear computational efficiency and capturing non-local spatial-temporal dependencies.24:34–26:48 · Guest teaching 4/10 Non-Linear Latent Spaces and Frequency Bounds in FNOs Brandon asks about frequency bounds and Nyquist limits in non-linear settings. Anima clarifies that lifting input signals into higher-dimensional non-linear latent spaces allows the network to bypass linear basis truncation constraints.26:48–31:52 · Guest teaching 6/10 Why Transformers Fail on High-Resolution Multi-Dimensional Physics Anima argues transformers are completely intractable for multi-dimensional physics due to context length scaling. RJ pushes back citing vision/video autoregressive codebooks, but Anima firmly rebuts that video models operate at perceptual low-res rather than high-fidelity physics grid requirements.31:53–38:38 · Guest teaching 5/10 FourCastNet: Revolutionizing Weather and Climate Modeling RJ asks about the genesis of FourCastNet and the shift in community thinking. Anima recounts defying skepticism from traditional meteorologists and achieving order-of-magnitude speedups while open-sourcing the model.38:39–44:44 · Guest teaching 5/10 Climate Rollouts, Probabilistic Ensembles, and Rare Physical Events Brandon and RJ question how deterministic rollouts handle long-term climate chaos and rare events. Anima highlights probabilistic ensemble forecasting and notes physical structure enables AI to generalize well on extreme events.44:44–50:16 · Guest teaching 5/10 Multi-Scale Physical Visualizations and Atmospheric Rivers Brandon connects the conversation to AlphaFold and structural biology while RJ asks to review visualizations. Anima presents multi-scale modeling visuals, explaining how soft physical constraints are incorporated into the loss function.50:16–55:57 · Guest teaching 4/10 Training on Reanalysis Data, Operational Deployment at ECMWF, and FourCastNet Iterations Brandon and RJ ask how training data is constructed from reanalysis data assimilation and how versions 1 through 3 of FourCastNet evolved. Anima notes early landfall predictions during Hurricane Lee at ECMWF.55:57–1:04:03 · Guest teaching 6/10 Long-Term Climate Rollouts and Spherical Geometry Stability RJ and Brandon closely question how ensemble rollouts maintain physical conservation laws and spherical harmonics without blowing up. Brandon spots polar singularities on the visual rollout, leading Anima to explain why spherical basis geometry stabilizes multi-month runs.1:04:04–1:07:14 · Guest teaching 5/10 Simulating Tokamak Plasma and Digital Twins for Fusion Energy Brandon demonstrates specific physics knowledge by bringing up magnetohydrodynamic (MHD) equations and plasma disruptions in tokamaks. Anima explains how AI digital twins can evaluate both tokamak and stellarator designs.1:07:14–1:10:23 · Guest teaching 4/10 Anima Anandkumar's Career Journey: From Tensor Methods to Principled AI Brandon asks about Anima's career evolution from theoretical tensor methods to applied deep learning. Anima explains why empirical AI is coming full circle back to principled mathematical inductive biases for extrapolation.1:10:24–1:14:20 · Guest teaching 5/10 Physics Foundation Models: Carbon Sequestration, Aerodynamics, and Inverse Design Brandon lists candidate physical PDE systems (electromagnetism, diffusion, heat dissipation). Anima showcases carbon sequestration modeling and topological deformation for car aerodynamics, introducing the concept of general physics foundation models.1:14:21–1:17:23 · Guest teaching 5/10 Multi-Physics Coupling and Inverse Design in Nanotechnology RJ asks if models can transfer to completely unseen physical phenomena. Anima clarifies that unseen physics cannot be invented from zero data, but coupled multi-physics (e.g. thermal-mechanical coupling and inverse lithography) can be learned via modular curricula.1:17:23–1:21:05 · Guest teaching 4/10 Open Source Ecosystem, UN Advisory Board, and AI Policy Brandon and RJ wrap up by discussing the open-source neuraloperator library, Anima's appointment to the UN Scientific Advisory Board, and her call for distinguishing scientific AI from language models in regulation.3:12–6:22 · Guest disagreement 1/10 Bridging AI and Science Through Symbolic Verification and TorchLean Brandon asks a solid introductory question linking differential equations, machine learning, and Anima's work on TorchLean. Anima explains her thesis on AI for science, symbolic verification, and neural network safety.6:22–10:57 · Guest disagreement 2/10 Formal Proofs, Certified Robustness, and Bounds in TorchLean RJ and Brandon probe how formal verification and bounds work on unconstrained neural networks like transformers. Anima explains certified robustness algorithms like CROWN and acknowledges Lean's current CPU/scalability limitations.10:57–15:50 · Guest disagreement 2/10 Physics-Informed Neural Networks vs. Data-Driven Neural Operators Brandon compares physics-informed neural networks (PINNs) to classical numerical analysis discretization errors. Anima educates the hosts on why PINNs fail on chaotic or time-dependent problems and why data-driven neural operators overcome optimization landscape bottlenecks.15:51–19:13 · Guest disagreement 1/10 Continuous Function Learning and Resolution Independence RJ asks about function fitting and regularization in under-constrained settings. Anima explains that neural operators represent continuous mappings between infinite-dimensional function spaces, enabling zero-shot super-resolution.19:13–24:34 · Guest disagreement 1/10 Fourier Neural Operators and Non-Local Dependency Modeling RJ asks why operating in the Fourier dual domain is advantageous and how architecture differs from standard models. Anima details the balance between quasi-linear computational efficiency and capturing non-local spatial-temporal dependencies.24:34–26:48 · Guest disagreement 1/10 Non-Linear Latent Spaces and Frequency Bounds in FNOs Brandon asks about frequency bounds and Nyquist limits in non-linear settings. Anima clarifies that lifting input signals into higher-dimensional non-linear latent spaces allows the network to bypass linear basis truncation constraints.26:48–31:52 · Guest disagreement 5/10 Why Transformers Fail on High-Resolution Multi-Dimensional Physics Anima argues transformers are completely intractable for multi-dimensional physics due to context length scaling. RJ pushes back citing vision/video autoregressive codebooks, but Anima firmly rebuts that video models operate at perceptual low-res rather than high-fidelity physics grid requirements.31:53–38:38 · Guest disagreement 2/10 FourCastNet: Revolutionizing Weather and Climate Modeling RJ asks about the genesis of FourCastNet and the shift in community thinking. Anima recounts defying skepticism from traditional meteorologists and achieving order-of-magnitude speedups while open-sourcing the model.38:39–44:44 · Guest disagreement 1/10 Climate Rollouts, Probabilistic Ensembles, and Rare Physical Events Brandon and RJ question how deterministic rollouts handle long-term climate chaos and rare events. Anima highlights probabilistic ensemble forecasting and notes physical structure enables AI to generalize well on extreme events.44:44–50:16 · Guest disagreement 1/10 Multi-Scale Physical Visualizations and Atmospheric Rivers Brandon connects the conversation to AlphaFold and structural biology while RJ asks to review visualizations. Anima presents multi-scale modeling visuals, explaining how soft physical constraints are incorporated into the loss function.50:16–55:57 · Guest disagreement 1/10 Training on Reanalysis Data, Operational Deployment at ECMWF, and FourCastNet Iterations Brandon and RJ ask how training data is constructed from reanalysis data assimilation and how versions 1 through 3 of FourCastNet evolved. Anima notes early landfall predictions during Hurricane Lee at ECMWF.55:57–1:04:03 · Guest disagreement 3/10 Long-Term Climate Rollouts and Spherical Geometry Stability RJ and Brandon closely question how ensemble rollouts maintain physical conservation laws and spherical harmonics without blowing up. Brandon spots polar singularities on the visual rollout, leading Anima to explain why spherical basis geometry stabilizes multi-month runs.1:04:04–1:07:14 · Guest disagreement 1/10 Simulating Tokamak Plasma and Digital Twins for Fusion Energy Brandon demonstrates specific physics knowledge by bringing up magnetohydrodynamic (MHD) equations and plasma disruptions in tokamaks. Anima explains how AI digital twins can evaluate both tokamak and stellarator designs.1:07:14–1:10:23 · Guest disagreement 1/10 Anima Anandkumar's Career Journey: From Tensor Methods to Principled AI Brandon asks about Anima's career evolution from theoretical tensor methods to applied deep learning. Anima explains why empirical AI is coming full circle back to principled mathematical inductive biases for extrapolation.1:10:24–1:14:20 · Guest disagreement 1/10 Physics Foundation Models: Carbon Sequestration, Aerodynamics, and Inverse Design Brandon lists candidate physical PDE systems (electromagnetism, diffusion, heat dissipation). Anima showcases carbon sequestration modeling and topological deformation for car aerodynamics, introducing the concept of general physics foundation models.1:14:21–1:17:23 · Guest disagreement 2/10 Multi-Physics Coupling and Inverse Design in Nanotechnology RJ asks if models can transfer to completely unseen physical phenomena. Anima clarifies that unseen physics cannot be invented from zero data, but coupled multi-physics (e.g. thermal-mechanical coupling and inverse lithography) can be learned via modular curricula.1:17:23–1:21:05 · Guest disagreement 1/10 Open Source Ecosystem, UN Advisory Board, and AI Policy Brandon and RJ wrap up by discussing the open-source neuraloperator library, Anima's appointment to the UN Scientific Advisory Board, and her call for distinguishing scientific AI from language models in regulation.3:12–6:22 · The hosts pushing back 1/10 Bridging AI and Science Through Symbolic Verification and TorchLean Brandon asks a solid introductory question linking differential equations, machine learning, and Anima's work on TorchLean. Anima explains her thesis on AI for science, symbolic verification, and neural network safety.6:22–10:57 · The hosts pushing back 3/10 Formal Proofs, Certified Robustness, and Bounds in TorchLean RJ and Brandon probe how formal verification and bounds work on unconstrained neural networks like transformers. Anima explains certified robustness algorithms like CROWN and acknowledges Lean's current CPU/scalability limitations.10:57–15:50 · The hosts pushing back 2/10 Physics-Informed Neural Networks vs. Data-Driven Neural Operators Brandon compares physics-informed neural networks (PINNs) to classical numerical analysis discretization errors. Anima educates the hosts on why PINNs fail on chaotic or time-dependent problems and why data-driven neural operators overcome optimization landscape bottlenecks.15:51–19:13 · The hosts pushing back 2/10 Continuous Function Learning and Resolution Independence RJ asks about function fitting and regularization in under-constrained settings. Anima explains that neural operators represent continuous mappings between infinite-dimensional function spaces, enabling zero-shot super-resolution.19:13–24:34 · The hosts pushing back 2/10 Fourier Neural Operators and Non-Local Dependency Modeling RJ asks why operating in the Fourier dual domain is advantageous and how architecture differs from standard models. Anima details the balance between quasi-linear computational efficiency and capturing non-local spatial-temporal dependencies.24:34–26:48 · The hosts pushing back 1/10 Non-Linear Latent Spaces and Frequency Bounds in FNOs Brandon asks about frequency bounds and Nyquist limits in non-linear settings. Anima clarifies that lifting input signals into higher-dimensional non-linear latent spaces allows the network to bypass linear basis truncation constraints.26:48–31:52 · The hosts pushing back 6/10 Why Transformers Fail on High-Resolution Multi-Dimensional Physics Anima argues transformers are completely intractable for multi-dimensional physics due to context length scaling. RJ pushes back citing vision/video autoregressive codebooks, but Anima firmly rebuts that video models operate at perceptual low-res rather than high-fidelity physics grid requirements.31:53–38:38 · The hosts pushing back 2/10 FourCastNet: Revolutionizing Weather and Climate Modeling RJ asks about the genesis of FourCastNet and the shift in community thinking. Anima recounts defying skepticism from traditional meteorologists and achieving order-of-magnitude speedups while open-sourcing the model.38:39–44:44 · The hosts pushing back 2/10 Climate Rollouts, Probabilistic Ensembles, and Rare Physical Events Brandon and RJ question how deterministic rollouts handle long-term climate chaos and rare events. Anima highlights probabilistic ensemble forecasting and notes physical structure enables AI to generalize well on extreme events.44:44–50:16 · The hosts pushing back 1/10 Multi-Scale Physical Visualizations and Atmospheric Rivers Brandon connects the conversation to AlphaFold and structural biology while RJ asks to review visualizations. Anima presents multi-scale modeling visuals, explaining how soft physical constraints are incorporated into the loss function.50:16–55:57 · The hosts pushing back 2/10 Training on Reanalysis Data, Operational Deployment at ECMWF, and FourCastNet Iterations Brandon and RJ ask how training data is constructed from reanalysis data assimilation and how versions 1 through 3 of FourCastNet evolved. Anima notes early landfall predictions during Hurricane Lee at ECMWF.55:57–1:04:03 · The hosts pushing back 4/10 Long-Term Climate Rollouts and Spherical Geometry Stability RJ and Brandon closely question how ensemble rollouts maintain physical conservation laws and spherical harmonics without blowing up. Brandon spots polar singularities on the visual rollout, leading Anima to explain why spherical basis geometry stabilizes multi-month runs.1:04:04–1:07:14 · The hosts pushing back 2/10 Simulating Tokamak Plasma and Digital Twins for Fusion Energy Brandon demonstrates specific physics knowledge by bringing up magnetohydrodynamic (MHD) equations and plasma disruptions in tokamaks. Anima explains how AI digital twins can evaluate both tokamak and stellarator designs.1:07:14–1:10:23 · The hosts pushing back 1/10 Anima Anandkumar's Career Journey: From Tensor Methods to Principled AI Brandon asks about Anima's career evolution from theoretical tensor methods to applied deep learning. Anima explains why empirical AI is coming full circle back to principled mathematical inductive biases for extrapolation.1:10:24–1:14:20 · The hosts pushing back 1/10 Physics Foundation Models: Carbon Sequestration, Aerodynamics, and Inverse Design Brandon lists candidate physical PDE systems (electromagnetism, diffusion, heat dissipation). Anima showcases carbon sequestration modeling and topological deformation for car aerodynamics, introducing the concept of general physics foundation models.1:14:21–1:17:23 · The hosts pushing back 2/10 Multi-Physics Coupling and Inverse Design in Nanotechnology RJ asks if models can transfer to completely unseen physical phenomena. Anima clarifies that unseen physics cannot be invented from zero data, but coupled multi-physics (e.g. thermal-mechanical coupling and inverse lithography) can be learned via modular curricula.1:17:23–1:21:05 · The hosts pushing back 1/10 Open Source Ecosystem, UN Advisory Board, and AI Policy Brandon and RJ wrap up by discussing the open-source neuraloperator library, Anima's appointment to the UN Scientific Advisory Board, and her call for distinguishing scientific AI from language models in regulation.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 100%1:03:00 · the hosts 0% · guest 100%1:03:00 · the hosts 0% · guest 100%1:06:00 · the hosts 0% · guest 100%1:06:00 · the hosts 0% · guest 100%1:09:00 · the hosts 0% · guest 100%1:09:00 · the hosts 0% · guest 100%1:12:00 · the hosts 0% · guest 100%1:12:00 · the hosts 0% · guest 100%1:15:00 · the hosts 0% · guest 100%1:15:00 · the hosts 0% · guest 100%1:18:00 · the hosts 0% · guest 100%1:18:00 · the hosts 0% · guest 100%1:21:00 · the hosts 0% · guest 100%1:21:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 29:54 Anima rejects RJ's video model comparison

When RJ suggests autoregressive video and vision models solve multi-dimensional physical representations, Anima firmly interrupts to point out their resolution is far too low and superficial for real physical simulations.

Hardest push from the hosts ▶ 29:54 RJ challenges transformer limitations using video models

RJ pushes back against Anima's assertion that transformers cannot scale to physical dimensions by citing autoregressive video codebooks and latent compression techniques.

Biggest teaching moment ▶ 12:03 Anima explains PINN failures in turbulent time-dependent regimes

Anima educates Brandon on why physics-informed neural networks break down when solving partial differential equations from scratch on chaotic, turbulent systems.

The host holds their own ▶ 1:04:39 Brandon identifies MHD equations and plasma disruption mechanisms

Brandon demonstrates deep domain expertise in fusion physics by immediately identifying magnetohydrodynamic equations and detailing the containment risk of plasma disruptions in tokamaks.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Bridging AI and Science Through Symbolic Verification and TorchLean 5411 Brandon asks a solid introductory question linking differential equations, machine learning, and Anima's work on TorchLean. Anima explains her thesis on AI for science, symbolic verification, and neural network safety.
Formal Proofs, Certified Robustness, and Bounds in TorchLean 6523 RJ and Brandon probe how formal verification and bounds work on unconstrained neural networks like transformers. Anima explains certified robustness algorithms like CROWN and acknowledges Lean's current CPU/scalability limitations.
Physics-Informed Neural Networks vs. Data-Driven Neural Operators 6622 Brandon compares physics-informed neural networks (PINNs) to classical numerical analysis discretization errors. Anima educates the hosts on why PINNs fail on chaotic or time-dependent problems and why data-driven neural operators overcome optimization landscape bottlenecks.
Continuous Function Learning and Resolution Independence 5612 RJ asks about function fitting and regularization in under-constrained settings. Anima explains that neural operators represent continuous mappings between infinite-dimensional function spaces, enabling zero-shot super-resolution.
Fourier Neural Operators and Non-Local Dependency Modeling 6512 RJ asks why operating in the Fourier dual domain is advantageous and how architecture differs from standard models. Anima details the balance between quasi-linear computational efficiency and capturing non-local spatial-temporal dependencies.
Non-Linear Latent Spaces and Frequency Bounds in FNOs 6411 Brandon asks about frequency bounds and Nyquist limits in non-linear settings. Anima clarifies that lifting input signals into higher-dimensional non-linear latent spaces allows the network to bypass linear basis truncation constraints.
Why Transformers Fail on High-Resolution Multi-Dimensional Physics 7656 Anima argues transformers are completely intractable for multi-dimensional physics due to context length scaling. RJ pushes back citing vision/video autoregressive codebooks, but Anima firmly rebuts that video models operate at perceptual low-res rather than high-fidelity physics grid requirements.
FourCastNet: Revolutionizing Weather and Climate Modeling 5522 RJ asks about the genesis of FourCastNet and the shift in community thinking. Anima recounts defying skepticism from traditional meteorologists and achieving order-of-magnitude speedups while open-sourcing the model.
Climate Rollouts, Probabilistic Ensembles, and Rare Physical Events 6512 Brandon and RJ question how deterministic rollouts handle long-term climate chaos and rare events. Anima highlights probabilistic ensemble forecasting and notes physical structure enables AI to generalize well on extreme events.
Multi-Scale Physical Visualizations and Atmospheric Rivers 5511 Brandon connects the conversation to AlphaFold and structural biology while RJ asks to review visualizations. Anima presents multi-scale modeling visuals, explaining how soft physical constraints are incorporated into the loss function.
Training on Reanalysis Data, Operational Deployment at ECMWF, and FourCastNet Iterations 6412 Brandon and RJ ask how training data is constructed from reanalysis data assimilation and how versions 1 through 3 of FourCastNet evolved. Anima notes early landfall predictions during Hurricane Lee at ECMWF.
Long-Term Climate Rollouts and Spherical Geometry Stability 7634 RJ and Brandon closely question how ensemble rollouts maintain physical conservation laws and spherical harmonics without blowing up. Brandon spots polar singularities on the visual rollout, leading Anima to explain why spherical basis geometry stabilizes multi-month runs.
Simulating Tokamak Plasma and Digital Twins for Fusion Energy 7512 Brandon demonstrates specific physics knowledge by bringing up magnetohydrodynamic (MHD) equations and plasma disruptions in tokamaks. Anima explains how AI digital twins can evaluate both tokamak and stellarator designs.
Anima Anandkumar's Career Journey: From Tensor Methods to Principled AI 5411 Brandon asks about Anima's career evolution from theoretical tensor methods to applied deep learning. Anima explains why empirical AI is coming full circle back to principled mathematical inductive biases for extrapolation.
Physics Foundation Models: Carbon Sequestration, Aerodynamics, and Inverse Design 6511 Brandon lists candidate physical PDE systems (electromagnetism, diffusion, heat dissipation). Anima showcases carbon sequestration modeling and topological deformation for car aerodynamics, introducing the concept of general physics foundation models.
Multi-Physics Coupling and Inverse Design in Nanotechnology 5522 RJ asks if models can transfer to completely unseen physical phenomena. Anima clarifies that unseen physics cannot be invented from zero data, but coupled multi-physics (e.g. thermal-mechanical coupling and inverse lithography) can be learned via modular curricula.
Open Source Ecosystem, UN Advisory Board, and AI Policy 4411 Brandon and RJ wrap up by discussing the open-source neuraloperator library, Anima's appointment to the UN Scientific Advisory Board, and her call for distinguishing scientific AI from language models in regulation.

Statements from this episode (28)

Insight
Anandkumar: Scientific AI Bottleneck Is Real-World Testing, Not Hypothesis Generation
“Yes, you can do a lot of hypothesis generation. You can have ideas, but ideas are not enough, right? So you can have a lot of ideas. The bottleneck is going, testing, and verifying that they work in the real world.”
Anima Anandkumar Aug 26, 2026 ▶ 4:20
Assertion Supported
TorchLean enables defining PyTorch-like neural networks directly in Lean
“So what it really enables is that you can now write neural networks essentially in Lean. So instead of writing in, like, PyTorch, it's like a PyTorch-like abstraction, but you can, like, kind of, you know, write it in Lean, and so it can be fully formalized in…”
Anima Anandkumar Aug 26, 2026 ▶ 6:45
Assertion Supported
Lean faces CPU-bound scalability limits for verifying large neural networks
“So lean still has a lot of shortcomings there. It's CPU based and you know, it's not, Like, getting that onto the GPU has a lot of nuances there. So, you know, a lot of work needs to be done. So what we've started with is a framework, you know, making that mor…”
Anima Anandkumar Aug 26, 2026 ▶ 10:31
Insight
Physics-Informed Neural Networks fail on chaotic, time-dependent differential equations
“Optimization ends up being usually very difficult, especially for problems that are time-dependent, meaning it's not just stationary, you also have time, and the time component in many cases could be turbulent, like in the case of fluid dynamics, you know, you…”
Anima Anandkumar Aug 26, 2026 ▶ 12:24
Insight
Neural operators overcome PINN limitations by combining data with physics constraints
“Our idea of neural operators came as a way to overcome this, right? So saying, you know, we can't rely just on physics constraints alone to come up with answers. We have lots of data available. You know, I'll talk about the weather example where we even collec…”
Anima Anandkumar Aug 26, 2026 ▶ 13:25
Insight
Neural operators can evaluate at arbitrary resolutions at inference time
“Neural operators enable because they model inputs and outputs as continuous functions that can be infinitely resolved, that can have infinite discretization. And now we can have You know, at inference time, you can give it now inputs and ask for outputs at any…”
Anima Anandkumar Aug 26, 2026 ▶ 17:26
Insight
Physics constraints enable neural operators to exceed training data resolution
“If you now give it the model additional information in terms of, let's say, a physical loss, so you could give it partial differential equation constraints, conservation laws, and you can now enforce them at a finer resolution than the data you have, then ther…”
Anima Anandkumar Aug 26, 2026 ▶ 18:34
Insight
Fourier neural operators scale quasi-linearly, avoiding transformers' quadratic complexity
“If we were to use transformers and we require a very high resolution, it would become untenable because of the quadratic complexity and all to all connections. On the other hand, if you did that with Fourier transforms, we have like quasi linear complexity and…”
Anima Anandkumar Aug 26, 2026 ▶ 21:51
Insight
Non-linearities enable neural operators to expressively capture limited frequency modes
“So that's one way of thinking because, you know, first of all, we are lifting the signal to more dimensions, even if the signal is two or three dimensions we are now lifting it to much higher dimension. So in that space, the idea is it's easier to learn and we…”
Anima Anandkumar Aug 26, 2026 ▶ 26:08
Assertion Supported
Anandkumar's AI weather model trained on 50,000 global weather maps
“You know, our weather model, like, had about, like, 50,000 samples, right? 50,000 samples of fairly high resolution, like, world global weather maps, but it's nothing like what we see with language.”
Anima Anandkumar Aug 26, 2026 ▶ 28:29
Prediction Not checkable as stated
Transformers will never scale to high-resolution 4D physics simulations
“So forget ever having a transformer for anything of this scale. All of the world's compute will not be enough. And first of all, they all have to be co-located to be able to ever do this. So that's why we need other architectures.”
Anima Anandkumar Aug 26, 2026 ▶ 29:41
Assertion Supported
FourCastNet matches supercomputer weather accuracy 10,000 times faster on consumer GPUs
“To our surprise, we found that it's not only, you know, accurate, it's almost as close to the what the traditional weather models can do accurately, but also tens of thousands of times faster. So what would take a big supercomputer to run can now be run. And w…”
Anima Anandkumar Aug 26, 2026 ▶ 33:56
Assertion Supported
FourCastNet was the first AI weather model to be permissively open-sourced
“We were the first to actually open source our weather model for CastNet and do it permissively.”
Anima Anandkumar Aug 26, 2026 ▶ 34:34
Assertion Contradicted
Neural operators are the only AI architecture that works for climate emulation
“This is where the Allen AI Institute has now built climate models based on our neural operator architecture. And that's the only one that works As an AI emulator, right? None of the other architectures work for climate because climate requires us to assume the…”
Anima Anandkumar Aug 26, 2026 ▶ 37:44
Insight
Structured physical signatures enable AI to predict rare events with fewer samples
“But I think this is where more broadly the lesson is the physical world may be more forgiving because, you know, where there are extreme events like hurricanes that have very specific physical signature. Right. So it's like extreme, but in a very specific way.…”
Anima Anandkumar Aug 26, 2026 ▶ 43:23
Assertion Supported
AI models predict fusion reactor plasma disruption one million times faster
“You know, I talk about plasma and fusion reactor. You know, we barely have a few thousand samples, but we are able to accurately predict events like disruption very well. And we are able to do that a million times faster than what traditional simulations were …”
Anima Anandkumar Aug 26, 2026 ▶ 43:50
Insight
Enforcing physics as hard constraints in AI models is computationally intractable
“Making it a hard constraint is not tractable, whereas adding it as a loss function. And of course, there's still the balancing of that loss with the data we have.”
Anima Anandkumar Aug 26, 2026 ▶ 49:16
Assertion Supported
Neural operators accurately model non-local phenomena like atmospheric rivers
“These are like thousands of miles wide, so you really need non-local models that capture these very large span phenomena and do that accurately, and that's what our neural operators are able to do.”
Anima Anandkumar Aug 26, 2026 ▶ 49:56
Assertion Supported
FourCastNet predicted Hurricane Lee's landfall days earlier than standard weather models
“For instance, our forecast net was able to correctly predict that the hurricane making the landfall several days earlier compared to the standard weather forecasting models.”
Anima Anandkumar Aug 26, 2026 ▶ 52:28
Assertion Supported
Spherical AI weather models achieve longer autoregressive rollouts than flat models
“These models that we have are able to do the longest rollouts compared to any of the other weather models that completely ignore spherical assumption and a range of other things.”
Anima Anandkumar Aug 26, 2026 ▶ 57:12
Assertion Supported
FourCastNet models trained on six-hour steps produce stable months-long forecasts
“To predict for the next six hours. And a little bit of multi-step fine tuning... Now we are showing for several months that it's able to do that.”
Anima Anandkumar Aug 26, 2026 ▶ 1:00:15
Insight
Purely data-driven AI approaches are saturating in physical scientific discovery
“A lot of data, purely data driven approaches in a way seeing saturation, right? So now we want to ask, okay, either make them more hardware efficient, right? There's a lot of now room to kind of say, can we now, you know, make them much more energy efficient o…”
Anima Anandkumar Aug 26, 2026 ▶ 1:09:23
Assertion Supported
Neural operators simulate underground carbon sequestration faster than traditional methods
“So this was like, you know, being able to ask, can we sequester carbon dioxide underground and model how carbon dioxide Expands or, you know, what is the pressure buildup in these reservoirs? And, you know, can we kind of model how they migrate over several de…”
Anima Anandkumar Aug 26, 2026 ▶ 1:11:15
Insight
Latent space physics models enable generalization across diverse geometric shapes
“So the idea of like a latent space to handle all kinds of different geometries and be able to capture the physics there in the latent space well, means we can now have a model that generalizes across a lot of different geometries.”
Anima Anandkumar Aug 26, 2026 ▶ 1:12:10
Assertion Not checkable as stated
AI has foundation models for language and vision but not physics
“Because we have foundation models for language, maybe vision, but not for physics. So, you know, the idea is instead of like right now what we've seen are narrow surrogates and we're trying to broaden their scope more and more, but ideally we have much broader…”
Anima Anandkumar Aug 26, 2026 ▶ 1:12:57
Insight
Curriculum learning enables multi-physics AI fine-tuning with significantly fewer samples
“And now there's coupling, like because of heat, there's also stretching or kind of the joint phenomena. You could like now hope to fine tune with much fewer samples because it kind of individually knows this phenomena. Then combining them together, maybe it ca…”
Anima Anandkumar Aug 26, 2026 ▶ 1:15:09
Insight
Simulation-in-the-loop AI outperforms human intuition at highly nonlinear inverse design
“And humans are usually not good at this, right? We are not good at like looking at highly nonlinear phenomena and say, oh, somehow maybe this combination of all these gates coming together helps pull the electrons together in a quantum gate. And so our collabo…”
Anima Anandkumar Aug 26, 2026 ▶ 1:16:37
Opinion
Regulating AI for science like large language models creates serious problems
“A lot of regulatory frameworks equate AI with language models and Yes, language models can, you know, manipulate people, can have all these kinds of harmful impacts that we should think about controlling, but AI for science is different. So I think this one si…”
Anima Anandkumar Aug 26, 2026 ▶ 1:20:19
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.