Sep 4, 2026 · 27m · latent-space

Faster Chips That Don't Melt — Anima Anandkumar & Benedikt Jenik, Accelerated Understanding

Benedikt Jenik · 12m spoken Anima Anandkumar · 8m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of the Latent Space podcast, Dr. Anima Anandkumar and Benedikt Jenik introduce Accelerated Understanding, a startup building universal foundation models for the physical sciences. They discuss how neural operators, governing physical invariants, and trillion-token 4D computing architectures transform simulation for semiconductor design and clean energy.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 4.0 Guest teaching 4.9 Guest disagreement 0.4 The hosts pushing back 1.1
05100:0010:0020:000:55–4:44 · The hosts as informed peer 5/10 The Vision for Universal Physical Foundation Models RJ questions whether physical foundation models can actually transfer across drastically different domains like catheters and fusion reactors. Anima explains how fundamental principles such as Reynolds numbers, causality, and conservation laws provide universal mathematical commonalities across PDEs.4:45–8:42 · The hosts as informed peer 2/10 4D Spatiotemporal Scaling and Multi-Physics Model Convergence RJ asks for a progress update on building the model. Benedikt and Anima deliver an in-depth technical explanation of 4D spatiotemporal scaling, non-autoregressive rollouts, novel sharding strategies, and emergent cross-physics learning.8:43–12:23 · The hosts as informed peer 6/10 Neural Operators, Resolution Invariance, and Transformer Limitations Brandon asks a technically nuanced question regarding Fourier neural operators, inductive biases, and geometric transfer. Anima explains why neural operators enable resolution invariance while standard transformer architectures fail at multi-trillion token scales due to quadratic complexity.12:23–15:00 · The hosts as informed peer 4/10 Extreme Trillion-Token Infrastructure and Hardware Constraints RJ asks how the team technically reaches trillion-token context lengths across algorithms and infrastructure. Benedikt details the exact 4D memory arithmetic requiring 22 terabytes in accelerator memory and specialized interconnect sharding.15:01–18:31 · The hosts as informed peer 5/10 Synthetic Data, Curriculum Learning, and Dense Self-Improvement Brandon probes how the team balances numerical simulation data with real physical data to bridge the sim-to-real gap. Benedikt and Anima explain curriculum engineering with numerical simulators and highlight how dense physics-based loss feedback outperforms sparse human feedback in LLMs.18:32–24:54 · The hosts as informed peer 5/10 Multi-Scale Physical Modeling via Dynamic Latent Mixing Brandon and RJ explore multi-scale modeling and commercial use cases in semiconductor design and geothermal energy. RJ pushes to verify whether empirical performance transfer actually manifests in commercial deployments.24:54–25:52 · The hosts as informed peer 1/10 Company Scaling, Recruitment, and Enterprise Roadmap The hosts wrap up the conversation with casual questions regarding company hiring and Anima's inclusion on the Time 100 list.0:55–4:44 · Guest teaching 5/10 The Vision for Universal Physical Foundation Models RJ questions whether physical foundation models can actually transfer across drastically different domains like catheters and fusion reactors. Anima explains how fundamental principles such as Reynolds numbers, causality, and conservation laws provide universal mathematical commonalities across PDEs.4:45–8:42 · Guest teaching 6/10 4D Spatiotemporal Scaling and Multi-Physics Model Convergence RJ asks for a progress update on building the model. Benedikt and Anima deliver an in-depth technical explanation of 4D spatiotemporal scaling, non-autoregressive rollouts, novel sharding strategies, and emergent cross-physics learning.8:43–12:23 · Guest teaching 6/10 Neural Operators, Resolution Invariance, and Transformer Limitations Brandon asks a technically nuanced question regarding Fourier neural operators, inductive biases, and geometric transfer. Anima explains why neural operators enable resolution invariance while standard transformer architectures fail at multi-trillion token scales due to quadratic complexity.12:23–15:00 · Guest teaching 6/10 Extreme Trillion-Token Infrastructure and Hardware Constraints RJ asks how the team technically reaches trillion-token context lengths across algorithms and infrastructure. Benedikt details the exact 4D memory arithmetic requiring 22 terabytes in accelerator memory and specialized interconnect sharding.15:01–18:31 · Guest teaching 6/10 Synthetic Data, Curriculum Learning, and Dense Self-Improvement Brandon probes how the team balances numerical simulation data with real physical data to bridge the sim-to-real gap. Benedikt and Anima explain curriculum engineering with numerical simulators and highlight how dense physics-based loss feedback outperforms sparse human feedback in LLMs.18:32–24:54 · Guest teaching 4/10 Multi-Scale Physical Modeling via Dynamic Latent Mixing Brandon and RJ explore multi-scale modeling and commercial use cases in semiconductor design and geothermal energy. RJ pushes to verify whether empirical performance transfer actually manifests in commercial deployments.24:54–25:52 · Guest teaching 1/10 Company Scaling, Recruitment, and Enterprise Roadmap The hosts wrap up the conversation with casual questions regarding company hiring and Anima's inclusion on the Time 100 list.0:55–4:44 · Guest disagreement 1/10 The Vision for Universal Physical Foundation Models RJ questions whether physical foundation models can actually transfer across drastically different domains like catheters and fusion reactors. Anima explains how fundamental principles such as Reynolds numbers, causality, and conservation laws provide universal mathematical commonalities across PDEs.4:45–8:42 · Guest disagreement 0/10 4D Spatiotemporal Scaling and Multi-Physics Model Convergence RJ asks for a progress update on building the model. Benedikt and Anima deliver an in-depth technical explanation of 4D spatiotemporal scaling, non-autoregressive rollouts, novel sharding strategies, and emergent cross-physics learning.8:43–12:23 · Guest disagreement 1/10 Neural Operators, Resolution Invariance, and Transformer Limitations Brandon asks a technically nuanced question regarding Fourier neural operators, inductive biases, and geometric transfer. Anima explains why neural operators enable resolution invariance while standard transformer architectures fail at multi-trillion token scales due to quadratic complexity.12:23–15:00 · Guest disagreement 0/10 Extreme Trillion-Token Infrastructure and Hardware Constraints RJ asks how the team technically reaches trillion-token context lengths across algorithms and infrastructure. Benedikt details the exact 4D memory arithmetic requiring 22 terabytes in accelerator memory and specialized interconnect sharding.15:01–18:31 · Guest disagreement 0/10 Synthetic Data, Curriculum Learning, and Dense Self-Improvement Brandon probes how the team balances numerical simulation data with real physical data to bridge the sim-to-real gap. Benedikt and Anima explain curriculum engineering with numerical simulators and highlight how dense physics-based loss feedback outperforms sparse human feedback in LLMs.18:32–24:54 · Guest disagreement 1/10 Multi-Scale Physical Modeling via Dynamic Latent Mixing Brandon and RJ explore multi-scale modeling and commercial use cases in semiconductor design and geothermal energy. RJ pushes to verify whether empirical performance transfer actually manifests in commercial deployments.24:54–25:52 · Guest disagreement 0/10 Company Scaling, Recruitment, and Enterprise Roadmap The hosts wrap up the conversation with casual questions regarding company hiring and Anima's inclusion on the Time 100 list.0:55–4:44 · The hosts pushing back 2/10 The Vision for Universal Physical Foundation Models RJ questions whether physical foundation models can actually transfer across drastically different domains like catheters and fusion reactors. Anima explains how fundamental principles such as Reynolds numbers, causality, and conservation laws provide universal mathematical commonalities across PDEs.4:45–8:42 · The hosts pushing back 0/10 4D Spatiotemporal Scaling and Multi-Physics Model Convergence RJ asks for a progress update on building the model. Benedikt and Anima deliver an in-depth technical explanation of 4D spatiotemporal scaling, non-autoregressive rollouts, novel sharding strategies, and emergent cross-physics learning.8:43–12:23 · The hosts pushing back 2/10 Neural Operators, Resolution Invariance, and Transformer Limitations Brandon asks a technically nuanced question regarding Fourier neural operators, inductive biases, and geometric transfer. Anima explains why neural operators enable resolution invariance while standard transformer architectures fail at multi-trillion token scales due to quadratic complexity.12:23–15:00 · The hosts pushing back 1/10 Extreme Trillion-Token Infrastructure and Hardware Constraints RJ asks how the team technically reaches trillion-token context lengths across algorithms and infrastructure. Benedikt details the exact 4D memory arithmetic requiring 22 terabytes in accelerator memory and specialized interconnect sharding.15:01–18:31 · The hosts pushing back 0/10 Synthetic Data, Curriculum Learning, and Dense Self-Improvement Brandon probes how the team balances numerical simulation data with real physical data to bridge the sim-to-real gap. Benedikt and Anima explain curriculum engineering with numerical simulators and highlight how dense physics-based loss feedback outperforms sparse human feedback in LLMs.18:32–24:54 · The hosts pushing back 3/10 Multi-Scale Physical Modeling via Dynamic Latent Mixing Brandon and RJ explore multi-scale modeling and commercial use cases in semiconductor design and geothermal energy. RJ pushes to verify whether empirical performance transfer actually manifests in commercial deployments.24:54–25:52 · The hosts pushing back 0/10 Company Scaling, Recruitment, and Enterprise Roadmap The hosts wrap up the conversation with casual questions regarding company hiring and Anima's inclusion on the Time 100 list.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 2:43 Dismissing pure data-driven AI for science

Anima firmly reframes the discussion, arguing that purely data-driven AI is fundamentally insufficient for physical discovery because internet-scale training data does not exist for novel science.

Hardest push from the hosts ▶ 24:10 Demanding commercial proof of transfer

RJ directly pushes back on Benedikt's theoretical claims, demanding to know if empirical multi-physics transfer actually holds up in real-world commercial deployments.

Biggest teaching moment ▶ 10:30 Explaining why transformers fail on physical simulation

Anima educates the hosts on the structural constraints of standard architectures, explaining why quadratic transformer complexity makes multi-trillion token physical rollouts impossible without neural operators.

The host holds their own ▶ 8:42 Brandon drills into geometric inductive biases

Brandon demonstrates strong technical proficiency by questioning how Fourier neural operators maintain transferability across varying geometric domains and PDE classes.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
The Vision for Universal Physical Foundation Models 5512 RJ questions whether physical foundation models can actually transfer across drastically different domains like catheters and fusion reactors. Anima explains how fundamental principles such as Reynolds numbers, causality, and conservation laws provide universal mathematical commonalities across PDEs.
4D Spatiotemporal Scaling and Multi-Physics Model Convergence 2600 RJ asks for a progress update on building the model. Benedikt and Anima deliver an in-depth technical explanation of 4D spatiotemporal scaling, non-autoregressive rollouts, novel sharding strategies, and emergent cross-physics learning.
Neural Operators, Resolution Invariance, and Transformer Limitations 6612 Brandon asks a technically nuanced question regarding Fourier neural operators, inductive biases, and geometric transfer. Anima explains why neural operators enable resolution invariance while standard transformer architectures fail at multi-trillion token scales due to quadratic complexity.
Extreme Trillion-Token Infrastructure and Hardware Constraints 4601 RJ asks how the team technically reaches trillion-token context lengths across algorithms and infrastructure. Benedikt details the exact 4D memory arithmetic requiring 22 terabytes in accelerator memory and specialized interconnect sharding.
Synthetic Data, Curriculum Learning, and Dense Self-Improvement 5600 Brandon probes how the team balances numerical simulation data with real physical data to bridge the sim-to-real gap. Benedikt and Anima explain curriculum engineering with numerical simulators and highlight how dense physics-based loss feedback outperforms sparse human feedback in LLMs.
Multi-Scale Physical Modeling via Dynamic Latent Mixing 5413 Brandon and RJ explore multi-scale modeling and commercial use cases in semiconductor design and geothermal energy. RJ pushes to verify whether empirical performance transfer actually manifests in commercial deployments.
Company Scaling, Recruitment, and Enterprise Roadmap 1100 The hosts wrap up the conversation with casual questions regarding company hiring and Anima's inclusion on the Time 100 list.

Statements from this episode (16)

Disclosure
Anandkumar: Accelerated Understanding is building universal foundation models for physical simulation
“You know, that universality and scale that we've seen play out for language, what is that counterpart for the physical world? And that's the bet that accelerated understanding is making.”
Anima Anandkumar Sep 4, 2026 ▶ 1:52
Insight
Anandkumar: Purely data-driven AI fails at physical simulation without physics laws
“And so this reliance on just purely data-driven AI is not going to be enough. And that's where, you know, adding the loss of physics is really critical.”
Anima Anandkumar Sep 4, 2026 ▶ 3:07
Insight
Anandkumar: Neural models capture shared mathematical features across disparate physics domains
“So there's implicitly a lot of common features, even across physics that are gone by different equations. So that's how you see across different domains, across different mathematical models. There's a lot of Shared features that these neural models can pick u…”
Anima Anandkumar Sep 4, 2026 ▶ 4:27
Assertion Not checkable as stated
Jenik: Accelerated Understanding achieved 5-trillion token inference context length
“Like we're able to train up to a trillion context input. We're able to train With like, even inputs, outputs, both trillion context length, we are able to do inference at five trillion contexts.”
Benedikt Jenik Sep 4, 2026 ▶ 6:25
Assertion Not checkable as stated
Jenik: Accelerated Understanding has trained physics models up to 1-trillion parameters
“We've trained, done hundreds of training runs. We've trained up to a trillion parameter models. So we've really shown the stuff to take off.”
Benedikt Jenik Sep 4, 2026 ▶ 7:34
Assertion Open · timeframe Sep 2029
Anandkumar: Multi-physics models outperform single-physics models of equivalent parameter size
“And in fact, I was going to add that it turns out that having the model of the same size with multiple areas of physics does better than giving all of those parameters to each single physics. So if you had separate models and made them big enough as the origin…”
Anima Anandkumar Sep 4, 2026 ▶ 8:08
Assertion Contradicted
Anandkumar: Existing video and vision world models incorrectly assume fixed resolutions
“That immediately distinguishes us from other so-called world models, whether it's video models, vision models, they all assume during training and inference, it's a fixed resolution.”
Anima Anandkumar Sep 4, 2026 ▶ 10:15
Insight
Anandkumar: Standard Transformers cannot scale to 5-trillion context lengths for physics
“On the other hand, if you think about using transformer architectures that have worked so well for language, that just wouldn't be able to support a five trillion context length. No matter all the compute in the world is thrown at it. So that kind of quadratic…”
Anima Anandkumar Sep 4, 2026 ▶ 10:58
Assertion Not checkable as stated
Jenik: A single 5-trillion-token model run generated 22 terabytes of output
“Like, for example, our five trillion run that we did, the outputs were 22 terabytes, and you want that kind of stuff in accelerator memory.”
Benedikt Jenik Sep 4, 2026 ▶ 13:07
Assertion Not checkable as stated
Jenik: Accelerated Understanding's medium-sized physics models run inference on Mac Studios
“We have done inference on, like, we can do small models with smaller context, or even medium big models with smaller context fit on a MacBook or Mac Studio.”
Benedikt Jenik Sep 4, 2026 ▶ 13:32
Insight
Jenik: Broadening PDE training domains solves the physical sim-to-real gap
“One interesting thing with PDEs is we are actually fairly confident, like the math is known, that when you solve a PDE correctly, you're doing the physics correctly. Like, obviously, you still need to make sure that you're representing the task that you're try…”
Benedikt Jenik Sep 4, 2026 ▶ 16:05
Insight
Jenik: PDE feedback pushes physics AI models to exceed training data
“You can use those PDEs both for numerical simulators to generate data, but if you're clever about it, you can even use them as a training signal. You can check how well is my model actually doing on the PDEs themselves. And use that as an additional training s…”
Benedikt Jenik Sep 4, 2026 ▶ 17:33
Insight
Anandkumar: Dense physics feedback enables better AI self-improvement than sparse LLMs
“And the difference there is compared to language where self-improvement needs something like human feedback or other reward signals that are very sparse. They just tell you yes or no, thumbs up or down. We have dense feedback because the physics laws, there's …”
Anima Anandkumar Sep 4, 2026 ▶ 18:00
Insight
Anandkumar: Neural operators mix physical scales internally to boost data efficiency
“Neural operators have this flexibility because they can allow you to mix across different scales within the model rather than be prescribed externally like a lot of other Hybrid machine learning for physics do, and that, you know, allows us to be a lot more da…”
Anima Anandkumar Sep 4, 2026 ▶ 19:36
Disclosure
Jenik: Accelerated Understanding targets semiconductors and energy as primary commercial markets
“So if you look at areas that are obviously very interesting right now, because everybody needs improvement there, the big ones to everybody are semi and energy, and those are precisely right in our wheelhouse.”
Benedikt Jenik Sep 4, 2026 ▶ 20:39
Assertion Not checkable as stated
Jenik: Traditional chip design loses performance by separating digital logic and physics
“Especially when you look at the chip design itself, it was much more a, let's start in the digital, let's freeze the digital in, let's send it through some physics for a one time check, like the PDK dictates, I have to have The following feature, otherwise TSN…”
Benedikt Jenik Sep 4, 2026 ▶ 21:05
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.