Assertion Supported AI assessment confidence: 95% certainty 4/5 debate potential 2/5

Catanzaro: Nemotron 3's Latent MoE quadruples experts for same inference cost

Bryan Catanzaro · Inside Nemotron & NVIDIA’s AI Lab | Bryan Catanzaro · Jul 2, 2026 · at 46:00

Bryan Catanzaro, VP of Applied Deep Learning Research at NVIDIA, details the technical mechanics and efficiency gains of Latent Mixture of Experts (MoE) in the Nemotron 3 open model family.

0:00 / 0:34exact quote · 34.7s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“Latent MOE is a specific innovation that we have in NemoTron three family. And what it does is actually reduces the amount of communication that has to be sent through NVLink during MOE computations by basically down projecting it. So, you know, every token produces a vector and the idea is like, we're going to take that vector and learn a way to compress it. And then send that compressed thing through the network, and then we're gonna uncompress it at the other end. And as a result, we save on network bandwidth, and we also get four times the number of experts for the same inference cost.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Bryan Catanzaro

Opinion
Catanzaro: Chinese AI achievements are not driven by a copycat mentality
“I think it's absolutely false to say that you know the achievements of some other country are all being created by sort of, you know, copycat mentality. It's just not true.”
Bryan Catanzaro Jul 2, 2026 ▶ 8:53 Inside Nemotron & NVIDIA’s AI Lab | Bryan Catanzaro
Opinion
Catanzaro: China has been leading in open community-oriented AI development
“I think there's a chance for the rest of the world to catch up to China in the sense that you know, we can understand the benefits of working together as a community to build technologies for AI in a way that I think China has frankly been leading.”
Bryan Catanzaro Jul 2, 2026 ▶ 10:14 Inside Nemotron & NVIDIA’s AI Lab | Bryan Catanzaro
Opinion
Catanzaro: The technological singularity is a wrongheaded idea
“The singularity is, although it's an attractive idea, I think that it's a really a wrongheaded idea because it doesn't really take into account these other factors.”
Bryan Catanzaro Jul 2, 2026 ▶ 1:14:57 Inside Nemotron & NVIDIA’s AI Lab | Bryan Catanzaro
Opinion
Catanzaro: Open technologies are inherently the safest way to build AI
“I believe that open technologies for AI are inherently the safest way of building AI.”
Bryan Catanzaro Jul 2, 2026 ▶ 1:22:19 Inside Nemotron & NVIDIA’s AI Lab | Bryan Catanzaro
Assertion Not checkable as stated
Catanzaro: Moore's Law has been economically dead for five to ten years
“The original statement of Moore's law was economic, right? It was about, we can afford to put twice as many transistors on the same chip in every, whatever, 24 months, whatever the time period is. And these days that is, Absolutely not the case. It hasn't been…”
Bryan Catanzaro Jul 2, 2026 ▶ 24:36 Inside Nemotron & NVIDIA’s AI Lab | Bryan Catanzaro
Insight
Catanzaro: At AI compute limits, intelligence gains require higher efficiency
“If you accept as the truth that we're going to be running at the limit, then what that means is that the way to get more intelligence is to be more efficient. We can't get more intelligence by applying more force if we're already at the limit. We have to be mo…”
Bryan Catanzaro Jul 2, 2026 ▶ 37:50 Inside Nemotron & NVIDIA’s AI Lab | Bryan Catanzaro
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.