Jul 2, 2026 · 1h 23m · mad

Inside Nemotron & NVIDIA’s AI Lab | Bryan Catanzaro

Bryan Catanzaro · 1h 4m spoken Matt Turck · 12m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of The MAD Podcast, host Matt Turck interviews Bryan Catanzaro, VP of Applied Deep Learning Research at NVIDIA, covering the architectural innovations behind the open-source Nemotron model family. Catanzaro discusses hardware-software co-design, native 4-bit pre-training, open-source AI philosophy, NVIDIA's unique organizational culture, and the conceptualization of AI as humanity's external brain.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Matt holds 16.1% of the talking time here. How this is scored →

Matt as informed peer 4.1 Guest teaching 5.0 Guest disagreement 1.2 Matt pushing back 1.7
05100:0020:0040:001:00:001:20:001:33–3:43 · Matt as informed peer 3/10 Open Source Momentum vs. Closed Source Models Host Matt Turck opens with an overview of open source momentum and asks Bryan to assess the performance gap between open and closed source models. Bryan provides an agreeable overview, drawing a historical parallel between early closed portals like AOL and the open internet.3:43–6:10 · Matt as informed peer 4/10 Drivers of Open Source Progress & Global AI Dynamics Matt pushes Bryan with a provocative question about whether open source relies artificially on distillation from closed models. Bryan gently reframes the question, moving away from competitive scorekeeping to focus on global collective innovation.6:10–10:43 · Matt as informed peer 4/10 Global AI Innovation and Research Contributions from China Matt asks whether rapid Chinese AI progress is primarily copycat distillation or genuine novel research. Bryan explicitly rejects the copycat narrative as false, citing his personal career experience at Baidu alongside Andrew Ng and Dario Amodei.10:43–12:57 · Matt as informed peer 3/10 Enterprise Data Sovereignty & Bryan's Career Foundations Matt asks for the core enterprise case for open source models over proprietary APIs. Bryan explains data sovereignty, trade secret integration, and custom guardrails as key drivers.12:57–17:45 · Matt as informed peer 4/10 Early GPU Computing, ICML 2008, and Working with Dario Amodei Matt invites Bryan to reflect on his career path, noting that pursuing GPU compute in 2008 was a lonely quest. Bryan recounts how academics dismissed his 2008 ICML GPU paper and shares anecdotes about working with a young Dario Amodei.17:45–21:55 · Matt as informed peer 3/10 Rejoining NVIDIA: DLSS, Megatron, and System-Level AI Training Bryan describes returning to NVIDIA in 2016 to work on DLSS and starting Megatron in 2017. He explains how Megatron proved to the industry that GPUs could scale large transformer pre-training.21:55–26:41 · Matt as informed peer 4/10 Why NVIDIA Builds Frontier Models & Accelerated Computing Post-Moore's Law Matt asks if Moore's Law being dead is official statement. Bryan candidly responds that it has been dead for years economically, forcing NVIDIA to focus on accelerated co-design across hardware and software.26:41–30:23 · Matt as informed peer 3/10 Chronology of Nemotron Models: From Megatron-Turing to Nemotron 3 & 4 Bryan walks through the history of Nemotron releases from the 2021 Megatron-Turing 530B model through Llama-Nemotron and Nemotron 3. He humorously acknowledges internal model naming confusions.30:23–33:38 · Matt as informed peer 3/10 The Nemotron Coalition and Ecosystem Pre-Collaboration Matt asks about the Nemotron Coalition announced in early 2025. Bryan explains why NVIDIA actively collaborates with ecosystem partners prior to pre-training rather than dropping models unilaterally.33:38–39:25 · Matt as informed peer 4/10 Nemotron 3 Lineup & Pre-Training in 4-bit (NVFP4) Matt asks Bryan to clarify 4-bit quantization versus 16-bit precision. Bryan uses a visual posterization analogy to explain NVFP4 and details the mathematical hurdles overcome during 4-bit pre-training.39:25–42:42 · Matt as informed peer 5/10 Hybrid Architecture: Mamba State Space Models and Transformers Matt identifies Nemotron's hybrid state-space Mamba and transformer architecture. Bryan details the complementary strengths of constant-memory state-space summaries and lossy-free full attention lookup.42:42–47:26 · Matt as informed peer 6/10 Mixture of Experts (MoE), Blackwell Co-Design, and Latent MoE Matt proposes a corporate department analogy for Mixture of Experts (MoE), which Bryan validates and expands using a library analogy. Bryan explains how Blackwell NVL72 dynamic interconnects support latent MoE compression.47:26–52:48 · Matt as informed peer 6/10 Long Context Windows and Multi-Token Prediction (MTP) Speedups Matt brings up 1M context windows, context compaction, and multi-token prediction (MTP). Bryan delivers a deep explanation of memory bandwidth bottlenecks, showing how higher predictor accuracy directly accelerates inference speed.52:50–55:12 · Matt as informed peer 5/10 Multi-Teacher Distillation and Nemotron-3 Architecture Bryan outlines multi-domain on-policy distillation (MOPD) with 10-15 specialized teacher models. Matt perceptively notes that MOPD solves human organizational friction across 500 researchers as much as technical convergence.55:12–58:01 · Matt as informed peer 5/10 Post-Training Data Sources and Synthetic Data Generation Matt asks where post-training data comes from for non-verifiable enterprise domains. Bryan outlines commercial data acquisition and synthetic data generation, re-emphasizing NVIDIA's commitment to releasing data openly.58:01–1:00:16 · Matt as informed peer 5/10 Reinforcement Learning and Expanding Beyond Verifiable Domains Matt asks if reinforcement learning can expand beyond code and math into less verifiable fields. Bryan notes coding's unique status due to execution feedback loops and predicts increasingly complex RL training environments.1:00:16–1:04:03 · Matt as informed peer 4/10 NVIDIA's Internal AI Research and Organizational Structure Matt explores how NVIDIA structures its internal AI research. Bryan reveals that NVIDIA operates largely outside conventional org charts, drawing volunteer contributions across ten distinct internal divisions.1:04:03–1:06:53 · Matt as informed peer 4/10 GPU Resource Allocation and Research Budgeting Matt inquires how GPU compute gets allocated internally when research demand is infinite. Bryan describes a bi-weekly hierarchical review process that balances researcher conviction against physical limits.1:06:53–1:10:53 · Matt as informed peer 4/10 Balancing Exploratory Research and Bootstrapping Innovation Matt asks how NVIDIA balances applied engineering with high-risk exploratory moonshots. Bryan shares his philosophy of bootstrapping research through iterative small-scale validation.1:10:53–1:12:59 · Matt as informed peer 4/10 NVIDIA's Company Culture and Leadership Tenure Matt observes that NVIDIA maintains an entrepreneurial spirit despite its scale. Bryan attributes this to executive tenure and the cultural maxim 'no one fails alone' in accelerated computing.1:12:59–1:17:51 · Matt as informed peer 4/10 Skeptical Perspectives on the Singularity and AI as an External Brain Matt references Bryan's known skepticism regarding the Singularity. Bryan explains why he considers the concept wrongheaded, contrasting Math Olympiad talent with CEO capabilities and comparing AI to an external brain.1:17:51–1:19:18 · Matt as informed peer 4/10 Addressing Public Perception and AI Backlash Matt asks about public AI backlash and whether the tech industry faces a communication challenge. Bryan argues that public anxiety fades as AI becomes invisibly integrated into everyday tools like turn-by-turn navigation.1:33–3:43 · Guest teaching 3/10 Open Source Momentum vs. Closed Source Models Host Matt Turck opens with an overview of open source momentum and asks Bryan to assess the performance gap between open and closed source models. Bryan provides an agreeable overview, drawing a historical parallel between early closed portals like AOL and the open internet.3:43–6:10 · Guest teaching 4/10 Drivers of Open Source Progress & Global AI Dynamics Matt pushes Bryan with a provocative question about whether open source relies artificially on distillation from closed models. Bryan gently reframes the question, moving away from competitive scorekeeping to focus on global collective innovation.6:10–10:43 · Guest teaching 6/10 Global AI Innovation and Research Contributions from China Matt asks whether rapid Chinese AI progress is primarily copycat distillation or genuine novel research. Bryan explicitly rejects the copycat narrative as false, citing his personal career experience at Baidu alongside Andrew Ng and Dario Amodei.10:43–12:57 · Guest teaching 3/10 Enterprise Data Sovereignty & Bryan's Career Foundations Matt asks for the core enterprise case for open source models over proprietary APIs. Bryan explains data sovereignty, trade secret integration, and custom guardrails as key drivers.12:57–17:45 · Guest teaching 4/10 Early GPU Computing, ICML 2008, and Working with Dario Amodei Matt invites Bryan to reflect on his career path, noting that pursuing GPU compute in 2008 was a lonely quest. Bryan recounts how academics dismissed his 2008 ICML GPU paper and shares anecdotes about working with a young Dario Amodei.17:45–21:55 · Guest teaching 5/10 Rejoining NVIDIA: DLSS, Megatron, and System-Level AI Training Bryan describes returning to NVIDIA in 2016 to work on DLSS and starting Megatron in 2017. He explains how Megatron proved to the industry that GPUs could scale large transformer pre-training.21:55–26:41 · Guest teaching 6/10 Why NVIDIA Builds Frontier Models & Accelerated Computing Post-Moore's Law Matt asks if Moore's Law being dead is official statement. Bryan candidly responds that it has been dead for years economically, forcing NVIDIA to focus on accelerated co-design across hardware and software.26:41–30:23 · Guest teaching 5/10 Chronology of Nemotron Models: From Megatron-Turing to Nemotron 3 & 4 Bryan walks through the history of Nemotron releases from the 2021 Megatron-Turing 530B model through Llama-Nemotron and Nemotron 3. He humorously acknowledges internal model naming confusions.30:23–33:38 · Guest teaching 3/10 The Nemotron Coalition and Ecosystem Pre-Collaboration Matt asks about the Nemotron Coalition announced in early 2025. Bryan explains why NVIDIA actively collaborates with ecosystem partners prior to pre-training rather than dropping models unilaterally.33:38–39:25 · Guest teaching 6/10 Nemotron 3 Lineup & Pre-Training in 4-bit (NVFP4) Matt asks Bryan to clarify 4-bit quantization versus 16-bit precision. Bryan uses a visual posterization analogy to explain NVFP4 and details the mathematical hurdles overcome during 4-bit pre-training.39:25–42:42 · Guest teaching 5/10 Hybrid Architecture: Mamba State Space Models and Transformers Matt identifies Nemotron's hybrid state-space Mamba and transformer architecture. Bryan details the complementary strengths of constant-memory state-space summaries and lossy-free full attention lookup.42:42–47:26 · Guest teaching 6/10 Mixture of Experts (MoE), Blackwell Co-Design, and Latent MoE Matt proposes a corporate department analogy for Mixture of Experts (MoE), which Bryan validates and expands using a library analogy. Bryan explains how Blackwell NVL72 dynamic interconnects support latent MoE compression.47:26–52:48 · Guest teaching 7/10 Long Context Windows and Multi-Token Prediction (MTP) Speedups Matt brings up 1M context windows, context compaction, and multi-token prediction (MTP). Bryan delivers a deep explanation of memory bandwidth bottlenecks, showing how higher predictor accuracy directly accelerates inference speed.52:50–55:12 · Guest teaching 5/10 Multi-Teacher Distillation and Nemotron-3 Architecture Bryan outlines multi-domain on-policy distillation (MOPD) with 10-15 specialized teacher models. Matt perceptively notes that MOPD solves human organizational friction across 500 researchers as much as technical convergence.55:12–58:01 · Guest teaching 5/10 Post-Training Data Sources and Synthetic Data Generation Matt asks where post-training data comes from for non-verifiable enterprise domains. Bryan outlines commercial data acquisition and synthetic data generation, re-emphasizing NVIDIA's commitment to releasing data openly.58:01–1:00:16 · Guest teaching 5/10 Reinforcement Learning and Expanding Beyond Verifiable Domains Matt asks if reinforcement learning can expand beyond code and math into less verifiable fields. Bryan notes coding's unique status due to execution feedback loops and predicts increasingly complex RL training environments.1:00:16–1:04:03 · Guest teaching 6/10 NVIDIA's Internal AI Research and Organizational Structure Matt explores how NVIDIA structures its internal AI research. Bryan reveals that NVIDIA operates largely outside conventional org charts, drawing volunteer contributions across ten distinct internal divisions.1:04:03–1:06:53 · Guest teaching 5/10 GPU Resource Allocation and Research Budgeting Matt inquires how GPU compute gets allocated internally when research demand is infinite. Bryan describes a bi-weekly hierarchical review process that balances researcher conviction against physical limits.1:06:53–1:10:53 · Guest teaching 5/10 Balancing Exploratory Research and Bootstrapping Innovation Matt asks how NVIDIA balances applied engineering with high-risk exploratory moonshots. Bryan shares his philosophy of bootstrapping research through iterative small-scale validation.1:10:53–1:12:59 · Guest teaching 5/10 NVIDIA's Company Culture and Leadership Tenure Matt observes that NVIDIA maintains an entrepreneurial spirit despite its scale. Bryan attributes this to executive tenure and the cultural maxim 'no one fails alone' in accelerated computing.1:12:59–1:17:51 · Guest teaching 6/10 Skeptical Perspectives on the Singularity and AI as an External Brain Matt references Bryan's known skepticism regarding the Singularity. Bryan explains why he considers the concept wrongheaded, contrasting Math Olympiad talent with CEO capabilities and comparing AI to an external brain.1:17:51–1:19:18 · Guest teaching 4/10 Addressing Public Perception and AI Backlash Matt asks about public AI backlash and whether the tech industry faces a communication challenge. Bryan argues that public anxiety fades as AI becomes invisibly integrated into everyday tools like turn-by-turn navigation.1:33–3:43 · Guest disagreement 1/10 Open Source Momentum vs. Closed Source Models Host Matt Turck opens with an overview of open source momentum and asks Bryan to assess the performance gap between open and closed source models. Bryan provides an agreeable overview, drawing a historical parallel between early closed portals like AOL and the open internet.3:43–6:10 · Guest disagreement 2/10 Drivers of Open Source Progress & Global AI Dynamics Matt pushes Bryan with a provocative question about whether open source relies artificially on distillation from closed models. Bryan gently reframes the question, moving away from competitive scorekeeping to focus on global collective innovation.6:10–10:43 · Guest disagreement 3/10 Global AI Innovation and Research Contributions from China Matt asks whether rapid Chinese AI progress is primarily copycat distillation or genuine novel research. Bryan explicitly rejects the copycat narrative as false, citing his personal career experience at Baidu alongside Andrew Ng and Dario Amodei.10:43–12:57 · Guest disagreement 1/10 Enterprise Data Sovereignty & Bryan's Career Foundations Matt asks for the core enterprise case for open source models over proprietary APIs. Bryan explains data sovereignty, trade secret integration, and custom guardrails as key drivers.12:57–17:45 · Guest disagreement 1/10 Early GPU Computing, ICML 2008, and Working with Dario Amodei Matt invites Bryan to reflect on his career path, noting that pursuing GPU compute in 2008 was a lonely quest. Bryan recounts how academics dismissed his 2008 ICML GPU paper and shares anecdotes about working with a young Dario Amodei.17:45–21:55 · Guest disagreement 1/10 Rejoining NVIDIA: DLSS, Megatron, and System-Level AI Training Bryan describes returning to NVIDIA in 2016 to work on DLSS and starting Megatron in 2017. He explains how Megatron proved to the industry that GPUs could scale large transformer pre-training.21:55–26:41 · Guest disagreement 2/10 Why NVIDIA Builds Frontier Models & Accelerated Computing Post-Moore's Law Matt asks if Moore's Law being dead is official statement. Bryan candidly responds that it has been dead for years economically, forcing NVIDIA to focus on accelerated co-design across hardware and software.26:41–30:23 · Guest disagreement 1/10 Chronology of Nemotron Models: From Megatron-Turing to Nemotron 3 & 4 Bryan walks through the history of Nemotron releases from the 2021 Megatron-Turing 530B model through Llama-Nemotron and Nemotron 3. He humorously acknowledges internal model naming confusions.30:23–33:38 · Guest disagreement 1/10 The Nemotron Coalition and Ecosystem Pre-Collaboration Matt asks about the Nemotron Coalition announced in early 2025. Bryan explains why NVIDIA actively collaborates with ecosystem partners prior to pre-training rather than dropping models unilaterally.33:38–39:25 · Guest disagreement 1/10 Nemotron 3 Lineup & Pre-Training in 4-bit (NVFP4) Matt asks Bryan to clarify 4-bit quantization versus 16-bit precision. Bryan uses a visual posterization analogy to explain NVFP4 and details the mathematical hurdles overcome during 4-bit pre-training.39:25–42:42 · Guest disagreement 1/10 Hybrid Architecture: Mamba State Space Models and Transformers Matt identifies Nemotron's hybrid state-space Mamba and transformer architecture. Bryan details the complementary strengths of constant-memory state-space summaries and lossy-free full attention lookup.42:42–47:26 · Guest disagreement 1/10 Mixture of Experts (MoE), Blackwell Co-Design, and Latent MoE Matt proposes a corporate department analogy for Mixture of Experts (MoE), which Bryan validates and expands using a library analogy. Bryan explains how Blackwell NVL72 dynamic interconnects support latent MoE compression.47:26–52:48 · Guest disagreement 1/10 Long Context Windows and Multi-Token Prediction (MTP) Speedups Matt brings up 1M context windows, context compaction, and multi-token prediction (MTP). Bryan delivers a deep explanation of memory bandwidth bottlenecks, showing how higher predictor accuracy directly accelerates inference speed.52:50–55:12 · Guest disagreement 1/10 Multi-Teacher Distillation and Nemotron-3 Architecture Bryan outlines multi-domain on-policy distillation (MOPD) with 10-15 specialized teacher models. Matt perceptively notes that MOPD solves human organizational friction across 500 researchers as much as technical convergence.55:12–58:01 · Guest disagreement 1/10 Post-Training Data Sources and Synthetic Data Generation Matt asks where post-training data comes from for non-verifiable enterprise domains. Bryan outlines commercial data acquisition and synthetic data generation, re-emphasizing NVIDIA's commitment to releasing data openly.58:01–1:00:16 · Guest disagreement 1/10 Reinforcement Learning and Expanding Beyond Verifiable Domains Matt asks if reinforcement learning can expand beyond code and math into less verifiable fields. Bryan notes coding's unique status due to execution feedback loops and predicts increasingly complex RL training environments.1:00:16–1:04:03 · Guest disagreement 1/10 NVIDIA's Internal AI Research and Organizational Structure Matt explores how NVIDIA structures its internal AI research. Bryan reveals that NVIDIA operates largely outside conventional org charts, drawing volunteer contributions across ten distinct internal divisions.1:04:03–1:06:53 · Guest disagreement 1/10 GPU Resource Allocation and Research Budgeting Matt inquires how GPU compute gets allocated internally when research demand is infinite. Bryan describes a bi-weekly hierarchical review process that balances researcher conviction against physical limits.1:06:53–1:10:53 · Guest disagreement 1/10 Balancing Exploratory Research and Bootstrapping Innovation Matt asks how NVIDIA balances applied engineering with high-risk exploratory moonshots. Bryan shares his philosophy of bootstrapping research through iterative small-scale validation.1:10:53–1:12:59 · Guest disagreement 1/10 NVIDIA's Company Culture and Leadership Tenure Matt observes that NVIDIA maintains an entrepreneurial spirit despite its scale. Bryan attributes this to executive tenure and the cultural maxim 'no one fails alone' in accelerated computing.1:12:59–1:17:51 · Guest disagreement 2/10 Skeptical Perspectives on the Singularity and AI as an External Brain Matt references Bryan's known skepticism regarding the Singularity. Bryan explains why he considers the concept wrongheaded, contrasting Math Olympiad talent with CEO capabilities and comparing AI to an external brain.1:17:51–1:19:18 · Guest disagreement 1/10 Addressing Public Perception and AI Backlash Matt asks about public AI backlash and whether the tech industry faces a communication challenge. Bryan argues that public anxiety fades as AI becomes invisibly integrated into everyday tools like turn-by-turn navigation.1:33–3:43 · Matt pushing back 1/10 Open Source Momentum vs. Closed Source Models Host Matt Turck opens with an overview of open source momentum and asks Bryan to assess the performance gap between open and closed source models. Bryan provides an agreeable overview, drawing a historical parallel between early closed portals like AOL and the open internet.3:43–6:10 · Matt pushing back 4/10 Drivers of Open Source Progress & Global AI Dynamics Matt pushes Bryan with a provocative question about whether open source relies artificially on distillation from closed models. Bryan gently reframes the question, moving away from competitive scorekeeping to focus on global collective innovation.6:10–10:43 · Matt pushing back 4/10 Global AI Innovation and Research Contributions from China Matt asks whether rapid Chinese AI progress is primarily copycat distillation or genuine novel research. Bryan explicitly rejects the copycat narrative as false, citing his personal career experience at Baidu alongside Andrew Ng and Dario Amodei.10:43–12:57 · Matt pushing back 1/10 Enterprise Data Sovereignty & Bryan's Career Foundations Matt asks for the core enterprise case for open source models over proprietary APIs. Bryan explains data sovereignty, trade secret integration, and custom guardrails as key drivers.12:57–17:45 · Matt pushing back 1/10 Early GPU Computing, ICML 2008, and Working with Dario Amodei Matt invites Bryan to reflect on his career path, noting that pursuing GPU compute in 2008 was a lonely quest. Bryan recounts how academics dismissed his 2008 ICML GPU paper and shares anecdotes about working with a young Dario Amodei.17:45–21:55 · Matt pushing back 1/10 Rejoining NVIDIA: DLSS, Megatron, and System-Level AI Training Bryan describes returning to NVIDIA in 2016 to work on DLSS and starting Megatron in 2017. He explains how Megatron proved to the industry that GPUs could scale large transformer pre-training.21:55–26:41 · Matt pushing back 3/10 Why NVIDIA Builds Frontier Models & Accelerated Computing Post-Moore's Law Matt asks if Moore's Law being dead is official statement. Bryan candidly responds that it has been dead for years economically, forcing NVIDIA to focus on accelerated co-design across hardware and software.26:41–30:23 · Matt pushing back 1/10 Chronology of Nemotron Models: From Megatron-Turing to Nemotron 3 & 4 Bryan walks through the history of Nemotron releases from the 2021 Megatron-Turing 530B model through Llama-Nemotron and Nemotron 3. He humorously acknowledges internal model naming confusions.30:23–33:38 · Matt pushing back 1/10 The Nemotron Coalition and Ecosystem Pre-Collaboration Matt asks about the Nemotron Coalition announced in early 2025. Bryan explains why NVIDIA actively collaborates with ecosystem partners prior to pre-training rather than dropping models unilaterally.33:38–39:25 · Matt pushing back 2/10 Nemotron 3 Lineup & Pre-Training in 4-bit (NVFP4) Matt asks Bryan to clarify 4-bit quantization versus 16-bit precision. Bryan uses a visual posterization analogy to explain NVFP4 and details the mathematical hurdles overcome during 4-bit pre-training.39:25–42:42 · Matt pushing back 1/10 Hybrid Architecture: Mamba State Space Models and Transformers Matt identifies Nemotron's hybrid state-space Mamba and transformer architecture. Bryan details the complementary strengths of constant-memory state-space summaries and lossy-free full attention lookup.42:42–47:26 · Matt pushing back 1/10 Mixture of Experts (MoE), Blackwell Co-Design, and Latent MoE Matt proposes a corporate department analogy for Mixture of Experts (MoE), which Bryan validates and expands using a library analogy. Bryan explains how Blackwell NVL72 dynamic interconnects support latent MoE compression.47:26–52:48 · Matt pushing back 2/10 Long Context Windows and Multi-Token Prediction (MTP) Speedups Matt brings up 1M context windows, context compaction, and multi-token prediction (MTP). Bryan delivers a deep explanation of memory bandwidth bottlenecks, showing how higher predictor accuracy directly accelerates inference speed.52:50–55:12 · Matt pushing back 1/10 Multi-Teacher Distillation and Nemotron-3 Architecture Bryan outlines multi-domain on-policy distillation (MOPD) with 10-15 specialized teacher models. Matt perceptively notes that MOPD solves human organizational friction across 500 researchers as much as technical convergence.55:12–58:01 · Matt pushing back 2/10 Post-Training Data Sources and Synthetic Data Generation Matt asks where post-training data comes from for non-verifiable enterprise domains. Bryan outlines commercial data acquisition and synthetic data generation, re-emphasizing NVIDIA's commitment to releasing data openly.58:01–1:00:16 · Matt pushing back 1/10 Reinforcement Learning and Expanding Beyond Verifiable Domains Matt asks if reinforcement learning can expand beyond code and math into less verifiable fields. Bryan notes coding's unique status due to execution feedback loops and predicts increasingly complex RL training environments.1:00:16–1:04:03 · Matt pushing back 1/10 NVIDIA's Internal AI Research and Organizational Structure Matt explores how NVIDIA structures its internal AI research. Bryan reveals that NVIDIA operates largely outside conventional org charts, drawing volunteer contributions across ten distinct internal divisions.1:04:03–1:06:53 · Matt pushing back 2/10 GPU Resource Allocation and Research Budgeting Matt inquires how GPU compute gets allocated internally when research demand is infinite. Bryan describes a bi-weekly hierarchical review process that balances researcher conviction against physical limits.1:06:53–1:10:53 · Matt pushing back 2/10 Balancing Exploratory Research and Bootstrapping Innovation Matt asks how NVIDIA balances applied engineering with high-risk exploratory moonshots. Bryan shares his philosophy of bootstrapping research through iterative small-scale validation.1:10:53–1:12:59 · Matt pushing back 1/10 NVIDIA's Company Culture and Leadership Tenure Matt observes that NVIDIA maintains an entrepreneurial spirit despite its scale. Bryan attributes this to executive tenure and the cultural maxim 'no one fails alone' in accelerated computing.1:12:59–1:17:51 · Matt pushing back 2/10 Skeptical Perspectives on the Singularity and AI as an External Brain Matt references Bryan's known skepticism regarding the Singularity. Bryan explains why he considers the concept wrongheaded, contrasting Math Olympiad talent with CEO capabilities and comparing AI to an external brain.1:17:51–1:19:18 · Matt pushing back 2/10 Addressing Public Perception and AI Backlash Matt asks about public AI backlash and whether the tech industry faces a communication challenge. Bryan argues that public anxiety fades as AI becomes invisibly integrated into everyday tools like turn-by-turn navigation.

speaking balance: gold is Matt, purple is the guest (3 minute bins)

0:00 · Matt 54.6% · guest 45.4%0:00 · Matt 54.6% · guest 45.4%3:00 · Matt 36.9% · guest 63.1%3:00 · Matt 36.9% · guest 63.1%6:00 · Matt 26.8% · guest 73.2%6:00 · Matt 26.8% · guest 73.2%9:00 · Matt 6.6% · guest 93.4%9:00 · Matt 6.6% · guest 93.4%12:00 · Matt 13.8% · guest 86.2%12:00 · Matt 13.8% · guest 86.2%15:00 · Matt 6.8% · guest 93.2%15:00 · Matt 6.8% · guest 93.2%18:00 · Matt 0% · guest 100%18:00 · Matt 0% · guest 100%21:00 · Matt 13.7% · guest 86.3%21:00 · Matt 13.7% · guest 86.3%24:00 · Matt 21.7% · guest 78.3%24:00 · Matt 21.7% · guest 78.3%27:00 · Matt 1% · guest 99%27:00 · Matt 1% · guest 99%30:00 · Matt 18.9% · guest 81.1%30:00 · Matt 18.9% · guest 81.1%33:00 · Matt 12% · guest 88%33:00 · Matt 12% · guest 88%36:00 · Matt 2.9% · guest 97.1%36:00 · Matt 2.9% · guest 97.1%39:00 · Matt 12.6% · guest 87.4%39:00 · Matt 12.6% · guest 87.4%42:00 · Matt 15% · guest 85%42:00 · Matt 15% · guest 85%45:00 · Matt 13.6% · guest 86.4%45:00 · Matt 13.6% · guest 86.4%48:00 · Matt 11.5% · guest 88.5%48:00 · Matt 11.5% · guest 88.5%51:00 · Matt 6% · guest 94%51:00 · Matt 6% · guest 94%54:00 · Matt 43.7% · guest 56.3%54:00 · Matt 43.7% · guest 56.3%57:00 · Matt 19.3% · guest 80.7%57:00 · Matt 19.3% · guest 80.7%1:00:00 · Matt 8.9% · guest 91.1%1:00:00 · Matt 8.9% · guest 91.1%1:03:00 · Matt 16.4% · guest 83.6%1:03:00 · Matt 16.4% · guest 83.6%1:06:00 · Matt 11.9% · guest 88.1%1:06:00 · Matt 11.9% · guest 88.1%1:09:00 · Matt 18.3% · guest 81.7%1:09:00 · Matt 18.3% · guest 81.7%1:12:00 · Matt 15.5% · guest 84.5%1:12:00 · Matt 15.5% · guest 84.5%1:15:00 · Matt 5.3% · guest 94.7%1:15:00 · Matt 5.3% · guest 94.7%1:18:00 · Matt 16.4% · guest 83.6%1:18:00 · Matt 16.4% · guest 83.6%1:21:00 · Matt 24.5% · guest 75.5%1:21:00 · Matt 24.5% · guest 75.5%
Sharpest disagreement ▶ 8:21 Rejection of Chinese AI copycat myth

Bryan explicitly rejects the host's framing that Chinese model progress is largely copycat distillation, calling it 'absolutely false' based on his personal experience at Baidu.

Hardest push from Matt ▶ 5:29 Distillation dependence pushback

Matt Turck directly challenges the narrative of open-source momentum by asking whether progress relies heavily on distilling closed models like Anthropic and OpenAI.

Biggest teaching moment ▶ 1:13:30 Singularity reframing via Math Olympiad vs CEO

Bryan reframes the popular notion of AGI and the Singularity by illustrating that raw intelligence is multifaceted and contextual, contrasting Math Olympiad winners with effective company CEOs.

Matt holds his own ▶ 43:59 Host MoE specialist routing analogy

Matt Turck demonstrates sharp domain understanding by constructing an intuitive enterprise analogy for Mixture of Experts (MoE) routing, which Bryan immediately validates.

the scores for every segment, with the reasoning behind each
ChapterTopicMatt as informed peerGuest teachingGuest disagreementMatt pushing backWhy
Open Source Momentum vs. Closed Source Models 3311 Host Matt Turck opens with an overview of open source momentum and asks Bryan to assess the performance gap between open and closed source models. Bryan provides an agreeable overview, drawing a historical parallel between early closed portals like AOL and the open internet.
Drivers of Open Source Progress & Global AI Dynamics 4424 Matt pushes Bryan with a provocative question about whether open source relies artificially on distillation from closed models. Bryan gently reframes the question, moving away from competitive scorekeeping to focus on global collective innovation.
Global AI Innovation and Research Contributions from China 4634 Matt asks whether rapid Chinese AI progress is primarily copycat distillation or genuine novel research. Bryan explicitly rejects the copycat narrative as false, citing his personal career experience at Baidu alongside Andrew Ng and Dario Amodei.
Enterprise Data Sovereignty & Bryan's Career Foundations 3311 Matt asks for the core enterprise case for open source models over proprietary APIs. Bryan explains data sovereignty, trade secret integration, and custom guardrails as key drivers.
Early GPU Computing, ICML 2008, and Working with Dario Amodei 4411 Matt invites Bryan to reflect on his career path, noting that pursuing GPU compute in 2008 was a lonely quest. Bryan recounts how academics dismissed his 2008 ICML GPU paper and shares anecdotes about working with a young Dario Amodei.
Rejoining NVIDIA: DLSS, Megatron, and System-Level AI Training 3511 Bryan describes returning to NVIDIA in 2016 to work on DLSS and starting Megatron in 2017. He explains how Megatron proved to the industry that GPUs could scale large transformer pre-training.
Why NVIDIA Builds Frontier Models & Accelerated Computing Post-Moore's Law 4623 Matt asks if Moore's Law being dead is official statement. Bryan candidly responds that it has been dead for years economically, forcing NVIDIA to focus on accelerated co-design across hardware and software.
Chronology of Nemotron Models: From Megatron-Turing to Nemotron 3 & 4 3511 Bryan walks through the history of Nemotron releases from the 2021 Megatron-Turing 530B model through Llama-Nemotron and Nemotron 3. He humorously acknowledges internal model naming confusions.
The Nemotron Coalition and Ecosystem Pre-Collaboration 3311 Matt asks about the Nemotron Coalition announced in early 2025. Bryan explains why NVIDIA actively collaborates with ecosystem partners prior to pre-training rather than dropping models unilaterally.
Nemotron 3 Lineup & Pre-Training in 4-bit (NVFP4) 4612 Matt asks Bryan to clarify 4-bit quantization versus 16-bit precision. Bryan uses a visual posterization analogy to explain NVFP4 and details the mathematical hurdles overcome during 4-bit pre-training.
Hybrid Architecture: Mamba State Space Models and Transformers 5511 Matt identifies Nemotron's hybrid state-space Mamba and transformer architecture. Bryan details the complementary strengths of constant-memory state-space summaries and lossy-free full attention lookup.
Mixture of Experts (MoE), Blackwell Co-Design, and Latent MoE 6611 Matt proposes a corporate department analogy for Mixture of Experts (MoE), which Bryan validates and expands using a library analogy. Bryan explains how Blackwell NVL72 dynamic interconnects support latent MoE compression.
Long Context Windows and Multi-Token Prediction (MTP) Speedups 6712 Matt brings up 1M context windows, context compaction, and multi-token prediction (MTP). Bryan delivers a deep explanation of memory bandwidth bottlenecks, showing how higher predictor accuracy directly accelerates inference speed.
Multi-Teacher Distillation and Nemotron-3 Architecture 5511 Bryan outlines multi-domain on-policy distillation (MOPD) with 10-15 specialized teacher models. Matt perceptively notes that MOPD solves human organizational friction across 500 researchers as much as technical convergence.
Post-Training Data Sources and Synthetic Data Generation 5512 Matt asks where post-training data comes from for non-verifiable enterprise domains. Bryan outlines commercial data acquisition and synthetic data generation, re-emphasizing NVIDIA's commitment to releasing data openly.
Reinforcement Learning and Expanding Beyond Verifiable Domains 5511 Matt asks if reinforcement learning can expand beyond code and math into less verifiable fields. Bryan notes coding's unique status due to execution feedback loops and predicts increasingly complex RL training environments.
NVIDIA's Internal AI Research and Organizational Structure 4611 Matt explores how NVIDIA structures its internal AI research. Bryan reveals that NVIDIA operates largely outside conventional org charts, drawing volunteer contributions across ten distinct internal divisions.
GPU Resource Allocation and Research Budgeting 4512 Matt inquires how GPU compute gets allocated internally when research demand is infinite. Bryan describes a bi-weekly hierarchical review process that balances researcher conviction against physical limits.
Balancing Exploratory Research and Bootstrapping Innovation 4512 Matt asks how NVIDIA balances applied engineering with high-risk exploratory moonshots. Bryan shares his philosophy of bootstrapping research through iterative small-scale validation.
NVIDIA's Company Culture and Leadership Tenure 4511 Matt observes that NVIDIA maintains an entrepreneurial spirit despite its scale. Bryan attributes this to executive tenure and the cultural maxim 'no one fails alone' in accelerated computing.
Skeptical Perspectives on the Singularity and AI as an External Brain 4622 Matt references Bryan's known skepticism regarding the Singularity. Bryan explains why he considers the concept wrongheaded, contrasting Math Olympiad talent with CEO capabilities and comparing AI to an external brain.
Addressing Public Perception and AI Backlash 4412 Matt asks about public AI backlash and whether the tech industry faces a communication challenge. Bryan argues that public anxiety fades as AI becomes invisibly integrated into everyday tools like turn-by-turn navigation.

Statements from this episode (27)

Opinion
Catanzaro: Chinese AI achievements are not driven by a copycat mentality
“I think it's absolutely false to say that you know the achievements of some other country are all being created by sort of, you know, copycat mentality. It's just not true.”
Bryan Catanzaro Jul 2, 2026 ▶ 8:53
Opinion
Catanzaro: China has been leading in open community-oriented AI development
“I think there's a chance for the rest of the world to catch up to China in the sense that you know, we can understand the benefits of working together as a community to build technologies for AI in a way that I think China has frankly been leading.”
Bryan Catanzaro Jul 2, 2026 ▶ 10:14
Assertion Not checkable as stated
Catanzaro: Enterprise data sovereignty is spurring demand for open AI models
“This is really spurring a lot of demand for open technologies for AI.”
Bryan Catanzaro Jul 2, 2026 ▶ 12:35
Assertion Supported
Catanzaro: cuDNN was NVIDIA's first GPU deep learning product
“Then that led to the creation of QDNN, which was NVIDIA's first product for deep learning on the GPU.”
Bryan Catanzaro Jul 2, 2026 ▶ 14:41
Assertion Supported
Catanzaro: Dario Amodei worked in bioinformatics before deep learning
“At the time he had been working in bioinformatics, so he hadn't been working on deep learning or the things that we call AI these days.”
Bryan Catanzaro Jul 2, 2026 ▶ 15:50
Assertion Not checkable as stated
NVIDIA DLSS Is About 10 Times More Efficient Than Traditional Rendering
“DLSS is our real-time AI for graphics, and it makes a small GPU run like a big GPU. It's about 10 times more efficient because rather than computing the color of every pixel for every frame, we use AI to infer the color.”
Bryan Catanzaro Jul 2, 2026 ▶ 18:40
Assertion Supported
NVIDIA DLSS Generates 23 Out of Every 24 Pixels in Games
“These days, 23 out of every 24 pixels is being generated by our AI model when you're using DLSS to play games”
Bryan Catanzaro Jul 2, 2026 ▶ 19:09
Insight
Catanzaro: Any development or deployment of AI benefits NVIDIA's business
“Whenever AI is further developed and further deployed. It's an opportunity for our business. So, so this is you know, we're very explicitly trying to develop our ecosystem because that's good business for us.”
Bryan Catanzaro Jul 2, 2026 ▶ 23:51
Assertion Not checkable as stated
Catanzaro: Moore's Law has been economically dead for five to ten years
“The original statement of Moore's law was economic, right? It was about, we can afford to put twice as many transistors on the same chip in every, whatever, 24 months, whatever the time period is. And these days that is, Absolutely not the case. It hasn't been…”
Bryan Catanzaro Jul 2, 2026 ▶ 24:36
Assertion Partly supported
Catanzaro: NVIDIA's first major AI model was built jointly with Microsoft
“Our first big model that we trained, we did with Microsoft, right? It was a joint effort where NVIDIA and Microsoft researchers worked side by side to build that.”
Bryan Catanzaro Jul 2, 2026 ▶ 31:51
Assertion Supported
NVIDIA details Nemotron 3 parameter specs from 3B to 55B active
“Nano is a thirty billion total three billion active parameter model. Super is one 20 and 12, and Ultra is five 50 and 55.”
Bryan Catanzaro Jul 2, 2026 ▶ 33:50
Assertion Supported
NVIDIA pre-trained Nemotron Ultra and Super natively using 4-bit floating point
“NemoTron Ultra and Super, ah, were pre-trained using four-bit arithmetic. We pre-trained those in MVFP four”
Bryan Catanzaro Jul 2, 2026 ▶ 35:41
Insight
Catanzaro: At AI compute limits, intelligence gains require higher efficiency
“If you accept as the truth that we're going to be running at the limit, then what that means is that the way to get more intelligence is to be more efficient. We can't get more intelligence by applying more force if we're already at the limit. We have to be mo…”
Bryan Catanzaro Jul 2, 2026 ▶ 37:50
Assertion Supported
Catanzaro: Combining SSMs and transformers produces smarter AI models than either alone
“Using both of these together was actually better than using either one on their own. And that is independent of the speed benefit. That is just the model is smarter.”
Bryan Catanzaro Jul 2, 2026 ▶ 41:10
Assertion Partly supported
Catanzaro: Hybrid state-space transformer architectures are widely adopted in frontier AI
“It's become, I think, Quite widely adopted to use some sort of state space model in conjunction with full attention for the base architecture.”
Bryan Catanzaro Jul 2, 2026 ▶ 41:45
Assertion Supported
Catanzaro: Nemotron 3's Latent MoE quadruples experts for same inference cost
“Latent MOE is a specific innovation that we have in NemoTron three family. And what it does is actually reduces the amount of communication that has to be sent through NVLink during MOE computations by basically down projecting it. So, you know, every token pr…”
Bryan Catanzaro Jul 2, 2026 ▶ 46:00
Assertion Not checkable as stated
Catanzaro: MoEs have long been the default architecture in frontier AI
“Yeah, I believe MOEs have been the default in Frontier AI for a long time. They're just a really good combination of inference cost and intelligence.”
Bryan Catanzaro Jul 2, 2026 ▶ 46:54
Insight
Catanzaro: Dense models outperform MoE models under strict memory constraints
“You know, they take a lot more memory. If you have a very small amount of memory, a dense model is going to be smarter.”
Bryan Catanzaro Jul 2, 2026 ▶ 47:06
Insight
Catanzaro: Multi-token prediction lowers inference costs as model accuracy improves
“With multi-token prediction, the speed that you get is a function of the accuracy of your model. The more accurate your model is, the faster the inference is, the cheaper the inference is, the more accurate it is. That's not usually how it works, but in this c…”
Bryan Catanzaro Jul 2, 2026 ▶ 52:02
Assertion Supported
Catanzaro: NVIDIA post-trained Nemotron-3 Ultra using multi-domain on-policy distillation
“So with Nemo Tron three ultra, we did post-training using something called multi-domain on policy distillation.”
Bryan Catanzaro Jul 2, 2026 ▶ 52:59
Assertion Supported
Catanzaro: NVIDIA used 10 to 15 teacher models for Nemotron-3 distillation
“There's with NemoTron three, I think we had about 10, 10 or 15 of these teachers.”
Bryan Catanzaro Jul 2, 2026 ▶ 53:24
Disclosure
NVIDIA purchases commercial datasets and opens them when licensing permits
“We do purchase data from companies that that, you know, are building data sets that you can purchase. And to the extent that, you know, we have the rights to redistribute or to open up that data, we do as part of our Mnemotron data effort.”
Bryan Catanzaro Jul 2, 2026 ▶ 56:33
Disclosure
NVIDIA uses massive compute to generate and publicly release synthetic data
“We also are big believers in synthetic data generation. We use an enormous amount of compute running language models on our own systems to create synthetic data that then helps our models be better at solving problems in specific domains, and we release a lot …”
Bryan Catanzaro Jul 2, 2026 ▶ 57:20
Disclosure
Catanzaro: NVIDIA's Applied Deep Learning team sits inside GPU division
“My team, for example, is not part of the official NVIDIA research team. My team is actually part of the organization that builds the GPU.”
Bryan Catanzaro Jul 2, 2026 ▶ 1:00:43
Insight
Catanzaro: In accelerated computing, software failure destroys hardware value
“Accelerated computing is the composition of thousands of technologies. If any of them fail to deliver acceleration, the value is destroyed. It doesn't matter whether the chip is great if the compiler sucks.”
Bryan Catanzaro Jul 2, 2026 ▶ 1:12:14
Opinion
Catanzaro: The technological singularity is a wrongheaded idea
“The singularity is, although it's an attractive idea, I think that it's a really a wrongheaded idea because it doesn't really take into account these other factors.”
Bryan Catanzaro Jul 2, 2026 ▶ 1:14:57
Opinion
Catanzaro: Open technologies are inherently the safest way to build AI
“I believe that open technologies for AI are inherently the safest way of building AI.”
Bryan Catanzaro Jul 2, 2026 ▶ 1:22:19
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 400 conversations transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.