Nov 7, 2024 · 36m · no-priors

No Priors Ep. 89 | With NVIDIA CEO Jensen Huang

Jensen Huang · 28m spoken Sarah Guo · 2m spoken Elad Gil · 1m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

NVIDIA CEO Jensen Huang discusses the transition to full-stack accelerated computing, the rise of industrial-scale AI factories, and how agentic flywheels and embodied robotics are transforming science and technology.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 12.3% of the talking time here. How this is scored →

The hosts as informed peer 3.1 Guest teaching 3.4 Guest disagreement 1.1 The hosts pushing back 0.2
05100:0010:0020:0030:002:29–7:16 · The hosts as informed peer 5/10 Full-Stack Co-Design, Data Center Networking, and CUDA Acceleration Elad demonstrates domain expertise by introducing proprietary research on the 240x drop in token costs for GPT-4 equivalents over 18 months. Jensen elaborates on the underlying hardware dynamics, explaining Dennard scaling limits, full-stack co-design, and NVLink virtual GPUs.7:16–10:40 · The hosts as informed peer 3/10 Infrastructure Disaggregation and the Hierarchy of AI Models Sarah asks about infrastructure fungibility between training and inference workloads. Jensen educates the hosts on how older training infrastructure naturally cascades down into inference and model distillation.10:40–14:56 · The hosts as informed peer 2/10 The Data Center as the New Unit of Computing When Sarah asks if NVIDIA should build entire data centers, Jensen gently corrects the premise by stating they already build complete operational supercomputer data centers internally before disaggregating components for customers.14:56–18:53 · The hosts as informed peer 2/10 Rapid Engineering and Deployment of xAI's Colossus Supercluster The hosts prompt Jensen on the rapid bringup of xAI's Colossus cluster. Jensen provides an admiring, highly detailed breakdown of the staging, simulation, and hardware deployment logistics.18:53–22:20 · The hosts as informed peer 3/10 Scaling Bottlenecks and the Multi-Vendor Race for General Intelligence Sarah offers multiple choices for the primary scaling bottleneck (capital, energy, supply), but Jensen rejects the framing by declaring 'Everything' is abnormal. Sarah interjects to remind him that nothing is impossible.22:20–27:00 · The hosts as informed peer 4/10 The Paradigm Shift to AI Factories and Token Generation Elad frames NVIDIA's exponential market cap surge with detailed financial comparisons. Jensen reframes the business model, explaining that computing has shifted from storage data centers to token-producing AI factories.27:00–31:17 · The hosts as informed peer 2/10 Embodied Robotics and Enterprise Agent Ecosystems Jensen counters the common industry narrative that enterprise SaaS platforms will be disrupted by AI, arguing instead that incumbents like SAP and ServiceNow sit on gold mines for autonomous domain agents.31:18–35:03 · The hosts as informed peer 2/10 The Groundswell of Generative AI Across All Scientific Disciplines Jensen provides an authoritative monologue on the underappreciated groundswell of generative AI in hard sciences, drawing parallels to his early observations of AlexNet in computer vision.35:05–36:29 · The hosts as informed peer 5/10 Universal Algorithmic Foundations and Personal AI Utilization Sarah articulates a high-level thesis about universal algorithmic foundations bridging computer science across physical scientific disciplines, which Jensen enthusiastically validates before sharing his personal AI workflow.2:29–7:16 · Guest teaching 4/10 Full-Stack Co-Design, Data Center Networking, and CUDA Acceleration Elad demonstrates domain expertise by introducing proprietary research on the 240x drop in token costs for GPT-4 equivalents over 18 months. Jensen elaborates on the underlying hardware dynamics, explaining Dennard scaling limits, full-stack co-design, and NVLink virtual GPUs.7:16–10:40 · Guest teaching 4/10 Infrastructure Disaggregation and the Hierarchy of AI Models Sarah asks about infrastructure fungibility between training and inference workloads. Jensen educates the hosts on how older training infrastructure naturally cascades down into inference and model distillation.10:40–14:56 · Guest teaching 4/10 The Data Center as the New Unit of Computing When Sarah asks if NVIDIA should build entire data centers, Jensen gently corrects the premise by stating they already build complete operational supercomputer data centers internally before disaggregating components for customers.14:56–18:53 · Guest teaching 3/10 Rapid Engineering and Deployment of xAI's Colossus Supercluster The hosts prompt Jensen on the rapid bringup of xAI's Colossus cluster. Jensen provides an admiring, highly detailed breakdown of the staging, simulation, and hardware deployment logistics.18:53–22:20 · Guest teaching 3/10 Scaling Bottlenecks and the Multi-Vendor Race for General Intelligence Sarah offers multiple choices for the primary scaling bottleneck (capital, energy, supply), but Jensen rejects the framing by declaring 'Everything' is abnormal. Sarah interjects to remind him that nothing is impossible.22:20–27:00 · Guest teaching 4/10 The Paradigm Shift to AI Factories and Token Generation Elad frames NVIDIA's exponential market cap surge with detailed financial comparisons. Jensen reframes the business model, explaining that computing has shifted from storage data centers to token-producing AI factories.27:00–31:17 · Guest teaching 4/10 Embodied Robotics and Enterprise Agent Ecosystems Jensen counters the common industry narrative that enterprise SaaS platforms will be disrupted by AI, arguing instead that incumbents like SAP and ServiceNow sit on gold mines for autonomous domain agents.31:18–35:03 · Guest teaching 4/10 The Groundswell of Generative AI Across All Scientific Disciplines Jensen provides an authoritative monologue on the underappreciated groundswell of generative AI in hard sciences, drawing parallels to his early observations of AlexNet in computer vision.35:05–36:29 · Guest teaching 1/10 Universal Algorithmic Foundations and Personal AI Utilization Sarah articulates a high-level thesis about universal algorithmic foundations bridging computer science across physical scientific disciplines, which Jensen enthusiastically validates before sharing his personal AI workflow.2:29–7:16 · Guest disagreement 1/10 Full-Stack Co-Design, Data Center Networking, and CUDA Acceleration Elad demonstrates domain expertise by introducing proprietary research on the 240x drop in token costs for GPT-4 equivalents over 18 months. Jensen elaborates on the underlying hardware dynamics, explaining Dennard scaling limits, full-stack co-design, and NVLink virtual GPUs.7:16–10:40 · Guest disagreement 1/10 Infrastructure Disaggregation and the Hierarchy of AI Models Sarah asks about infrastructure fungibility between training and inference workloads. Jensen educates the hosts on how older training infrastructure naturally cascades down into inference and model distillation.10:40–14:56 · Guest disagreement 2/10 The Data Center as the New Unit of Computing When Sarah asks if NVIDIA should build entire data centers, Jensen gently corrects the premise by stating they already build complete operational supercomputer data centers internally before disaggregating components for customers.14:56–18:53 · Guest disagreement 0/10 Rapid Engineering and Deployment of xAI's Colossus Supercluster The hosts prompt Jensen on the rapid bringup of xAI's Colossus cluster. Jensen provides an admiring, highly detailed breakdown of the staging, simulation, and hardware deployment logistics.18:53–22:20 · Guest disagreement 2/10 Scaling Bottlenecks and the Multi-Vendor Race for General Intelligence Sarah offers multiple choices for the primary scaling bottleneck (capital, energy, supply), but Jensen rejects the framing by declaring 'Everything' is abnormal. Sarah interjects to remind him that nothing is impossible.22:20–27:00 · Guest disagreement 1/10 The Paradigm Shift to AI Factories and Token Generation Elad frames NVIDIA's exponential market cap surge with detailed financial comparisons. Jensen reframes the business model, explaining that computing has shifted from storage data centers to token-producing AI factories.27:00–31:17 · Guest disagreement 2/10 Embodied Robotics and Enterprise Agent Ecosystems Jensen counters the common industry narrative that enterprise SaaS platforms will be disrupted by AI, arguing instead that incumbents like SAP and ServiceNow sit on gold mines for autonomous domain agents.31:18–35:03 · Guest disagreement 1/10 The Groundswell of Generative AI Across All Scientific Disciplines Jensen provides an authoritative monologue on the underappreciated groundswell of generative AI in hard sciences, drawing parallels to his early observations of AlexNet in computer vision.35:05–36:29 · Guest disagreement 0/10 Universal Algorithmic Foundations and Personal AI Utilization Sarah articulates a high-level thesis about universal algorithmic foundations bridging computer science across physical scientific disciplines, which Jensen enthusiastically validates before sharing his personal AI workflow.2:29–7:16 · The hosts pushing back 0/10 Full-Stack Co-Design, Data Center Networking, and CUDA Acceleration Elad demonstrates domain expertise by introducing proprietary research on the 240x drop in token costs for GPT-4 equivalents over 18 months. Jensen elaborates on the underlying hardware dynamics, explaining Dennard scaling limits, full-stack co-design, and NVLink virtual GPUs.7:16–10:40 · The hosts pushing back 0/10 Infrastructure Disaggregation and the Hierarchy of AI Models Sarah asks about infrastructure fungibility between training and inference workloads. Jensen educates the hosts on how older training infrastructure naturally cascades down into inference and model distillation.10:40–14:56 · The hosts pushing back 0/10 The Data Center as the New Unit of Computing When Sarah asks if NVIDIA should build entire data centers, Jensen gently corrects the premise by stating they already build complete operational supercomputer data centers internally before disaggregating components for customers.14:56–18:53 · The hosts pushing back 0/10 Rapid Engineering and Deployment of xAI's Colossus Supercluster The hosts prompt Jensen on the rapid bringup of xAI's Colossus cluster. Jensen provides an admiring, highly detailed breakdown of the staging, simulation, and hardware deployment logistics.18:53–22:20 · The hosts pushing back 2/10 Scaling Bottlenecks and the Multi-Vendor Race for General Intelligence Sarah offers multiple choices for the primary scaling bottleneck (capital, energy, supply), but Jensen rejects the framing by declaring 'Everything' is abnormal. Sarah interjects to remind him that nothing is impossible.22:20–27:00 · The hosts pushing back 0/10 The Paradigm Shift to AI Factories and Token Generation Elad frames NVIDIA's exponential market cap surge with detailed financial comparisons. Jensen reframes the business model, explaining that computing has shifted from storage data centers to token-producing AI factories.27:00–31:17 · The hosts pushing back 0/10 Embodied Robotics and Enterprise Agent Ecosystems Jensen counters the common industry narrative that enterprise SaaS platforms will be disrupted by AI, arguing instead that incumbents like SAP and ServiceNow sit on gold mines for autonomous domain agents.31:18–35:03 · The hosts pushing back 0/10 The Groundswell of Generative AI Across All Scientific Disciplines Jensen provides an authoritative monologue on the underappreciated groundswell of generative AI in hard sciences, drawing parallels to his early observations of AlexNet in computer vision.35:05–36:29 · The hosts pushing back 0/10 Universal Algorithmic Foundations and Personal AI Utilization Sarah articulates a high-level thesis about universal algorithmic foundations bridging computer science across physical scientific disciplines, which Jensen enthusiastically validates before sharing his personal AI workflow.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 28.8% · guest 71.2%0:00 · the hosts 28.8% · guest 71.2%3:00 · the hosts 0.6% · guest 99.4%3:00 · the hosts 0.6% · guest 99.4%6:00 · the hosts 16.5% · guest 83.5%6:00 · the hosts 16.5% · guest 83.5%9:00 · the hosts 9.1% · guest 90.9%9:00 · the hosts 9.1% · guest 90.9%12:00 · the hosts 1.7% · guest 98.3%12:00 · the hosts 1.7% · guest 98.3%15:00 · the hosts 5.5% · guest 94.5%15:00 · the hosts 5.5% · guest 94.5%18:00 · the hosts 22.2% · guest 77.8%18:00 · the hosts 22.2% · guest 77.8%21:00 · the hosts 30.5% · guest 69.5%21:00 · the hosts 30.5% · guest 69.5%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 5.5% · guest 94.5%27:00 · the hosts 5.5% · guest 94.5%30:00 · the hosts 5.9% · guest 94.1%30:00 · the hosts 5.9% · guest 94.1%33:00 · the hosts 14.1% · guest 85.9%33:00 · the hosts 14.1% · guest 85.9%36:00 · the hosts 45.4% · guest 54.6%36:00 · the hosts 45.4% · guest 54.6%
Sharpest disagreement ▶ 19:07 Jensen rejects single-bottleneck framing

When Sarah presents a multiple-choice list of potential scaling bottlenecks, Jensen bluntly dismisses narrowing it down by stating 'Everything' is hard and nothing about these scales is normal.

Hardest push from the hosts ▶ 19:15 Sarah counters Jensen's bottleneck pessimism

After Jensen insists that nothing about megacluster scaling is normal or easy, Sarah immediately pushes back with 'But nothing is impossible.'

Biggest teaching moment ▶ 10:56 Jensen clarifies NVIDIA already builds full data centers

In response to Sarah asking if NVIDIA might eventually build full data centers, Jensen explains that NVIDIA already builds end-to-end supercomputing centers to validate software rather than relying on PowerPoint specs.

The host holds their own ▶ 5:58 Elad presents proprietary token price compression data

Elad demonstrates rigorous market analysis by citing his team's internal findings that GPT-4 equivalent inference token costs dropped 240x in just 18 months.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Full-Stack Co-Design, Data Center Networking, and CUDA Acceleration 5410 Elad demonstrates domain expertise by introducing proprietary research on the 240x drop in token costs for GPT-4 equivalents over 18 months. Jensen elaborates on the underlying hardware dynamics, explaining Dennard scaling limits, full-stack co-design, and NVLink virtual GPUs.
Infrastructure Disaggregation and the Hierarchy of AI Models 3410 Sarah asks about infrastructure fungibility between training and inference workloads. Jensen educates the hosts on how older training infrastructure naturally cascades down into inference and model distillation.
The Data Center as the New Unit of Computing 2420 When Sarah asks if NVIDIA should build entire data centers, Jensen gently corrects the premise by stating they already build complete operational supercomputer data centers internally before disaggregating components for customers.
Rapid Engineering and Deployment of xAI's Colossus Supercluster 2300 The hosts prompt Jensen on the rapid bringup of xAI's Colossus cluster. Jensen provides an admiring, highly detailed breakdown of the staging, simulation, and hardware deployment logistics.
Scaling Bottlenecks and the Multi-Vendor Race for General Intelligence 3322 Sarah offers multiple choices for the primary scaling bottleneck (capital, energy, supply), but Jensen rejects the framing by declaring 'Everything' is abnormal. Sarah interjects to remind him that nothing is impossible.
The Paradigm Shift to AI Factories and Token Generation 4410 Elad frames NVIDIA's exponential market cap surge with detailed financial comparisons. Jensen reframes the business model, explaining that computing has shifted from storage data centers to token-producing AI factories.
Embodied Robotics and Enterprise Agent Ecosystems 2420 Jensen counters the common industry narrative that enterprise SaaS platforms will be disrupted by AI, arguing instead that incumbents like SAP and ServiceNow sit on gold mines for autonomous domain agents.
The Groundswell of Generative AI Across All Scientific Disciplines 2410 Jensen provides an authoritative monologue on the underappreciated groundswell of generative AI in hard sciences, drawing parallels to his early observations of AlexNet in computer vision.
Universal Algorithmic Foundations and Personal AI Utilization 5100 Sarah articulates a high-level thesis about universal algorithmic foundations bridging computer science across physical scientific disciplines, which Jensen enthusiastically validates before sharing his personal AI workflow.

Statements from this episode (21)

Prediction Open · timeframe Nov 2034
Huang: AI compute will follow hyper Moore's Law over the next decade
“Over the next 10 years our hope is that we could double or triple performance every year at scale. Not at chip. At scale. And to be able to, therefore, drive the cost down by a factor of two or three, drive the energy down by a factor of two, three every singl…”
Jensen Huang Nov 7, 2024 ▶ 1:45
Insight
Huang: Scaling AI Compute Requires Controlling Both Algorithms and System Hardware
“The new way of doing scaling are all kinds of things associated with co-design unless you can modify or change the algorithm to reflect the architecture of the system or change and then change the system to reflect the architecture of the new software and go b…”
Jensen Huang Nov 7, 2024 ▶ 3:00
Disclosure
Huang: NVIDIA Bought Mellanox to Treat Data Center Networks as Compute Fabrics
“That, that's the reason why we bought Mellanox and started fusing InfiniBand and MVLink in such an aggressive way.”
Jensen Huang Nov 7, 2024 ▶ 4:04
Assertion Supported
Gil: GPT-4 Equivalent Token Costs Dropped 240x in 18 Months
“Over the last 18 months or so, the cost of a million tokens going into a GPT-IV equivalent model has basically dropped 240 X.”
Elad Gil Nov 7, 2024 ▶ 6:05
Assertion Supported
Huang: NVIDIA Improved Hopper Performance on LLaMA 5x in One Year
“CUDA made it possible for us to iterate so quickly, just in the last year, and then we just went back and benchmarked when Lama first came out, we've improved the performance of hopper by a factor of five without the layer on top ever changing.”
Jensen Huang Nov 7, 2024 ▶ 6:44
Assertion Not checkable as stated
Huang: Sam Altman recently decommissioned OpenAI's Volta GPUs
“Sam was just telling me that he had decommissioned Volta just recently.”
Jensen Huang Nov 7, 2024 ▶ 7:33
Assertion Not checkable as stated
Huang: ChatGPT is largely inferenced on same systems used for training
“And most of ChatGPT I believe are inferenced on the same type of systems that we're trained on just recently.”
Jensen Huang Nov 7, 2024 ▶ 7:59
Prediction Open · timeframe Nov 2029
Huang: Tiny language models will achieve superhuman performance in narrow domains
“And we're going to see superhuman tasks in one little tiny domain from a little tiny, tiny, tiny model.”
Jensen Huang Nov 7, 2024 ▶ 9:33
Insight
Jensen Huang: The new unit of computing is the entire data center
“You know, I say that the new unit of computing is the data center.”
Jensen Huang Nov 7, 2024 ▶ 11:49
Disclosure
Huang: NVIDIA plans to build at least five more supercomputers in 2025
“Next year, we're going to build easily five more.”
Jensen Huang Nov 7, 2024 ▶ 12:14
Assertion Contradicted
Huang: NVIDIA has never abandoned a single piece of software
“We've never given up on a piece of software.”
Jensen Huang Nov 7, 2024 ▶ 13:59
Assertion Supported
Huang: xAI's 100,000 GPU cluster is the largest single unit built
“Decide to build this 100,000 GPU super cluster, which is, you know, the largest of its kind in, in one unit.”
Jensen Huang Nov 7, 2024 ▶ 15:24
Disclosure
Jensen Huang: NVIDIA can set up complete data centers in 30 days
“If you're interested in a data center and just have to give me a space and some power, some cooling, You know, and we'll, we'll help you set it up within, call it, 30 days.”
Jensen Huang Nov 7, 2024 ▶ 18:41
Opinion
Huang: No laws of physics limit scaling superclusters to 1M GPUs
“Nothing is, yeah, no laws of physics limits but everything is going to be hard.”
Jensen Huang Nov 7, 2024 ▶ 19:18
What-if
Huang: NVIDIA could not have built Hopper without AI chip designers
“We can't, we couldn't build Hopper without it.”
Jensen Huang Nov 7, 2024 ▶ 21:16
Assertion Contradicted
Jensen Huang: Computing marginal costs fell 1,000,000x over the last decade
“We've driven down the marginal cost of computing down probably by a million X in the last 10 years, to the point that we just, hey, let's just let the computer go exhaustively write the software.”
Jensen Huang Nov 7, 2024 ▶ 23:46
Insight
Jensen Huang: AI data centers are token-producing factories, not storage
“These new data centers we're creating are not data centers. They don't, they're not multi-tenant. They tend to be single tenant. They're not storing any of our files. They're just, they're producing something and they're producing tokens.”
Jensen Huang Nov 7, 2024 ▶ 24:41
Prediction Not checkable as stated
Huang: The world is close to artificial general robotics
“In a lot of ways, we've, we're close to artificial general intelligence, but we're also close to artificial general robotics.”
Jensen Huang Nov 7, 2024 ▶ 27:06
Opinion
Huang: Incumbent SaaS platforms will flourish with AI agents, not be disrupted
“People, people say that these SaaS platforms are going to be disrupted. I actually think the opposite. That they're sitting on a gold mine, that, that they're going to be this flourishing of agents that are going to be specialized in Salesforce specialized in,…”
Jensen Huang Nov 7, 2024 ▶ 30:35
Prediction Not checkable as stated
Huang: In 2-3 years, every science breakthrough will use generative AI
“If we give ourselves another couple, two, three years, the world's going to change. There's not going to be one paper There's not going to be one breakthrough in science, one breakthrough in engineering where generative AI isn't at the foundation of it. I'm fa…”
Jensen Huang Nov 7, 2024 ▶ 34:00
Disclosure
Jensen Huang: I don't learn anything without first consulting AI
“I don't learn anything without first going to an AI, you know. Why learn the hard way? Just go directly to an AI. I go directly to ChatGPT or, you know, sometimes I do perplexity just depending on just the formulation of my questions, and I just start learning…”
Jensen Huang Nov 7, 2024 ▶ 35:47
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.