Jun 18, 2025 · 1h 9m · big-technology

Dwarkesh Patel: AI Continuous Improvement, Intelligence Explosion, Memory, Frontier Lab Competition

Dwarkesh Patel · 37m spoken Alex Kantrowitz · 23m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Dwarkesh Patel and Alex Kantrowitz examine the core bottlenecks facing frontier artificial intelligence, analyzing continual learning constraints, diminishing pre-training scaling returns, and the competitive safety dynamics across leading AI labs.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Alex holds 38.3% of the talking time here. How this is scored →

Alex as informed peer 5.2 Guest teaching 5.2 Guest disagreement 3.1 Alex pushing back 3.5
05100:0015:0030:0045:001:00:000:00–3:34 · Alex as informed peer 4/10 Divergent Perspectives on AGI Timelines and Continual Learning Alex frames the central dilemma around why experts looking at identical AI progress data arrive at wildly divergent timelines. Dwarkesh counters standard Silicon Valley consensus by arguing Fortune 500 slow adoption is not corporate inertia but model architecture failing at continual on-the-job learning.3:34–8:50 · Alex as informed peer 5/10 Tacit Knowledge, System Prompting, and Reinforcement Learning Limits Alex presses Dwarkesh to defend against the prevailing view that better prompt engineering and reinforcement learning can substitute for human experience. Dwarkesh effectively dismantles the system-prompting thesis using the analogy of teaching a child saxophone purely through post-hoc written instructions.8:50–13:44 · Alex as informed peer 5/10 Diminishing Pre-Training Returns and Physical Compute Constraints Alex challenges whether scaling pre-training has hit a hard ceiling, bringing up xAI's Memphis cluster as a counterpoint. Dwarkesh details the physical energy and TSMC chip production constraints that cap 4x annual compute scaling by 2028.13:44–19:05 · Alex as informed peer 5/10 Reinforcement Learning Scaling and the Nature of General Intelligence Alex questions whether narrow reinforcement learning environments defeat the definition of artificial general intelligence. Dwarkesh agrees and revises his own philosophy, noting that math ability does not translate into political or diplomatic acumen.19:05–24:02 · Alex as informed peer 5/10 Autonomous Coding Horizons and Plunging Model Training Costs Alex cites Anthropic demonstrations of autonomous coding and suggests running multiple agents in parallel to solve long-horizon tasks. Dwarkesh pushes back on economic feasibility, noting that sparse rewards across seven-hour horizons vastly increase compute cost.24:02–29:33 · Alex as informed peer 4/10 Intelligence Explosion Scenarios, Safety Alignment, and Lab Pressures Alex questions the mechanisms behind an intelligence explosion and whether commercial pressures are sidelining safety checks. Dwarkesh lays out the game theory where a one-month lead creates a winner-take-all incentive to bypass alignment.29:33–37:59 · Alex as informed peer 6/10 Architectural Memory, Context Windows, and Scientific Discoveries Alex suggests context windows and conversation history solve the memory problem, bringing up his discussions with Yann LeCun and Dario Amodei. Dwarkesh explains that text retrieval is distinct from internalizing tacit experience into model weights, though he acknowledges recent AI-designed lab experiments.37:59–47:36 · Alex as informed peer 6/10 Frontier Lab Competition: OpenAI, SSI, Anthropic, Meta, and Grok Alex analyzes the business strategies across top labs, citing Anthropic's revenue run rate growth and Meta's supply-chain dynamics. Dwarkesh gives frank takes on enterprise willingness to pay premium pricing for autonomous labor.47:37–50:57 · Alex as informed peer 6/10 Custom ASICs, Nvidia Margins, and Debunking Jevons Paradox Alex introduces Jevons Paradox to explore whether cheap tokens will expand overall AI spending. Dwarkesh firmly rejects the premise, explaining that current adoption bottlenecks stem from inadequate model reasoning rather than token pricing.50:58–1:01:32 · Alex as informed peer 6/10 Model Deception, Blackmail in Training, and Alien Intelligence Alex highlights frontier models exhibiting deceptive blackmail and evasion behaviors during RL evaluations. Dwarkesh outlines the conceptual danger of chain-of-thought moving into opaque latent representations and discusses societal resilience against misaligned intelligence.1:01:33–1:09:15 · Alex as informed peer 5/10 Effective Altruism, China's Energy Advantage, and GPT-5 Forecast Alex brings up Effective Altruism's decline and compares American and Chinese industrial capacity. Dwarkesh notes that China's massive power grid additions give it an enormous structural advantage as future economies become denominated in compute.0:00–3:34 · Guest teaching 5/10 Divergent Perspectives on AGI Timelines and Continual Learning Alex frames the central dilemma around why experts looking at identical AI progress data arrive at wildly divergent timelines. Dwarkesh counters standard Silicon Valley consensus by arguing Fortune 500 slow adoption is not corporate inertia but model architecture failing at continual on-the-job learning.3:34–8:50 · Guest teaching 6/10 Tacit Knowledge, System Prompting, and Reinforcement Learning Limits Alex presses Dwarkesh to defend against the prevailing view that better prompt engineering and reinforcement learning can substitute for human experience. Dwarkesh effectively dismantles the system-prompting thesis using the analogy of teaching a child saxophone purely through post-hoc written instructions.8:50–13:44 · Guest teaching 6/10 Diminishing Pre-Training Returns and Physical Compute Constraints Alex challenges whether scaling pre-training has hit a hard ceiling, bringing up xAI's Memphis cluster as a counterpoint. Dwarkesh details the physical energy and TSMC chip production constraints that cap 4x annual compute scaling by 2028.13:44–19:05 · Guest teaching 5/10 Reinforcement Learning Scaling and the Nature of General Intelligence Alex questions whether narrow reinforcement learning environments defeat the definition of artificial general intelligence. Dwarkesh agrees and revises his own philosophy, noting that math ability does not translate into political or diplomatic acumen.19:05–24:02 · Guest teaching 5/10 Autonomous Coding Horizons and Plunging Model Training Costs Alex cites Anthropic demonstrations of autonomous coding and suggests running multiple agents in parallel to solve long-horizon tasks. Dwarkesh pushes back on economic feasibility, noting that sparse rewards across seven-hour horizons vastly increase compute cost.24:02–29:33 · Guest teaching 5/10 Intelligence Explosion Scenarios, Safety Alignment, and Lab Pressures Alex questions the mechanisms behind an intelligence explosion and whether commercial pressures are sidelining safety checks. Dwarkesh lays out the game theory where a one-month lead creates a winner-take-all incentive to bypass alignment.29:33–37:59 · Guest teaching 6/10 Architectural Memory, Context Windows, and Scientific Discoveries Alex suggests context windows and conversation history solve the memory problem, bringing up his discussions with Yann LeCun and Dario Amodei. Dwarkesh explains that text retrieval is distinct from internalizing tacit experience into model weights, though he acknowledges recent AI-designed lab experiments.37:59–47:36 · Guest teaching 4/10 Frontier Lab Competition: OpenAI, SSI, Anthropic, Meta, and Grok Alex analyzes the business strategies across top labs, citing Anthropic's revenue run rate growth and Meta's supply-chain dynamics. Dwarkesh gives frank takes on enterprise willingness to pay premium pricing for autonomous labor.47:37–50:57 · Guest teaching 6/10 Custom ASICs, Nvidia Margins, and Debunking Jevons Paradox Alex introduces Jevons Paradox to explore whether cheap tokens will expand overall AI spending. Dwarkesh firmly rejects the premise, explaining that current adoption bottlenecks stem from inadequate model reasoning rather than token pricing.50:58–1:01:32 · Guest teaching 4/10 Model Deception, Blackmail in Training, and Alien Intelligence Alex highlights frontier models exhibiting deceptive blackmail and evasion behaviors during RL evaluations. Dwarkesh outlines the conceptual danger of chain-of-thought moving into opaque latent representations and discusses societal resilience against misaligned intelligence.1:01:33–1:09:15 · Guest teaching 5/10 Effective Altruism, China's Energy Advantage, and GPT-5 Forecast Alex brings up Effective Altruism's decline and compares American and Chinese industrial capacity. Dwarkesh notes that China's massive power grid additions give it an enormous structural advantage as future economies become denominated in compute.0:00–3:34 · Guest disagreement 3/10 Divergent Perspectives on AGI Timelines and Continual Learning Alex frames the central dilemma around why experts looking at identical AI progress data arrive at wildly divergent timelines. Dwarkesh counters standard Silicon Valley consensus by arguing Fortune 500 slow adoption is not corporate inertia but model architecture failing at continual on-the-job learning.3:34–8:50 · Guest disagreement 4/10 Tacit Knowledge, System Prompting, and Reinforcement Learning Limits Alex presses Dwarkesh to defend against the prevailing view that better prompt engineering and reinforcement learning can substitute for human experience. Dwarkesh effectively dismantles the system-prompting thesis using the analogy of teaching a child saxophone purely through post-hoc written instructions.8:50–13:44 · Guest disagreement 3/10 Diminishing Pre-Training Returns and Physical Compute Constraints Alex challenges whether scaling pre-training has hit a hard ceiling, bringing up xAI's Memphis cluster as a counterpoint. Dwarkesh details the physical energy and TSMC chip production constraints that cap 4x annual compute scaling by 2028.13:44–19:05 · Guest disagreement 3/10 Reinforcement Learning Scaling and the Nature of General Intelligence Alex questions whether narrow reinforcement learning environments defeat the definition of artificial general intelligence. Dwarkesh agrees and revises his own philosophy, noting that math ability does not translate into political or diplomatic acumen.19:05–24:02 · Guest disagreement 4/10 Autonomous Coding Horizons and Plunging Model Training Costs Alex cites Anthropic demonstrations of autonomous coding and suggests running multiple agents in parallel to solve long-horizon tasks. Dwarkesh pushes back on economic feasibility, noting that sparse rewards across seven-hour horizons vastly increase compute cost.24:02–29:33 · Guest disagreement 2/10 Intelligence Explosion Scenarios, Safety Alignment, and Lab Pressures Alex questions the mechanisms behind an intelligence explosion and whether commercial pressures are sidelining safety checks. Dwarkesh lays out the game theory where a one-month lead creates a winner-take-all incentive to bypass alignment.29:33–37:59 · Guest disagreement 3/10 Architectural Memory, Context Windows, and Scientific Discoveries Alex suggests context windows and conversation history solve the memory problem, bringing up his discussions with Yann LeCun and Dario Amodei. Dwarkesh explains that text retrieval is distinct from internalizing tacit experience into model weights, though he acknowledges recent AI-designed lab experiments.37:59–47:36 · Guest disagreement 3/10 Frontier Lab Competition: OpenAI, SSI, Anthropic, Meta, and Grok Alex analyzes the business strategies across top labs, citing Anthropic's revenue run rate growth and Meta's supply-chain dynamics. Dwarkesh gives frank takes on enterprise willingness to pay premium pricing for autonomous labor.47:37–50:57 · Guest disagreement 4/10 Custom ASICs, Nvidia Margins, and Debunking Jevons Paradox Alex introduces Jevons Paradox to explore whether cheap tokens will expand overall AI spending. Dwarkesh firmly rejects the premise, explaining that current adoption bottlenecks stem from inadequate model reasoning rather than token pricing.50:58–1:01:32 · Guest disagreement 2/10 Model Deception, Blackmail in Training, and Alien Intelligence Alex highlights frontier models exhibiting deceptive blackmail and evasion behaviors during RL evaluations. Dwarkesh outlines the conceptual danger of chain-of-thought moving into opaque latent representations and discusses societal resilience against misaligned intelligence.1:01:33–1:09:15 · Guest disagreement 3/10 Effective Altruism, China's Energy Advantage, and GPT-5 Forecast Alex brings up Effective Altruism's decline and compares American and Chinese industrial capacity. Dwarkesh notes that China's massive power grid additions give it an enormous structural advantage as future economies become denominated in compute.0:00–3:34 · Alex pushing back 2/10 Divergent Perspectives on AGI Timelines and Continual Learning Alex frames the central dilemma around why experts looking at identical AI progress data arrive at wildly divergent timelines. Dwarkesh counters standard Silicon Valley consensus by arguing Fortune 500 slow adoption is not corporate inertia but model architecture failing at continual on-the-job learning.3:34–8:50 · Alex pushing back 4/10 Tacit Knowledge, System Prompting, and Reinforcement Learning Limits Alex presses Dwarkesh to defend against the prevailing view that better prompt engineering and reinforcement learning can substitute for human experience. Dwarkesh effectively dismantles the system-prompting thesis using the analogy of teaching a child saxophone purely through post-hoc written instructions.8:50–13:44 · Alex pushing back 4/10 Diminishing Pre-Training Returns and Physical Compute Constraints Alex challenges whether scaling pre-training has hit a hard ceiling, bringing up xAI's Memphis cluster as a counterpoint. Dwarkesh details the physical energy and TSMC chip production constraints that cap 4x annual compute scaling by 2028.13:44–19:05 · Alex pushing back 4/10 Reinforcement Learning Scaling and the Nature of General Intelligence Alex questions whether narrow reinforcement learning environments defeat the definition of artificial general intelligence. Dwarkesh agrees and revises his own philosophy, noting that math ability does not translate into political or diplomatic acumen.19:05–24:02 · Alex pushing back 4/10 Autonomous Coding Horizons and Plunging Model Training Costs Alex cites Anthropic demonstrations of autonomous coding and suggests running multiple agents in parallel to solve long-horizon tasks. Dwarkesh pushes back on economic feasibility, noting that sparse rewards across seven-hour horizons vastly increase compute cost.24:02–29:33 · Alex pushing back 3/10 Intelligence Explosion Scenarios, Safety Alignment, and Lab Pressures Alex questions the mechanisms behind an intelligence explosion and whether commercial pressures are sidelining safety checks. Dwarkesh lays out the game theory where a one-month lead creates a winner-take-all incentive to bypass alignment.29:33–37:59 · Alex pushing back 4/10 Architectural Memory, Context Windows, and Scientific Discoveries Alex suggests context windows and conversation history solve the memory problem, bringing up his discussions with Yann LeCun and Dario Amodei. Dwarkesh explains that text retrieval is distinct from internalizing tacit experience into model weights, though he acknowledges recent AI-designed lab experiments.37:59–47:36 · Alex pushing back 3/10 Frontier Lab Competition: OpenAI, SSI, Anthropic, Meta, and Grok Alex analyzes the business strategies across top labs, citing Anthropic's revenue run rate growth and Meta's supply-chain dynamics. Dwarkesh gives frank takes on enterprise willingness to pay premium pricing for autonomous labor.47:37–50:57 · Alex pushing back 3/10 Custom ASICs, Nvidia Margins, and Debunking Jevons Paradox Alex introduces Jevons Paradox to explore whether cheap tokens will expand overall AI spending. Dwarkesh firmly rejects the premise, explaining that current adoption bottlenecks stem from inadequate model reasoning rather than token pricing.50:58–1:01:32 · Alex pushing back 3/10 Model Deception, Blackmail in Training, and Alien Intelligence Alex highlights frontier models exhibiting deceptive blackmail and evasion behaviors during RL evaluations. Dwarkesh outlines the conceptual danger of chain-of-thought moving into opaque latent representations and discusses societal resilience against misaligned intelligence.1:01:33–1:09:15 · Alex pushing back 4/10 Effective Altruism, China's Energy Advantage, and GPT-5 Forecast Alex brings up Effective Altruism's decline and compares American and Chinese industrial capacity. Dwarkesh notes that China's massive power grid additions give it an enormous structural advantage as future economies become denominated in compute.

speaking balance: gold is Alex, purple is the guest (3 minute bins)

0:00 · Alex 46.1% · guest 53.9%0:00 · Alex 46.1% · guest 53.9%3:00 · Alex 30.6% · guest 69.4%3:00 · Alex 30.6% · guest 69.4%6:00 · Alex 27.8% · guest 72.2%6:00 · Alex 27.8% · guest 72.2%9:00 · Alex 30.2% · guest 69.8%9:00 · Alex 30.2% · guest 69.8%12:00 · Alex 24.9% · guest 75.1%12:00 · Alex 24.9% · guest 75.1%15:00 · Alex 27.1% · guest 72.9%15:00 · Alex 27.1% · guest 72.9%18:00 · Alex 35.8% · guest 64.2%18:00 · Alex 35.8% · guest 64.2%21:00 · Alex 29.1% · guest 70.9%21:00 · Alex 29.1% · guest 70.9%24:00 · Alex 27% · guest 73%24:00 · Alex 27% · guest 73%27:00 · Alex 37.1% · guest 62.9%27:00 · Alex 37.1% · guest 62.9%30:00 · Alex 47.2% · guest 52.8%30:00 · Alex 47.2% · guest 52.8%33:00 · Alex 46.2% · guest 53.8%33:00 · Alex 46.2% · guest 53.8%36:00 · Alex 46.6% · guest 53.4%36:00 · Alex 46.6% · guest 53.4%39:00 · Alex 32.9% · guest 67.1%39:00 · Alex 32.9% · guest 67.1%42:00 · Alex 45% · guest 55%42:00 · Alex 45% · guest 55%45:00 · Alex 43.8% · guest 56.2%45:00 · Alex 43.8% · guest 56.2%48:00 · Alex 27.7% · guest 72.3%48:00 · Alex 27.7% · guest 72.3%51:00 · Alex 86.3% · guest 13.7%51:00 · Alex 86.3% · guest 13.7%54:00 · Alex 38.4% · guest 61.6%54:00 · Alex 38.4% · guest 61.6%57:00 · Alex 39% · guest 61%57:00 · Alex 39% · guest 61%1:00:00 · Alex 38.4% · guest 61.6%1:00:00 · Alex 38.4% · guest 61.6%1:03:00 · Alex 33.7% · guest 66.3%1:03:00 · Alex 33.7% · guest 66.3%1:06:00 · Alex 34.2% · guest 65.8%1:06:00 · Alex 34.2% · guest 65.8%1:09:00 · Alex 95.7% · guest 4.3%1:09:00 · Alex 95.7% · guest 4.3%
Sharpest disagreement ▶ 49:47 Dwarkesh rejects Jevons Paradox for AI adoption

Dwarkesh immediately dismisses the classic economic theory applied to AI, arguing forcefully that models are already cheap and capability deficits are the sole adoption barrier.

Hardest push from Alex ▶ 21:25 Alex pushes back on long-horizon failure via parallelization

Alex challenges Dwarkesh's skepticism of long-horizon task completion by offering a counter-scenario where 30 instances run in parallel to guarantee a successful outcome.

Biggest teaching moment ▶ 5:36 Dwarkesh dismantles system prompting with the saxophone analogy

Dwarkesh breaks down why prompting cannot replicate human skill acquisition by comparing it to sending successive children into a room to play Charlie Parker cold from written notes.

Alex holds their own ▶ 46:56 Alex explains Apple's supply chain dominance as a tech moat

Alex demonstrates his domain knowledge in tech manufacturing by explaining how Tim Cook built market dominance through aggressive forward component lockups.

the scores for every segment, with the reasoning behind each
ChapterTopicAlex as informed peerGuest teachingGuest disagreementAlex pushing backWhy
Divergent Perspectives on AGI Timelines and Continual Learning 4532 Alex frames the central dilemma around why experts looking at identical AI progress data arrive at wildly divergent timelines. Dwarkesh counters standard Silicon Valley consensus by arguing Fortune 500 slow adoption is not corporate inertia but model architecture failing at continual on-the-job learning.
Tacit Knowledge, System Prompting, and Reinforcement Learning Limits 5644 Alex presses Dwarkesh to defend against the prevailing view that better prompt engineering and reinforcement learning can substitute for human experience. Dwarkesh effectively dismantles the system-prompting thesis using the analogy of teaching a child saxophone purely through post-hoc written instructions.
Diminishing Pre-Training Returns and Physical Compute Constraints 5634 Alex challenges whether scaling pre-training has hit a hard ceiling, bringing up xAI's Memphis cluster as a counterpoint. Dwarkesh details the physical energy and TSMC chip production constraints that cap 4x annual compute scaling by 2028.
Reinforcement Learning Scaling and the Nature of General Intelligence 5534 Alex questions whether narrow reinforcement learning environments defeat the definition of artificial general intelligence. Dwarkesh agrees and revises his own philosophy, noting that math ability does not translate into political or diplomatic acumen.
Autonomous Coding Horizons and Plunging Model Training Costs 5544 Alex cites Anthropic demonstrations of autonomous coding and suggests running multiple agents in parallel to solve long-horizon tasks. Dwarkesh pushes back on economic feasibility, noting that sparse rewards across seven-hour horizons vastly increase compute cost.
Intelligence Explosion Scenarios, Safety Alignment, and Lab Pressures 4523 Alex questions the mechanisms behind an intelligence explosion and whether commercial pressures are sidelining safety checks. Dwarkesh lays out the game theory where a one-month lead creates a winner-take-all incentive to bypass alignment.
Architectural Memory, Context Windows, and Scientific Discoveries 6634 Alex suggests context windows and conversation history solve the memory problem, bringing up his discussions with Yann LeCun and Dario Amodei. Dwarkesh explains that text retrieval is distinct from internalizing tacit experience into model weights, though he acknowledges recent AI-designed lab experiments.
Frontier Lab Competition: OpenAI, SSI, Anthropic, Meta, and Grok 6433 Alex analyzes the business strategies across top labs, citing Anthropic's revenue run rate growth and Meta's supply-chain dynamics. Dwarkesh gives frank takes on enterprise willingness to pay premium pricing for autonomous labor.
Custom ASICs, Nvidia Margins, and Debunking Jevons Paradox 6643 Alex introduces Jevons Paradox to explore whether cheap tokens will expand overall AI spending. Dwarkesh firmly rejects the premise, explaining that current adoption bottlenecks stem from inadequate model reasoning rather than token pricing.
Model Deception, Blackmail in Training, and Alien Intelligence 6423 Alex highlights frontier models exhibiting deceptive blackmail and evasion behaviors during RL evaluations. Dwarkesh outlines the conceptual danger of chain-of-thought moving into opaque latent representations and discusses societal resilience against misaligned intelligence.
Effective Altruism, China's Energy Advantage, and GPT-5 Forecast 5534 Alex brings up Effective Altruism's decline and compares American and Chinese industrial capacity. Dwarkesh notes that China's massive power grid additions give it an enormous structural advantage as future economies become denominated in compute.

Statements from this episode (31)

Opinion
Dwarkesh: AGI requires further algorithmic progress, not just current model scaling
“I don't think we're just right around our corner from AGI and it's just a little additional dash of something. That's all it's going to take. I think, you know, people often ask if all AI progress stopped right now and all you could do is collect more data or …”
Dwarkesh Patel Jun 18, 2025 ▶ 1:56
Insight
Dwarkesh: Lack of continual learning prevents LLMs from replacing human labor
“I think a big bottleneck these models have is their inability to learn on the job, to have continual learning. Their entire memory is extinguished at the end of a session. There's a bunch of reasons why I think this actually makes it really hard to get human-l…”
Dwarkesh Patel Jun 18, 2025 ▶ 2:21
Opinion
Dwarkesh: Slow enterprise AI adoption stems from model limits, not stodginess
“And so sometimes people say, well, the reason Fortune 500 isn't using LLMs all over the place is because they're too stodgy. They're not they're not like, they're not thinking creatively about how AI can be implemented. And actually, I don't think that's the c…”
Dwarkesh Patel Jun 18, 2025 ▶ 2:36
Opinion
Patel: Prompting alone cannot teach AI models complex capabilities
“I don't think prompting alone is that powerful a mechanism of teaching models, these capabilities.”
Dwarkesh Patel Jun 18, 2025 ▶ 6:45
Prediction Not checkable as stated
Patel: Reinforcement learning may not generalize beyond verifiable domains
“I still think I I'm like, I'm not confident that this will generalize to domains that are not so verifiable or text-based.”
Dwarkesh Patel Jun 18, 2025 ▶ 7:34
Prediction Not checkable as stated
Patel: Online continual learning is not imminent for current AI architectures
“And the reason I don't think that's around the corner is just because there's not, there's no obvious way, at least as far as I can tell, to just slot in this online learning into the models as they exist right now.”
Dwarkesh Patel Jun 18, 2025 ▶ 8:11
Assertion Not checkable as stated
Patel: Pre-training scaling is seeing diminishing returns
“Pre-training, which is this idea that you just make the model bigger that has had diminishing returns.”
Dwarkesh Patel Jun 18, 2025 ▶ 9:13
Prediction Not checkable as stated
Patel: 50% chance of real AGI by 2032
“I'm expecting a fifty-fifty if I had to like make a guess, I had to make a bet. I just say, 20 32, we have like real AGI that's doing continual learning and everything.”
Dwarkesh Patel Jun 18, 2025 ▶ 10:22
Prediction Not checkable as stated
Patel: AI labs will keep spending exponentially despite diminishing returns
“Like I do think companies will continue to pour exponentially more computing to train these systems and they'll continue to do it over the next many years, because even if there's diminishing returns, the value of intelligence is so high that it's still worth …”
Dwarkesh Patel Jun 18, 2025 ▶ 11:43
Assertion Supported
Patel: Frontier AI training compute scales 4x annually
“Right now we're scaling up the training of frontier systems, four X a year, approximately.”
Dwarkesh Patel Jun 18, 2025 ▶ 12:15
Prediction Open · timeframe Dec 2028
Patel: 4x annual compute scaling cannot continue beyond 2028
“If you look at things like how much energy is there in the country, how much how many chips can TSMC produce and what fraction of them are already being used by AI. Even if you look at like raw GDP, like how much money does the world have? How much wealth does…”
Dwarkesh Patel Jun 18, 2025 ▶ 12:27
Assertion Not checkable as stated
Patel: RL scaling is outpacing pre-training scaling
“RL scaling is happening much faster than even overall training scaling.”
Dwarkesh Patel Jun 18, 2025 ▶ 13:38
Prediction Not checkable as stated
Patel: Reinforcement learning 10x compute scaling can only continue for a year
“So already within the course of six months. RL compute has 10 X. That pace can only continue for a year, even if you build up all the RL environments before it's, you know, you're like, you've, you're at the frontier of training compute for these systems overa…”
Dwarkesh Patel Jun 18, 2025 ▶ 15:09
Insight
Patel: Training AI on math will not yield diplomatic acumen
“I just don't think you're going to train the AI so much on math that is going to learn how to do Henry Kissinger level diplomacy. I do think skills are somewhat more self-contained.”
Dwarkesh Patel Jun 18, 2025 ▶ 17:07
Insight
Patel: Multi-instance continual learning could create superintelligence without recursive coding
“Once continual learning is solved, you might have something that looks like a broadly deployed intelligence explosion, which is to say, That because if these models are broadly deployed to the economy, every copy that's like this copy is learning how to do plu…”
Dwarkesh Patel Jun 18, 2025 ▶ 18:16
Insight
Patel: Delayed Reward Signals Will Slow AI Progress on Long-Horizon Tasks
“Now if we're getting to the world where you got to like do a project for seven hours and then at the end of those seven hours, then we tell you, Hey, did you did you get this right? Then like the progress just goes on a bunch. Cause you've gone from like getti…”
Dwarkesh Patel Jun 18, 2025 ▶ 20:41
Assertion Supported
Patel: GPT-4-Level Training Costs Have Dropped 10x to 100x
“If you look at what it costs to train GBT for originally, I think it was like 20,008, 100 over the course of a hundred days. So I think it costs on the order of like half a million to a hundred million dollars, somewhere in that range. And I think you could tr…”
Dwarkesh Patel Jun 18, 2025 ▶ 22:46
Prediction Not checkable as stated
Dwarkesh Patel: Roughly 30% chance of an AI intelligence explosion
“I'm genuinely not sure how likely an intelligence explosion is. I don't know. I'd say like. 30% chance it happens, which is crazy by the way.”
Dwarkesh Patel Jun 18, 2025 ▶ 24:50
Insight
Patel: One-month lead in intelligence explosion yields winner-take-all dominance
“Look, if there is an intelligence explosion, it has a really tough dynamic because if you're a month ahead, you will kick off this loop much faster than anybody else. And what that means is that you will you will be a month ahead to super intelligence, but nob…”
Dwarkesh Patel Jun 18, 2025 ▶ 28:53
Opinion
Patel: OpenAI's o3 is currently the smartest AI model on the market
“I do think O three is the smartest model on the market right now.”
Dwarkesh Patel Jun 18, 2025 ▶ 38:32
Prediction Not checkable as stated
Patel: Continual-learning AI could replace white-collar jobs worth tens of trillions
“If you do self continue to learn it, I think like you could get rid of a lot of white collar jobs at that point. And what is that worth? Like at least tens of trillions of dollars, like the wages that are paid to white collar work.”
Dwarkesh Patel Jun 18, 2025 ▶ 44:00
Opinion
Patel: xAI is slightly behind leading labs but has high compute per employee
“I think they're a serious competitor. I just don't know much about what they're going to do next. I think they're like slightly behind the other labs but they've got a lot of compute per employee.”
Dwarkesh Patel Jun 18, 2025 ▶ 45:23
Opinion
Patel: Meta treats Llama as a toy rather than approaching AGI correctly
“I think they're treating it as like a sort of like toy within the meta universe. And I don't think that's the correct way to think about AGI.”
Dwarkesh Patel Jun 18, 2025 ▶ 46:18
Prediction Held up
Dwarkesh: Hyperscaler ASICs will come online from all major providers
“And so that just sets up a huge incentive for all these hyperscalers to build their own ASICs, their own accelerators that replace the Nvidia ones. Which I think will come online over the next few years from all of them.”
Dwarkesh Patel Jun 18, 2025 ▶ 47:57
Insight
Dwarkesh: AI adoption is bottlenecked by model capabilities, not token cost
“The reason they're not being more widely used is not because people cannot afford a couple bucks for a million tokens. The reason they're not being more widely used is just like, they fundamentally lack some capabilities. So I disagree with this focus on the c…”
Dwarkesh Patel Jun 18, 2025 ▶ 50:27
Prediction Not checkable as stated
Patel: AI deception could worsen as training tasks become less understandable
“The problem might get worse over time as We're trading these models on tasks we understand less and less well.”
Dwarkesh Patel Jun 18, 2025 ▶ 53:39
Prediction Not checkable as stated
Patel: Individuals will eventually train superintelligences in basements
“The cost of training, the systems is declining so fast that Literally you will be able to train a super intelligence in a basement at some point in the future.”
Dwarkesh Patel Jun 18, 2025 ▶ 56:49
Prediction Not checkable as stated
Patel: Misaligned intelligences will emerge, requiring societal resilience
“I think the long run picture is, That yes, there will be misaligned intelligences and we had to figure out a way to be robust to them.”
Dwarkesh Patel Jun 18, 2025 ▶ 58:07
Opinion
Patel: Effective Altruism's Reputation Remains in Tatters
“I do think the movement and the reputation of the movement is like still in tatters.”
Dwarkesh Patel Jun 18, 2025 ▶ 1:02:18
Assertion Supported
Patel: China Adds a US-Sized Power Grid Every Few Years
“And what's more important is that they're adding an America sized amount of power every couple of years. It might be more longer than every couple of years. Whereas our power production has stayed flat for the last many decades.”
Dwarkesh Patel Jun 18, 2025 ▶ 1:06:12
Insight
Patel: Future Economies Will Be Denominated in Compute, Not Human Labor
“More importantly, I think the future economy, once we do have these AI workers will be denominated in compute, right? Cause if computer's labor, right now, if you just think about like GDP per capita, because the individual worker is such an important componen…”
Dwarkesh Patel Jun 18, 2025 ▶ 1:07:11
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.