Mar 28, 2025 · 1h 15m · big-technology

AI's Rising Risks: Hacking, Virology, Loss of Control — With Dan Hendrycks

Dan Hendrycks · 52m spoken Alex Kantrowitz · 17m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of the Big Technology Podcast, host Alex Kantrowitz interviews AI safety researcher Dan Hendrycks about the full spectrum of artificial intelligence risks, ranging from near-term threats like bioweapon enablement and critical infrastructure cyberattacks to long-term existential challenges surrounding autonomous recursive self-improvement and geopolitical competition.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Alex holds 25.8% of the talking time here. How this is scored →

Alex as informed peer 5.0 Guest teaching 6.4 Guest disagreement 3.0 Alex pushing back 3.9
05100:0020:0040:001:00:005:59–9:18 · Alex as informed peer 3/10 Categorizing and Ranking AI Threat Timelines Kantrowitz asks Hendrycks to rank AI risks by severity. Hendrycks educates the host on distinguishing near-term malicious use (cyber and virology) from long-term loss of control driven by automated AI R&D loops.9:19–18:43 · Alex as informed peer 4/10 Virology Risks and Multimodal Wet Lab Guidance Kantrowitz questions how LLMs could create bioweapons if they cannot discover new compounds beyond their training set. Hendrycks schools him with upcoming research showing multimodal reasoning models achieving 90th percentile wet-lab protocol execution alongside MIT and Harvard virologists.18:43–22:19 · Alex as informed peer 5/10 Benchmarking Frontier AI with Humanity's Last Exam Kantrowitz inquires about Scale AI's post-training data collection with PhDs. Hendrycks explains the Humanity's Last Exam benchmark and notes how virology questions were deliberately excluded to avoid incentivizing dangerous model capabilities.22:19–29:17 · Alex as informed peer 6/10 Scaling Limits, the Reasoning Paradigm, and World Models Kantrowitz brings up Yann LeCun's arguments regarding scaling limits and lack of real-world physical models, citing video generator failures like haystacks exploding. Hendrycks brushes aside philosophical definitions of understanding as a 'no true Scotsman' fallacy and points out the steep slope of RL-driven reasoning models.29:19–36:40 · Alex as informed peer 5/10 Autonomous Cyberattacks, Critical Infrastructure, and Dual-Use Balance Kantrowitz pushes an optimistic angle on dual-use technology, arguing that offensive cyber capabilities imply equal defensive power. Hendrycks counters by detailing why critical infrastructure legacy systems make cyber offense-dominant.36:41–46:03 · Alex as informed peer 7/10 Corporate Arms Races, Game Theory, and xAI Mission Kantrowitz challenges the premise that AI labs build advanced models purely for safety and specifically pushes Hendrycks on Elon Musk's claims of political neutrality and anti-censorship on X. Hendrycks relies on game theory and security dilemmas, eventually declining to serve as Musk's corporate spokesperson.46:05–54:20 · Alex as informed peer 5/10 Superintelligence Strategy and Geopolitical Deterrence Hendrycks introduces his superintelligence deterrence strategy paper with Eric Schmidt and Alexandr Wang. Kantrowitz pushes back skeptically on whether China would ever cooperate and asks if Hendrycks aligns with Yudkowsky's data center bombing stance, prompting Hendrycks to advocate non-kinetic cyber deterrence.54:20–1:05:00 · Alex as informed peer 5/10 Mid-Roll Break and Podcast Re-Introduction Following the mid-roll break, Kantrowitz probes the mechanics of recursive self-improvement, White House briefings, and deceptive alignment. Hendrycks discusses circuit breakers and his recent paper demonstrating models lying under pressure.1:05:01–1:11:26 · Alex as informed peer 5/10 Center for AI Safety Funding and Critiquing Effective Altruism Kantrowitz asks about the Center for AI Safety's funding from Jaan Tallinn and ties to Effective Altruism. Hendrycks emphatically distances himself from EA, condemning the Berkeley alignment scene as a suffocating, fad-driven orthodoxy that dismissed malicious use.1:11:27–1:14:32 · Alex as informed peer 5/10 Open Source Weights, DeepSeek, and International Norms Kantrowitz asks how AI safety can survive when open-source models like DeepSeek release cutting-edge weights publicly. Hendrycks suggests establishing international nonproliferation norms restricting open weights once models achieve expert-level virology or critical cyber skills.5:59–9:18 · Guest teaching 6/10 Categorizing and Ranking AI Threat Timelines Kantrowitz asks Hendrycks to rank AI risks by severity. Hendrycks educates the host on distinguishing near-term malicious use (cyber and virology) from long-term loss of control driven by automated AI R&D loops.9:19–18:43 · Guest teaching 8/10 Virology Risks and Multimodal Wet Lab Guidance Kantrowitz questions how LLMs could create bioweapons if they cannot discover new compounds beyond their training set. Hendrycks schools him with upcoming research showing multimodal reasoning models achieving 90th percentile wet-lab protocol execution alongside MIT and Harvard virologists.18:43–22:19 · Guest teaching 6/10 Benchmarking Frontier AI with Humanity's Last Exam Kantrowitz inquires about Scale AI's post-training data collection with PhDs. Hendrycks explains the Humanity's Last Exam benchmark and notes how virology questions were deliberately excluded to avoid incentivizing dangerous model capabilities.22:19–29:17 · Guest teaching 7/10 Scaling Limits, the Reasoning Paradigm, and World Models Kantrowitz brings up Yann LeCun's arguments regarding scaling limits and lack of real-world physical models, citing video generator failures like haystacks exploding. Hendrycks brushes aside philosophical definitions of understanding as a 'no true Scotsman' fallacy and points out the steep slope of RL-driven reasoning models.29:19–36:40 · Guest teaching 7/10 Autonomous Cyberattacks, Critical Infrastructure, and Dual-Use Balance Kantrowitz pushes an optimistic angle on dual-use technology, arguing that offensive cyber capabilities imply equal defensive power. Hendrycks counters by detailing why critical infrastructure legacy systems make cyber offense-dominant.36:41–46:03 · Guest teaching 5/10 Corporate Arms Races, Game Theory, and xAI Mission Kantrowitz challenges the premise that AI labs build advanced models purely for safety and specifically pushes Hendrycks on Elon Musk's claims of political neutrality and anti-censorship on X. Hendrycks relies on game theory and security dilemmas, eventually declining to serve as Musk's corporate spokesperson.46:05–54:20 · Guest teaching 6/10 Superintelligence Strategy and Geopolitical Deterrence Hendrycks introduces his superintelligence deterrence strategy paper with Eric Schmidt and Alexandr Wang. Kantrowitz pushes back skeptically on whether China would ever cooperate and asks if Hendrycks aligns with Yudkowsky's data center bombing stance, prompting Hendrycks to advocate non-kinetic cyber deterrence.54:20–1:05:00 · Guest teaching 6/10 Mid-Roll Break and Podcast Re-Introduction Following the mid-roll break, Kantrowitz probes the mechanics of recursive self-improvement, White House briefings, and deceptive alignment. Hendrycks discusses circuit breakers and his recent paper demonstrating models lying under pressure.1:05:01–1:11:26 · Guest teaching 7/10 Center for AI Safety Funding and Critiquing Effective Altruism Kantrowitz asks about the Center for AI Safety's funding from Jaan Tallinn and ties to Effective Altruism. Hendrycks emphatically distances himself from EA, condemning the Berkeley alignment scene as a suffocating, fad-driven orthodoxy that dismissed malicious use.1:11:27–1:14:32 · Guest teaching 6/10 Open Source Weights, DeepSeek, and International Norms Kantrowitz asks how AI safety can survive when open-source models like DeepSeek release cutting-edge weights publicly. Hendrycks suggests establishing international nonproliferation norms restricting open weights once models achieve expert-level virology or critical cyber skills.5:59–9:18 · Guest disagreement 2/10 Categorizing and Ranking AI Threat Timelines Kantrowitz asks Hendrycks to rank AI risks by severity. Hendrycks educates the host on distinguishing near-term malicious use (cyber and virology) from long-term loss of control driven by automated AI R&D loops.9:19–18:43 · Guest disagreement 3/10 Virology Risks and Multimodal Wet Lab Guidance Kantrowitz questions how LLMs could create bioweapons if they cannot discover new compounds beyond their training set. Hendrycks schools him with upcoming research showing multimodal reasoning models achieving 90th percentile wet-lab protocol execution alongside MIT and Harvard virologists.18:43–22:19 · Guest disagreement 1/10 Benchmarking Frontier AI with Humanity's Last Exam Kantrowitz inquires about Scale AI's post-training data collection with PhDs. Hendrycks explains the Humanity's Last Exam benchmark and notes how virology questions were deliberately excluded to avoid incentivizing dangerous model capabilities.22:19–29:17 · Guest disagreement 5/10 Scaling Limits, the Reasoning Paradigm, and World Models Kantrowitz brings up Yann LeCun's arguments regarding scaling limits and lack of real-world physical models, citing video generator failures like haystacks exploding. Hendrycks brushes aside philosophical definitions of understanding as a 'no true Scotsman' fallacy and points out the steep slope of RL-driven reasoning models.29:19–36:40 · Guest disagreement 2/10 Autonomous Cyberattacks, Critical Infrastructure, and Dual-Use Balance Kantrowitz pushes an optimistic angle on dual-use technology, arguing that offensive cyber capabilities imply equal defensive power. Hendrycks counters by detailing why critical infrastructure legacy systems make cyber offense-dominant.36:41–46:03 · Guest disagreement 4/10 Corporate Arms Races, Game Theory, and xAI Mission Kantrowitz challenges the premise that AI labs build advanced models purely for safety and specifically pushes Hendrycks on Elon Musk's claims of political neutrality and anti-censorship on X. Hendrycks relies on game theory and security dilemmas, eventually declining to serve as Musk's corporate spokesperson.46:05–54:20 · Guest disagreement 3/10 Superintelligence Strategy and Geopolitical Deterrence Hendrycks introduces his superintelligence deterrence strategy paper with Eric Schmidt and Alexandr Wang. Kantrowitz pushes back skeptically on whether China would ever cooperate and asks if Hendrycks aligns with Yudkowsky's data center bombing stance, prompting Hendrycks to advocate non-kinetic cyber deterrence.54:20–1:05:00 · Guest disagreement 2/10 Mid-Roll Break and Podcast Re-Introduction Following the mid-roll break, Kantrowitz probes the mechanics of recursive self-improvement, White House briefings, and deceptive alignment. Hendrycks discusses circuit breakers and his recent paper demonstrating models lying under pressure.1:05:01–1:11:26 · Guest disagreement 6/10 Center for AI Safety Funding and Critiquing Effective Altruism Kantrowitz asks about the Center for AI Safety's funding from Jaan Tallinn and ties to Effective Altruism. Hendrycks emphatically distances himself from EA, condemning the Berkeley alignment scene as a suffocating, fad-driven orthodoxy that dismissed malicious use.1:11:27–1:14:32 · Guest disagreement 2/10 Open Source Weights, DeepSeek, and International Norms Kantrowitz asks how AI safety can survive when open-source models like DeepSeek release cutting-edge weights publicly. Hendrycks suggests establishing international nonproliferation norms restricting open weights once models achieve expert-level virology or critical cyber skills.5:59–9:18 · Alex pushing back 2/10 Categorizing and Ranking AI Threat Timelines Kantrowitz asks Hendrycks to rank AI risks by severity. Hendrycks educates the host on distinguishing near-term malicious use (cyber and virology) from long-term loss of control driven by automated AI R&D loops.9:19–18:43 · Alex pushing back 4/10 Virology Risks and Multimodal Wet Lab Guidance Kantrowitz questions how LLMs could create bioweapons if they cannot discover new compounds beyond their training set. Hendrycks schools him with upcoming research showing multimodal reasoning models achieving 90th percentile wet-lab protocol execution alongside MIT and Harvard virologists.18:43–22:19 · Alex pushing back 2/10 Benchmarking Frontier AI with Humanity's Last Exam Kantrowitz inquires about Scale AI's post-training data collection with PhDs. Hendrycks explains the Humanity's Last Exam benchmark and notes how virology questions were deliberately excluded to avoid incentivizing dangerous model capabilities.22:19–29:17 · Alex pushing back 5/10 Scaling Limits, the Reasoning Paradigm, and World Models Kantrowitz brings up Yann LeCun's arguments regarding scaling limits and lack of real-world physical models, citing video generator failures like haystacks exploding. Hendrycks brushes aside philosophical definitions of understanding as a 'no true Scotsman' fallacy and points out the steep slope of RL-driven reasoning models.29:19–36:40 · Alex pushing back 4/10 Autonomous Cyberattacks, Critical Infrastructure, and Dual-Use Balance Kantrowitz pushes an optimistic angle on dual-use technology, arguing that offensive cyber capabilities imply equal defensive power. Hendrycks counters by detailing why critical infrastructure legacy systems make cyber offense-dominant.36:41–46:03 · Alex pushing back 7/10 Corporate Arms Races, Game Theory, and xAI Mission Kantrowitz challenges the premise that AI labs build advanced models purely for safety and specifically pushes Hendrycks on Elon Musk's claims of political neutrality and anti-censorship on X. Hendrycks relies on game theory and security dilemmas, eventually declining to serve as Musk's corporate spokesperson.46:05–54:20 · Alex pushing back 6/10 Superintelligence Strategy and Geopolitical Deterrence Hendrycks introduces his superintelligence deterrence strategy paper with Eric Schmidt and Alexandr Wang. Kantrowitz pushes back skeptically on whether China would ever cooperate and asks if Hendrycks aligns with Yudkowsky's data center bombing stance, prompting Hendrycks to advocate non-kinetic cyber deterrence.54:20–1:05:00 · Alex pushing back 3/10 Mid-Roll Break and Podcast Re-Introduction Following the mid-roll break, Kantrowitz probes the mechanics of recursive self-improvement, White House briefings, and deceptive alignment. Hendrycks discusses circuit breakers and his recent paper demonstrating models lying under pressure.1:05:01–1:11:26 · Alex pushing back 3/10 Center for AI Safety Funding and Critiquing Effective Altruism Kantrowitz asks about the Center for AI Safety's funding from Jaan Tallinn and ties to Effective Altruism. Hendrycks emphatically distances himself from EA, condemning the Berkeley alignment scene as a suffocating, fad-driven orthodoxy that dismissed malicious use.1:11:27–1:14:32 · Alex pushing back 3/10 Open Source Weights, DeepSeek, and International Norms Kantrowitz asks how AI safety can survive when open-source models like DeepSeek release cutting-edge weights publicly. Hendrycks suggests establishing international nonproliferation norms restricting open weights once models achieve expert-level virology or critical cyber skills.

speaking balance: gold is Alex, purple is the guest (3 minute bins)

0:00 · Alex 56.3% · guest 43.7%0:00 · Alex 56.3% · guest 43.7%3:00 · Alex 28.1% · guest 71.9%3:00 · Alex 28.1% · guest 71.9%6:00 · Alex 4% · guest 96%6:00 · Alex 4% · guest 96%9:00 · Alex 50.7% · guest 49.3%9:00 · Alex 50.7% · guest 49.3%12:00 · Alex 26.1% · guest 73.9%12:00 · Alex 26.1% · guest 73.9%15:00 · Alex 22.4% · guest 77.6%15:00 · Alex 22.4% · guest 77.6%18:00 · Alex 25.1% · guest 74.9%18:00 · Alex 25.1% · guest 74.9%21:00 · Alex 37% · guest 63%21:00 · Alex 37% · guest 63%24:00 · Alex 13.7% · guest 86.3%24:00 · Alex 13.7% · guest 86.3%27:00 · Alex 47.3% · guest 52.7%27:00 · Alex 47.3% · guest 52.7%30:00 · Alex 8.5% · guest 91.5%30:00 · Alex 8.5% · guest 91.5%33:00 · Alex 23.3% · guest 76.7%33:00 · Alex 23.3% · guest 76.7%36:00 · Alex 18.7% · guest 81.3%36:00 · Alex 18.7% · guest 81.3%39:00 · Alex 42.5% · guest 57.5%39:00 · Alex 42.5% · guest 57.5%42:00 · Alex 36.5% · guest 63.5%42:00 · Alex 36.5% · guest 63.5%45:00 · Alex 60.1% · guest 39.9%45:00 · Alex 60.1% · guest 39.9%48:00 · Alex 2% · guest 98%48:00 · Alex 2% · guest 98%51:00 · Alex 10.5% · guest 89.5%51:00 · Alex 10.5% · guest 89.5%54:00 · Alex 24% · guest 76%54:00 · Alex 24% · guest 76%57:00 · Alex 5.9% · guest 94.1%57:00 · Alex 5.9% · guest 94.1%1:00:00 · Alex 24.8% · guest 75.2%1:00:00 · Alex 24.8% · guest 75.2%1:03:00 · Alex 18.9% · guest 81.1%1:03:00 · Alex 18.9% · guest 81.1%1:06:00 · Alex 23.7% · guest 76.3%1:06:00 · Alex 23.7% · guest 76.3%1:09:00 · Alex 19.6% · guest 80.4%1:09:00 · Alex 19.6% · guest 80.4%1:12:00 · Alex 16.3% · guest 83.7%1:12:00 · Alex 16.3% · guest 83.7%1:15:00 · Alex 17.6% · guest 82.4%1:15:00 · Alex 17.6% · guest 82.4%
Sharpest disagreement ▶ 1:08:15 Hendrycks forcefully denounces the Berkeley EA alignment orthodoxy

Hendrycks forcefully rejects the Effective Altruism consensus, describing Berkeley's alignment community as a suffocating monolith obsessed with speculative fads like ELK while ignoring practical malicious use.

Hardest push from Alex ▶ 43:20 Kantrowitz rejects framing of Musk as an unbiased truth-seeker

Kantrowitz directly refuses Hendrycks's framing of xAI and Grok as politically neutral, citing empirical evidence of Musk algorithmically promoting his own political views and boosting Trump on X.

Biggest teaching moment ▶ 10:31 Hendrycks reframes bio risks using multimodal wet-lab guidance data

Hendrycks dismantles Kantrowitz's assumption that LLMs only do web searches by revealing new empirical benchmark data showing reasoning models scoring in the 90th percentile on wet-lab virus manipulation protocols.

Alex holds their own ▶ 27:05 Kantrowitz pushes Yann LeCun's critique of physical world model limits

Kantrowitz demonstrates technical familiarity with the frontier debate by channeling Yann LeCun's thesis on the absence of real-world physics models and providing concrete examples from video generation failures.

the scores for every segment, with the reasoning behind each
ChapterTopicAlex as informed peerGuest teachingGuest disagreementAlex pushing backWhy
Categorizing and Ranking AI Threat Timelines 3622 Kantrowitz asks Hendrycks to rank AI risks by severity. Hendrycks educates the host on distinguishing near-term malicious use (cyber and virology) from long-term loss of control driven by automated AI R&D loops.
Virology Risks and Multimodal Wet Lab Guidance 4834 Kantrowitz questions how LLMs could create bioweapons if they cannot discover new compounds beyond their training set. Hendrycks schools him with upcoming research showing multimodal reasoning models achieving 90th percentile wet-lab protocol execution alongside MIT and Harvard virologists.
Benchmarking Frontier AI with Humanity's Last Exam 5612 Kantrowitz inquires about Scale AI's post-training data collection with PhDs. Hendrycks explains the Humanity's Last Exam benchmark and notes how virology questions were deliberately excluded to avoid incentivizing dangerous model capabilities.
Scaling Limits, the Reasoning Paradigm, and World Models 6755 Kantrowitz brings up Yann LeCun's arguments regarding scaling limits and lack of real-world physical models, citing video generator failures like haystacks exploding. Hendrycks brushes aside philosophical definitions of understanding as a 'no true Scotsman' fallacy and points out the steep slope of RL-driven reasoning models.
Autonomous Cyberattacks, Critical Infrastructure, and Dual-Use Balance 5724 Kantrowitz pushes an optimistic angle on dual-use technology, arguing that offensive cyber capabilities imply equal defensive power. Hendrycks counters by detailing why critical infrastructure legacy systems make cyber offense-dominant.
Corporate Arms Races, Game Theory, and xAI Mission 7547 Kantrowitz challenges the premise that AI labs build advanced models purely for safety and specifically pushes Hendrycks on Elon Musk's claims of political neutrality and anti-censorship on X. Hendrycks relies on game theory and security dilemmas, eventually declining to serve as Musk's corporate spokesperson.
Superintelligence Strategy and Geopolitical Deterrence 5636 Hendrycks introduces his superintelligence deterrence strategy paper with Eric Schmidt and Alexandr Wang. Kantrowitz pushes back skeptically on whether China would ever cooperate and asks if Hendrycks aligns with Yudkowsky's data center bombing stance, prompting Hendrycks to advocate non-kinetic cyber deterrence.
Mid-Roll Break and Podcast Re-Introduction 5623 Following the mid-roll break, Kantrowitz probes the mechanics of recursive self-improvement, White House briefings, and deceptive alignment. Hendrycks discusses circuit breakers and his recent paper demonstrating models lying under pressure.
Center for AI Safety Funding and Critiquing Effective Altruism 5763 Kantrowitz asks about the Center for AI Safety's funding from Jaan Tallinn and ties to Effective Altruism. Hendrycks emphatically distances himself from EA, condemning the Berkeley alignment scene as a suffocating, fad-driven orthodoxy that dismissed malicious use.
Open Source Weights, DeepSeek, and International Norms 5623 Kantrowitz asks how AI safety can survive when open-source models like DeepSeek release cutting-edge weights publicly. Hendrycks suggests establishing international nonproliferation norms restricting open weights once models achieve expert-level virology or critical cyber skills.

Statements from this episode (27)

Opinion
Hendrycks: Malicious human misuse is a bigger near-term AI risk than rogue AI
“I think those risks would potentially grow in time. I don't think they're as substantial now compared to just the malicious use sorts of risks.”
Dan Hendrycks Mar 28, 2025 ▶ 2:00
Prediction Not checkable as stated
Hendrycks: Automated AI R&D could compress a decade of progress into one year
“If you could have one AI do that, then you could make, you know, a 100,000 copies of these and have them perform research simultaneously. So that could lead to some very substantial acceleration in the rate of development. You might get a decade's worth of AI …”
Dan Hendrycks Mar 28, 2025 ▶ 4:04
Insight
Hendrycks: Fast-moving automated AI R&D loops are nearly impossible to de-risk
“Meanwhile, some automated AI research and development loop that goes extremely quickly with very little human oversight, that seems hard to de-risk and get the risks to be at a negligible level, just because it's so fast and so ah, and there's so little human …”
Dan Hendrycks Mar 28, 2025 ▶ 4:56
Opinion
Hendrycks: AI poses no immediate extinction risk because it cannot make PowerPoints
“So, I don't think AI poses a risk of extinction like today, ok? I don't think that they're powerful enough to do that. They, because they can't make PowerPoints yet, right? They don't have Agential skills. They can't accomplish tasks that require many hours to…”
Dan Hendrycks Mar 28, 2025 ▶ 6:12
Prediction Not checkable as stated
Hendrycks: Cyberattacks and bioweapons are top malicious AI risks within two years
“In the shorter term when AIs get more agential, I'd be concerned about AIs causing cyber attacks on critical infrastructure, possibly by as directed by a rogue actor. There'd also be the risk of AIs facilitating the development of bioweapons in particular pand…”
Dan Hendrycks Mar 28, 2025 ▶ 6:48
Insight
Hendrycks: Loss-of-Control Risk Stems from Competitive Pressure to Fully Automate AI R&D
“Loss of control risks, which I think primarily stem from people, an AI company trying to automate all of AI research and development, and they can't have humans check in on that process because that would slow them down too much. If you have a human do a week …”
Dan Hendrycks Mar 28, 2025 ▶ 7:25
Assertion Supported
Hendrycks: Recent AI reasoning models score in 90th percentile on wet-lab guidance
“We are finding that with the most recent reasoning models quite unlike the models from two years ago, like the initial GPT-IV, the most recent reasoning models are getting around 90th percentile compared to these expert level virologists in their area of exper…”
Dan Hendrycks Mar 28, 2025 ▶ 11:16
Assertion Not checkable as stated
Hendrycks: AI Has Brainstormed Ways to Make Viruses More Dangerous for Over a Year
“Them doing brainstorming to come up with ways to make viruses more dangerous, I think that's a capability that they've had for over a year, the brainstorming part, but the implementation part seems to be fairly different.”
Dan Hendrycks Mar 28, 2025 ▶ 11:54
Prediction Not checkable as stated
Hendrycks: Consensus on Expert-Level AI Bio Capabilities Coming Within Months
“So, I think in bio, actually, the I would not be surprised if in a few months there's a consensus that there expert level in many relevant ways in that we need to be doing something about that.”
Dan Hendrycks Mar 28, 2025 ▶ 12:06
Prediction Not checkable as stated
Hendrycks: Mastering Humanity's Last Exam will signal superhuman math capabilities
“When there's very high performance on, on that benchmark that would be suggestive of something that has, say, in the ballpark of superhuman mathematician capabilities. And so I think that would revolutionize the, academy quite substantially. Because all the th…”
Dan Hendrycks Mar 28, 2025 ▶ 21:12
Assertion Supported
Hendrycks: Top AI models score only 10% to 20% on Humanity's Last Exam
“They're in the ballpark of, like, 10 to 20% overall. They're the very best models.”
Dan Hendrycks Mar 28, 2025 ▶ 22:04
Opinion
Hendrycks: The next-token prediction paradigm in AI is running out of steam
“That sort of paradigm does seem like it's running out of steam. It has held for many, many orders of magnitude but the returns on doing that are lower.”
Dan Hendrycks Mar 28, 2025 ▶ 24:03
Assertion Not checkable as stated
Hendrycks: RL-based reasoning models are improving faster than pre-training did
“That is separate from the new reasoning paradigm that has emerged in the past year which is where you train models to on math and coding types of questions with reinforcement learning, and that has a very steep slope, and I don't see any signs of that slowing …”
Dan Hendrycks Mar 28, 2025 ▶ 24:15
Assertion Not checkable as stated
Hendrycks: Reasoning training on code and math generalizes to domains like law
“There has been a fair amount of generalization from training on coding and mathematics to other sorts of domains like law, for instance.”
Dan Hendrycks Mar 28, 2025 ▶ 26:15
Insight
Hendrycks: Biological AI capabilities are offense-dominant over defense
“So in, in bio, I think that is offense dominant. If somebody creates a virus, there's not necessarily a cure that it will immediately find for it. If it would help a rogue actor make A somewhat compelling virus. Now that could be enough to cause many millions …”
Dan Hendrycks Mar 28, 2025 ▶ 35:29
Insight
Hendrycks: Cyber AI in critical infrastructure is offense-dominant due to patching barriers
“Where in the context of critical infrastructure there the software is not updated rapidly. So even if you identify various vulnerabilities, there will not necessarily be a patch, because the system needs to always be on, or there are interoperability constrain…”
Dan Hendrycks Mar 28, 2025 ▶ 36:08
Insight
Hendrycks: AI labs spend over 90% of intellectual energy on scaling compute
“Over, you know, 90% of the intellectual energies that they're going to spend is actually, how can we afford the 10 X larger supercomputer? And, ah, that means being very competitive, speeding this up and making safety be some priority, but not necessarily a su…”
Dan Hendrycks Mar 28, 2025 ▶ 37:49
Insight
Hendrycks: Unilateral AI development pause makes no game theoretic sense
“An individual company pausing their development while others race ahead doesn't make game theoretic sense.”
Dan Hendrycks Mar 28, 2025 ▶ 39:16
Opinion
Hendrycks: Elon Musk's X successfully reduced self-censorship in US discourse
“I think that overall in terms of cultural influence and people being more disagreeable and doing less self-censoring has been has been successful. I think that was the main objective of it. And so I think I think that X had a large role to play there. So I don…”
Dan Hendrycks Mar 28, 2025 ▶ 44:14
Prediction Not checkable as stated
Hendrycks: Superintelligence race will incentivize preemptive sabotage of rival AI projects
“I think that form of competition would be very dangerous, and because there's a risk of loss of control, and because it might incentivize states to engage in preventive sabotage or preemptive sabotage to disable these sorts of projects.”
Dan Hendrycks Mar 28, 2025 ▶ 48:36
Prediction Not checkable as stated
Hendrycks: Russia and China may threaten US data centers over superintelligence progress
“If the US were pulling ahead both Russia and China may have a substantial interest in saying, hey, cut this out. Pulling ahead to develop super intelligence, which could give it a huge advantage and an ability to crush crush them. They'd say, you don't get to …”
Dan Hendrycks Mar 28, 2025 ▶ 49:57
Insight
Hendrycks: Automated AI R&D creates an insurmountable lead for early starters
“When you get to a different paradigm, like automated AI R&D, the slope might be extremely high, such that if the competitor starts to do automated AI R&D a year later, they may never catch up just because you're so far ahead and your gains are compounding on y…”
Dan Hendrycks Mar 28, 2025 ▶ 52:47
Prediction Not checkable as stated
Hendrycks: Geopolitical deterrence will delay superintelligence development
“This is why I think it might take a while for superintelligence to be developed, because there'll be deterrence around it later on.”
Dan Hendrycks Mar 28, 2025 ▶ 57:27
Assertion Supported
Hendrycks: AI models lied 20% to 60% of the time under pressure
“So we have a paper out last week, we're just measuring the extent to which they're deceptive. And in the scenarios we have, like all the models were in these sorts of scenarios under, you know, slight pressure to lie, not being told to lie, but just some sligh…”
Dan Hendrycks Mar 28, 2025 ▶ 1:03:58
Disclosure
Hendrycks: I earn $1 a year from xAI and $12 from Scale
“My appointment at XAI, I get a dollar a year. At Scale, I've at Scale AI, I've increased my salary exponentially to where I get 12 dollars a year, a dollar per month from Scale.”
Dan Hendrycks Mar 28, 2025 ▶ 1:05:40
Prediction Not checkable as stated
Hendrycks: AI labs will break voluntary commitments when facing commercial competition
“I think voluntary commitments from AI companies are also a distraction because the companies will, you should expect most of them by default to just break those sorts of commitments if they end up going up against economic competitiveness.”
Dan Hendrycks Mar 28, 2025 ▶ 1:09:18
Opinion
Hendrycks: Open-weight AI releases should be restricted for critical infrastructure cyber risks
“If for instance they have these cyber capabilities later on yeah, I think that, or I think that would be a potential place to be drawing the line on on open weight releases personally. In particular the ones that could cause damage to critical infrastructure.”
Dan Hendrycks Mar 28, 2025 ▶ 1:12:50
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.