Jan 28, 2026 · 1h 13m · latent-space

🔬 From Red Teaming GPT-4 to Automating Drug Discovery: The Future of AI in Science — Andrew White

Andrew White · 54m spoken Brandon Anderson · 8m spoken RJ Haneke · 4m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In the inaugural episode of the AI for Science podcast, hosts Brandon and RJ Haneke interview researcher-entrepreneur Andrew White on how autonomous agents, empirical verifiers, and natural language architectures are transforming the scientific method and outperforming traditional computational modeling.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The hosts as informed peer 4.4 Guest teaching 4.0 Guest disagreement 1.7 The hosts pushing back 1.9
05100:0015:0030:0045:001:00:000:00–2:20 · The hosts as informed peer 3/10 Molecular Dynamics vs. AlphaFold Paradigm Shift The episode opens with Andrew White reflecting on DE Shaw Research versus AlphaFold, followed by Brandon and RJ introducing the premise of the AI for Science podcast.2:22–6:39 · The hosts as informed peer 2/10 Introducing Andrew White and His Startup Ventures Andrew White recounts his academic path through biomaterials, molecular dynamics simulations, and maximum entropy modeling while the hosts ask conversational opening questions.6:39–9:12 · The hosts as informed peer 2/10 The Machine Learning Sabbatical and Early Chemistry Transformers Andrew explains his sabbatical at IPAM UCLA, authoring a textbook on geometric deep learning in chemistry, and early transformer benchmarks.9:12–11:35 · The hosts as informed peer 2/10 Red-Teaming GPT-4 and the Emergence of ChemCrow Andrew details red-teaming GPT-4 before release, developing ChemCrow with cloud labs, and briefing national security advisors at the White House.11:35–14:03 · The hosts as informed peer 3/10 Founding Future House, Resigning Tenure, and Launching Edison Andrew explains co-founding Future House with Sam Rodriques as a focused research organization, resigning tenure, and spinning out Edison Scientific.14:04–17:20 · The hosts as informed peer 4/10 The Accelerating Pace of Automating the Scientific Method Brandon presses Andrew on what 'automating science' actually means, prompting Andrew to differentiate between biological simulation models and automating cognitive discovery loops.17:20–20:05 · The hosts as informed peer 5/10 Identifying Scientific Bottlenecks and the Concept of Scientific Taste RJ frames the scientific method through systems bottlenecks in physical labs, leading Andrew to discuss lab-in-the-loop agents and the bottleneck of scientific taste.20:05–22:34 · The hosts as informed peer 4/10 Evaluating Hypotheses: RLHF Pitfalls and Objective Feedback Andrew describes why RLHF failed on hypothesis generation because humans overweighted prose and actionable details rather than world-changing impact.22:35–25:29 · The hosts as informed peer 4/10 Google Co-Scientist vs. Robin: Validating Drug Repurposing in the Wet Lab Andrew contrasts Google Co-Scientist tournament ranking with Future House's Robin agent, showing how ophthalmologists' preferred hypotheses failed while an unheralded candidate succeeded in the lab.25:29–29:23 · The hosts as informed peer 6/10 Data Provenance, Enumeration, and Verifier-in-the-Loop Discovery Brandon points out that biological hypotheses are cheap and experimental runway is scarce, challenging Andrew on how agents enrich for high-probability ideas.29:24–33:01 · The hosts as informed peer 6/10 Scientific Disagreement, Data Imputation, and Medicinal Chemistry Biases Brandon references Heather Kulik's interview regarding raw data contradicting paper conclusions. Andrew agrees and details human disagreements in benchmarks and medicinal chemistry superstitions around boron and fluorine.33:02–40:04 · The hosts as informed peer 5/10 Cosmos Benchmarks and the Subjectivity of Result Interpretation RJ pushes Andrew on Cosmos getting 50% accuracy on certain benchmark tasks. Andrew explains that was subjective human interpretation before detailing the lineage of tools leading to Cosmos.40:04–43:07 · The hosts as informed peer 5/10 The Critique of Molecular Dynamics and Density Functional Theory Andrew attacks molecular dynamics and DFT as overrated time-sinks that overfit to validation data, while Brandon backs him up with compute estimates on water simulations.43:07–46:26 · The hosts as informed peer 5/10 First Principles vs. Data-Driven ML: AlphaFold vs. DESRES Andrew compares DE Shaw's first-principles hardware approach against DeepMind's data-driven AlphaFold, highlighting that empirical data defeated physics simulations.46:26–49:17 · The hosts as informed peer 4/10 AI Safety and CBRN Risk Assessment in Chemistry and Biology Andrew reviews CBRN threats, arguing initial worries about text models uncovering secret bioweapon recipes were overblown because most information is already public.49:17–52:13 · The hosts as informed peer 5/10 Physical Constraints, Dual-Use Logistics, and Biosecurity Realities Brandon frames biosecurity around expert versus non-expert capabilities, while Andrew highlights physical centrifuge bottlenecks and subtle dual-use logistics like KYC circumvention.52:13–56:12 · The hosts as informed peer 5/10 Focused Research Organizations and Capital Allocation in Science Andrew discusses FRO economics and VC burn rates, while Brandon pushes back that seven-figure salaries remain a minor percentage compared to compute cluster expenditures.56:12–59:38 · The hosts as informed peer 6/10 The Essential Human Element in Scientific Interpretation and Curiosity RJ pushes back directly on Andrew's claim that human scientists remain indispensable, questioning why humans are fundamentally needed if end-to-end automation works.59:39–1:05:42 · The hosts as informed peer 7/10 Deployment and Enterprise Scaling of the Cosmos Platform Brandon challenges Andrew's assertion that chemistry is purely language, arguing that chemists naturally reason using graphs, bonds, and spatial geometry. Andrew illustrates why physical modeling escalates infinitely.1:05:42–1:08:15 · The hosts as informed peer 6/10 Expressing Complex Physics in Language and the Value of Strong Opinions RJ argues that quantum mechanics cannot be understood with natural language and requires math. Andrew defends linguistic formulation and the strategic power of taking strong opinions over optionality.1:08:15–1:13:32 · The hosts as informed peer 4/10 EtherZero: Verifiable Chemistry Rewards and Hilarious Reward Hacking Andrew shares funny reward-hacking anecdotes from training EtherZero, including the model inventing explosive six-nitrogen chains and stuffing inert purchasable nitrogen into reactions.0:00–2:20 · Guest teaching 1/10 Molecular Dynamics vs. AlphaFold Paradigm Shift The episode opens with Andrew White reflecting on DE Shaw Research versus AlphaFold, followed by Brandon and RJ introducing the premise of the AI for Science podcast.2:22–6:39 · Guest teaching 4/10 Introducing Andrew White and His Startup Ventures Andrew White recounts his academic path through biomaterials, molecular dynamics simulations, and maximum entropy modeling while the hosts ask conversational opening questions.6:39–9:12 · Guest teaching 3/10 The Machine Learning Sabbatical and Early Chemistry Transformers Andrew explains his sabbatical at IPAM UCLA, authoring a textbook on geometric deep learning in chemistry, and early transformer benchmarks.9:12–11:35 · Guest teaching 4/10 Red-Teaming GPT-4 and the Emergence of ChemCrow Andrew details red-teaming GPT-4 before release, developing ChemCrow with cloud labs, and briefing national security advisors at the White House.11:35–14:03 · Guest teaching 3/10 Founding Future House, Resigning Tenure, and Launching Edison Andrew explains co-founding Future House with Sam Rodriques as a focused research organization, resigning tenure, and spinning out Edison Scientific.14:04–17:20 · Guest teaching 4/10 The Accelerating Pace of Automating the Scientific Method Brandon presses Andrew on what 'automating science' actually means, prompting Andrew to differentiate between biological simulation models and automating cognitive discovery loops.17:20–20:05 · Guest teaching 4/10 Identifying Scientific Bottlenecks and the Concept of Scientific Taste RJ frames the scientific method through systems bottlenecks in physical labs, leading Andrew to discuss lab-in-the-loop agents and the bottleneck of scientific taste.20:05–22:34 · Guest teaching 4/10 Evaluating Hypotheses: RLHF Pitfalls and Objective Feedback Andrew describes why RLHF failed on hypothesis generation because humans overweighted prose and actionable details rather than world-changing impact.22:35–25:29 · Guest teaching 5/10 Google Co-Scientist vs. Robin: Validating Drug Repurposing in the Wet Lab Andrew contrasts Google Co-Scientist tournament ranking with Future House's Robin agent, showing how ophthalmologists' preferred hypotheses failed while an unheralded candidate succeeded in the lab.25:29–29:23 · Guest teaching 4/10 Data Provenance, Enumeration, and Verifier-in-the-Loop Discovery Brandon points out that biological hypotheses are cheap and experimental runway is scarce, challenging Andrew on how agents enrich for high-probability ideas.29:24–33:01 · Guest teaching 4/10 Scientific Disagreement, Data Imputation, and Medicinal Chemistry Biases Brandon references Heather Kulik's interview regarding raw data contradicting paper conclusions. Andrew agrees and details human disagreements in benchmarks and medicinal chemistry superstitions around boron and fluorine.33:02–40:04 · Guest teaching 5/10 Cosmos Benchmarks and the Subjectivity of Result Interpretation RJ pushes Andrew on Cosmos getting 50% accuracy on certain benchmark tasks. Andrew explains that was subjective human interpretation before detailing the lineage of tools leading to Cosmos.40:04–43:07 · Guest teaching 4/10 The Critique of Molecular Dynamics and Density Functional Theory Andrew attacks molecular dynamics and DFT as overrated time-sinks that overfit to validation data, while Brandon backs him up with compute estimates on water simulations.43:07–46:26 · Guest teaching 5/10 First Principles vs. Data-Driven ML: AlphaFold vs. DESRES Andrew compares DE Shaw's first-principles hardware approach against DeepMind's data-driven AlphaFold, highlighting that empirical data defeated physics simulations.46:26–49:17 · Guest teaching 5/10 AI Safety and CBRN Risk Assessment in Chemistry and Biology Andrew reviews CBRN threats, arguing initial worries about text models uncovering secret bioweapon recipes were overblown because most information is already public.49:17–52:13 · Guest teaching 4/10 Physical Constraints, Dual-Use Logistics, and Biosecurity Realities Brandon frames biosecurity around expert versus non-expert capabilities, while Andrew highlights physical centrifuge bottlenecks and subtle dual-use logistics like KYC circumvention.52:13–56:12 · Guest teaching 4/10 Focused Research Organizations and Capital Allocation in Science Andrew discusses FRO economics and VC burn rates, while Brandon pushes back that seven-figure salaries remain a minor percentage compared to compute cluster expenditures.56:12–59:38 · Guest teaching 4/10 The Essential Human Element in Scientific Interpretation and Curiosity RJ pushes back directly on Andrew's claim that human scientists remain indispensable, questioning why humans are fundamentally needed if end-to-end automation works.59:39–1:05:42 · Guest teaching 5/10 Deployment and Enterprise Scaling of the Cosmos Platform Brandon challenges Andrew's assertion that chemistry is purely language, arguing that chemists naturally reason using graphs, bonds, and spatial geometry. Andrew illustrates why physical modeling escalates infinitely.1:05:42–1:08:15 · Guest teaching 4/10 Expressing Complex Physics in Language and the Value of Strong Opinions RJ argues that quantum mechanics cannot be understood with natural language and requires math. Andrew defends linguistic formulation and the strategic power of taking strong opinions over optionality.1:08:15–1:13:32 · Guest teaching 5/10 EtherZero: Verifiable Chemistry Rewards and Hilarious Reward Hacking Andrew shares funny reward-hacking anecdotes from training EtherZero, including the model inventing explosive six-nitrogen chains and stuffing inert purchasable nitrogen into reactions.0:00–2:20 · Guest disagreement 0/10 Molecular Dynamics vs. AlphaFold Paradigm Shift The episode opens with Andrew White reflecting on DE Shaw Research versus AlphaFold, followed by Brandon and RJ introducing the premise of the AI for Science podcast.2:22–6:39 · Guest disagreement 0/10 Introducing Andrew White and His Startup Ventures Andrew White recounts his academic path through biomaterials, molecular dynamics simulations, and maximum entropy modeling while the hosts ask conversational opening questions.6:39–9:12 · Guest disagreement 0/10 The Machine Learning Sabbatical and Early Chemistry Transformers Andrew explains his sabbatical at IPAM UCLA, authoring a textbook on geometric deep learning in chemistry, and early transformer benchmarks.9:12–11:35 · Guest disagreement 0/10 Red-Teaming GPT-4 and the Emergence of ChemCrow Andrew details red-teaming GPT-4 before release, developing ChemCrow with cloud labs, and briefing national security advisors at the White House.11:35–14:03 · Guest disagreement 0/10 Founding Future House, Resigning Tenure, and Launching Edison Andrew explains co-founding Future House with Sam Rodriques as a focused research organization, resigning tenure, and spinning out Edison Scientific.14:04–17:20 · Guest disagreement 1/10 The Accelerating Pace of Automating the Scientific Method Brandon presses Andrew on what 'automating science' actually means, prompting Andrew to differentiate between biological simulation models and automating cognitive discovery loops.17:20–20:05 · Guest disagreement 1/10 Identifying Scientific Bottlenecks and the Concept of Scientific Taste RJ frames the scientific method through systems bottlenecks in physical labs, leading Andrew to discuss lab-in-the-loop agents and the bottleneck of scientific taste.20:05–22:34 · Guest disagreement 1/10 Evaluating Hypotheses: RLHF Pitfalls and Objective Feedback Andrew describes why RLHF failed on hypothesis generation because humans overweighted prose and actionable details rather than world-changing impact.22:35–25:29 · Guest disagreement 1/10 Google Co-Scientist vs. Robin: Validating Drug Repurposing in the Wet Lab Andrew contrasts Google Co-Scientist tournament ranking with Future House's Robin agent, showing how ophthalmologists' preferred hypotheses failed while an unheralded candidate succeeded in the lab.25:29–29:23 · Guest disagreement 1/10 Data Provenance, Enumeration, and Verifier-in-the-Loop Discovery Brandon points out that biological hypotheses are cheap and experimental runway is scarce, challenging Andrew on how agents enrich for high-probability ideas.29:24–33:01 · Guest disagreement 2/10 Scientific Disagreement, Data Imputation, and Medicinal Chemistry Biases Brandon references Heather Kulik's interview regarding raw data contradicting paper conclusions. Andrew agrees and details human disagreements in benchmarks and medicinal chemistry superstitions around boron and fluorine.33:02–40:04 · Guest disagreement 2/10 Cosmos Benchmarks and the Subjectivity of Result Interpretation RJ pushes Andrew on Cosmos getting 50% accuracy on certain benchmark tasks. Andrew explains that was subjective human interpretation before detailing the lineage of tools leading to Cosmos.40:04–43:07 · Guest disagreement 6/10 The Critique of Molecular Dynamics and Density Functional Theory Andrew attacks molecular dynamics and DFT as overrated time-sinks that overfit to validation data, while Brandon backs him up with compute estimates on water simulations.43:07–46:26 · Guest disagreement 3/10 First Principles vs. Data-Driven ML: AlphaFold vs. DESRES Andrew compares DE Shaw's first-principles hardware approach against DeepMind's data-driven AlphaFold, highlighting that empirical data defeated physics simulations.46:26–49:17 · Guest disagreement 2/10 AI Safety and CBRN Risk Assessment in Chemistry and Biology Andrew reviews CBRN threats, arguing initial worries about text models uncovering secret bioweapon recipes were overblown because most information is already public.49:17–52:13 · Guest disagreement 1/10 Physical Constraints, Dual-Use Logistics, and Biosecurity Realities Brandon frames biosecurity around expert versus non-expert capabilities, while Andrew highlights physical centrifuge bottlenecks and subtle dual-use logistics like KYC circumvention.52:13–56:12 · Guest disagreement 2/10 Focused Research Organizations and Capital Allocation in Science Andrew discusses FRO economics and VC burn rates, while Brandon pushes back that seven-figure salaries remain a minor percentage compared to compute cluster expenditures.56:12–59:38 · Guest disagreement 3/10 The Essential Human Element in Scientific Interpretation and Curiosity RJ pushes back directly on Andrew's claim that human scientists remain indispensable, questioning why humans are fundamentally needed if end-to-end automation works.59:39–1:05:42 · Guest disagreement 3/10 Deployment and Enterprise Scaling of the Cosmos Platform Brandon challenges Andrew's assertion that chemistry is purely language, arguing that chemists naturally reason using graphs, bonds, and spatial geometry. Andrew illustrates why physical modeling escalates infinitely.1:05:42–1:08:15 · Guest disagreement 4/10 Expressing Complex Physics in Language and the Value of Strong Opinions RJ argues that quantum mechanics cannot be understood with natural language and requires math. Andrew defends linguistic formulation and the strategic power of taking strong opinions over optionality.1:08:15–1:13:32 · Guest disagreement 2/10 EtherZero: Verifiable Chemistry Rewards and Hilarious Reward Hacking Andrew shares funny reward-hacking anecdotes from training EtherZero, including the model inventing explosive six-nitrogen chains and stuffing inert purchasable nitrogen into reactions.0:00–2:20 · The hosts pushing back 0/10 Molecular Dynamics vs. AlphaFold Paradigm Shift The episode opens with Andrew White reflecting on DE Shaw Research versus AlphaFold, followed by Brandon and RJ introducing the premise of the AI for Science podcast.2:22–6:39 · The hosts pushing back 0/10 Introducing Andrew White and His Startup Ventures Andrew White recounts his academic path through biomaterials, molecular dynamics simulations, and maximum entropy modeling while the hosts ask conversational opening questions.6:39–9:12 · The hosts pushing back 0/10 The Machine Learning Sabbatical and Early Chemistry Transformers Andrew explains his sabbatical at IPAM UCLA, authoring a textbook on geometric deep learning in chemistry, and early transformer benchmarks.9:12–11:35 · The hosts pushing back 0/10 Red-Teaming GPT-4 and the Emergence of ChemCrow Andrew details red-teaming GPT-4 before release, developing ChemCrow with cloud labs, and briefing national security advisors at the White House.11:35–14:03 · The hosts pushing back 0/10 Founding Future House, Resigning Tenure, and Launching Edison Andrew explains co-founding Future House with Sam Rodriques as a focused research organization, resigning tenure, and spinning out Edison Scientific.14:04–17:20 · The hosts pushing back 3/10 The Accelerating Pace of Automating the Scientific Method Brandon presses Andrew on what 'automating science' actually means, prompting Andrew to differentiate between biological simulation models and automating cognitive discovery loops.17:20–20:05 · The hosts pushing back 2/10 Identifying Scientific Bottlenecks and the Concept of Scientific Taste RJ frames the scientific method through systems bottlenecks in physical labs, leading Andrew to discuss lab-in-the-loop agents and the bottleneck of scientific taste.20:05–22:34 · The hosts pushing back 1/10 Evaluating Hypotheses: RLHF Pitfalls and Objective Feedback Andrew describes why RLHF failed on hypothesis generation because humans overweighted prose and actionable details rather than world-changing impact.22:35–25:29 · The hosts pushing back 1/10 Google Co-Scientist vs. Robin: Validating Drug Repurposing in the Wet Lab Andrew contrasts Google Co-Scientist tournament ranking with Future House's Robin agent, showing how ophthalmologists' preferred hypotheses failed while an unheralded candidate succeeded in the lab.25:29–29:23 · The hosts pushing back 3/10 Data Provenance, Enumeration, and Verifier-in-the-Loop Discovery Brandon points out that biological hypotheses are cheap and experimental runway is scarce, challenging Andrew on how agents enrich for high-probability ideas.29:24–33:01 · The hosts pushing back 2/10 Scientific Disagreement, Data Imputation, and Medicinal Chemistry Biases Brandon references Heather Kulik's interview regarding raw data contradicting paper conclusions. Andrew agrees and details human disagreements in benchmarks and medicinal chemistry superstitions around boron and fluorine.33:02–40:04 · The hosts pushing back 4/10 Cosmos Benchmarks and the Subjectivity of Result Interpretation RJ pushes Andrew on Cosmos getting 50% accuracy on certain benchmark tasks. Andrew explains that was subjective human interpretation before detailing the lineage of tools leading to Cosmos.40:04–43:07 · The hosts pushing back 1/10 The Critique of Molecular Dynamics and Density Functional Theory Andrew attacks molecular dynamics and DFT as overrated time-sinks that overfit to validation data, while Brandon backs him up with compute estimates on water simulations.43:07–46:26 · The hosts pushing back 2/10 First Principles vs. Data-Driven ML: AlphaFold vs. DESRES Andrew compares DE Shaw's first-principles hardware approach against DeepMind's data-driven AlphaFold, highlighting that empirical data defeated physics simulations.46:26–49:17 · The hosts pushing back 1/10 AI Safety and CBRN Risk Assessment in Chemistry and Biology Andrew reviews CBRN threats, arguing initial worries about text models uncovering secret bioweapon recipes were overblown because most information is already public.49:17–52:13 · The hosts pushing back 2/10 Physical Constraints, Dual-Use Logistics, and Biosecurity Realities Brandon frames biosecurity around expert versus non-expert capabilities, while Andrew highlights physical centrifuge bottlenecks and subtle dual-use logistics like KYC circumvention.52:13–56:12 · The hosts pushing back 3/10 Focused Research Organizations and Capital Allocation in Science Andrew discusses FRO economics and VC burn rates, while Brandon pushes back that seven-figure salaries remain a minor percentage compared to compute cluster expenditures.56:12–59:38 · The hosts pushing back 5/10 The Essential Human Element in Scientific Interpretation and Curiosity RJ pushes back directly on Andrew's claim that human scientists remain indispensable, questioning why humans are fundamentally needed if end-to-end automation works.59:39–1:05:42 · The hosts pushing back 4/10 Deployment and Enterprise Scaling of the Cosmos Platform Brandon challenges Andrew's assertion that chemistry is purely language, arguing that chemists naturally reason using graphs, bonds, and spatial geometry. Andrew illustrates why physical modeling escalates infinitely.1:05:42–1:08:15 · The hosts pushing back 4/10 Expressing Complex Physics in Language and the Value of Strong Opinions RJ argues that quantum mechanics cannot be understood with natural language and requires math. Andrew defends linguistic formulation and the strategic power of taking strong opinions over optionality.1:08:15–1:13:32 · The hosts pushing back 1/10 EtherZero: Verifiable Chemistry Rewards and Hilarious Reward Hacking Andrew shares funny reward-hacking anecdotes from training EtherZero, including the model inventing explosive six-nitrogen chains and stuffing inert purchasable nitrogen into reactions.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 0% · guest 100%0:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%6:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%9:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%12:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%15:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%18:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%21:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%24:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%27:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%30:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%33:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%36:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%39:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%42:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%45:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%48:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%51:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%54:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%57:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 100%1:00:00 · the hosts 0% · guest 100%1:03:00 · the hosts 0% · guest 100%1:03:00 · the hosts 0% · guest 100%1:06:00 · the hosts 0% · guest 100%1:06:00 · the hosts 0% · guest 100%1:09:00 · the hosts 0% · guest 100%1:09:00 · the hosts 0% · guest 100%1:12:00 · the hosts 0% · guest 100%1:12:00 · the hosts 0% · guest 100%
Sharpest disagreement ▶ 40:09 Andrew's broadside against MD and DFT simulation fields

Andrew bluntly dismisses molecular dynamics and density functional theory as overrated fields that have devoured countless PhD careers without delivering de novo real-world discoveries.

Hardest push from the hosts ▶ 57:20 RJ challenges the necessity of human involvement in science

RJ explicitly adopts a contrarian stance and rejects Andrew's premise that humans must remain in the scientific loop once autonomous discovery systems scale.

Biggest teaching moment ▶ 23:45 Expert clinical intuition fails empirical wet lab validation

Andrew reveals that expert ophthalmologists completely misjudged which drug repurposing hypothesis would work, proving that empirical lab loops outperform top human scientific opinion.

The host holds their own ▶ 1:02:35 Brandon challenges the 'chemistry as language' doctrine with graph representations

Brandon uses his domain expertise in structural biology to challenge Andrew's thesis, demonstrating that chemists fundamentally think in graphical representations and geometric abstractions rather than text strings.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Molecular Dynamics vs. AlphaFold Paradigm Shift 3100 The episode opens with Andrew White reflecting on DE Shaw Research versus AlphaFold, followed by Brandon and RJ introducing the premise of the AI for Science podcast.
Introducing Andrew White and His Startup Ventures 2400 Andrew White recounts his academic path through biomaterials, molecular dynamics simulations, and maximum entropy modeling while the hosts ask conversational opening questions.
The Machine Learning Sabbatical and Early Chemistry Transformers 2300 Andrew explains his sabbatical at IPAM UCLA, authoring a textbook on geometric deep learning in chemistry, and early transformer benchmarks.
Red-Teaming GPT-4 and the Emergence of ChemCrow 2400 Andrew details red-teaming GPT-4 before release, developing ChemCrow with cloud labs, and briefing national security advisors at the White House.
Founding Future House, Resigning Tenure, and Launching Edison 3300 Andrew explains co-founding Future House with Sam Rodriques as a focused research organization, resigning tenure, and spinning out Edison Scientific.
The Accelerating Pace of Automating the Scientific Method 4413 Brandon presses Andrew on what 'automating science' actually means, prompting Andrew to differentiate between biological simulation models and automating cognitive discovery loops.
Identifying Scientific Bottlenecks and the Concept of Scientific Taste 5412 RJ frames the scientific method through systems bottlenecks in physical labs, leading Andrew to discuss lab-in-the-loop agents and the bottleneck of scientific taste.
Evaluating Hypotheses: RLHF Pitfalls and Objective Feedback 4411 Andrew describes why RLHF failed on hypothesis generation because humans overweighted prose and actionable details rather than world-changing impact.
Google Co-Scientist vs. Robin: Validating Drug Repurposing in the Wet Lab 4511 Andrew contrasts Google Co-Scientist tournament ranking with Future House's Robin agent, showing how ophthalmologists' preferred hypotheses failed while an unheralded candidate succeeded in the lab.
Data Provenance, Enumeration, and Verifier-in-the-Loop Discovery 6413 Brandon points out that biological hypotheses are cheap and experimental runway is scarce, challenging Andrew on how agents enrich for high-probability ideas.
Scientific Disagreement, Data Imputation, and Medicinal Chemistry Biases 6422 Brandon references Heather Kulik's interview regarding raw data contradicting paper conclusions. Andrew agrees and details human disagreements in benchmarks and medicinal chemistry superstitions around boron and fluorine.
Cosmos Benchmarks and the Subjectivity of Result Interpretation 5524 RJ pushes Andrew on Cosmos getting 50% accuracy on certain benchmark tasks. Andrew explains that was subjective human interpretation before detailing the lineage of tools leading to Cosmos.
The Critique of Molecular Dynamics and Density Functional Theory 5461 Andrew attacks molecular dynamics and DFT as overrated time-sinks that overfit to validation data, while Brandon backs him up with compute estimates on water simulations.
First Principles vs. Data-Driven ML: AlphaFold vs. DESRES 5532 Andrew compares DE Shaw's first-principles hardware approach against DeepMind's data-driven AlphaFold, highlighting that empirical data defeated physics simulations.
AI Safety and CBRN Risk Assessment in Chemistry and Biology 4521 Andrew reviews CBRN threats, arguing initial worries about text models uncovering secret bioweapon recipes were overblown because most information is already public.
Physical Constraints, Dual-Use Logistics, and Biosecurity Realities 5412 Brandon frames biosecurity around expert versus non-expert capabilities, while Andrew highlights physical centrifuge bottlenecks and subtle dual-use logistics like KYC circumvention.
Focused Research Organizations and Capital Allocation in Science 5423 Andrew discusses FRO economics and VC burn rates, while Brandon pushes back that seven-figure salaries remain a minor percentage compared to compute cluster expenditures.
The Essential Human Element in Scientific Interpretation and Curiosity 6435 RJ pushes back directly on Andrew's claim that human scientists remain indispensable, questioning why humans are fundamentally needed if end-to-end automation works.
Deployment and Enterprise Scaling of the Cosmos Platform 7534 Brandon challenges Andrew's assertion that chemistry is purely language, arguing that chemists naturally reason using graphs, bonds, and spatial geometry. Andrew illustrates why physical modeling escalates infinitely.
Expressing Complex Physics in Language and the Value of Strong Opinions 6444 RJ argues that quantum mechanics cannot be understood with natural language and requires math. Andrew defends linguistic formulation and the strategic power of taking strong opinions over optionality.
EtherZero: Verifiable Chemistry Rewards and Hilarious Reward Hacking 4521 Andrew shares funny reward-hacking anecdotes from training EtherZero, including the model inventing explosive six-nitrogen chains and stuffing inert purchasable nitrogen into reactions.

Statements from this episode (50)

Assertion Not checkable as stated
White: DE Shaw Research Had Similar or Greater Funding Than DeepMind
“They had, you know, similar funding to DeepMind, probably more actually.”
Andrew White Jan 28, 2026 ▶ 0:09
Opinion
White: Molecular dynamics is reviving to generate simulation data
“Molecular dynamics which I think is actually suddenly becoming interesting again, as everyone's looking for ways to generate data from first principle simulation, and molecular dynamics, you know, covers basically everything that's molecules moving around in d…”
Andrew White Jan 28, 2026 ▶ 3:18
Insight
White: Maximum entropy modeling is the inverse of machine learning
“This theory called maximum entropy. And it's about like, how do you take complex simulations and match them to limited observations? And it's like the inverse of machine learning. Machine learning is like, you have simple models, you're going to a lot of data …”
Andrew White Jan 28, 2026 ▶ 5:44
Insight
White: Machine learning in chemistry centers on graphs, symmetry, and geometry
“In chemistry, it's all about graphs, right? It's all about how do you represent these graph structures. It's all about symmetry and geometry.”
Andrew White Jan 28, 2026 ▶ 7:41
Assertion Supported
White: OpenAI reached out to red team new models after reading his chemistry paper
“And then opening eye, some people there Lama was there. She saw this paper, and they reached out, like, hey, we're building this new model, and we think it'd be great to red team it to see, like, what could happen with these models if they're applied to chemis…”
Andrew White Jan 28, 2026 ▶ 8:58
Disclosure
Andrew White Began Red-Teaming GPT-4 in August 2022
“And so I was a red teamer for GPT-IV and I was using it like nine months with me for release with August. So GPT-IV came out in March and I was using it in August.”
Andrew White Jan 28, 2026 ▶ 9:13
Assertion Open · timeframe Jan 2026
White: ChemCrow paper was presented to U.S. President in 30-minute block
“I ended up visiting the white house. I guess my paper was like the only time a preprint or peer review paper was presented to the president on like their schedule for like a 30 minute block.”
Andrew White Jan 28, 2026 ▶ 10:21
Assertion Supported
White: Sam Rodriques Conceived FRO Model with Schmidt and Kalil
“Sam had been talking to Eric Schmidt and Tom Khalil, who was also a National Security Counsel at the Obama Administration, about how to like, scale up these ideas. And so Sam had this concept of, like, focused research organizations”
Andrew White Jan 28, 2026 ▶ 11:36
Disclosure
Andrew White Resigned University Tenure to Co-Found Edison
“I did resign my tenure position when we co-founded Edison.”
Andrew White Jan 28, 2026 ▶ 12:41
Opinion
White: AI for Science Is Difficult in Academia and Demands Bigger Bets
“AI over science is just, I think, A, difficult to do in academia, and B, so exciting, but I think you can take bigger bets, and I think having a tenured position and writing research grants is maybe not the biggest bet you can take on, on a field.”
Andrew White Jan 28, 2026 ▶ 13:08
Disclosure
White: Venture-Backed Startup Edison Was Spun Out of Future House
“Now we have a venture-backed startup called Edison, which we spun out of Future House and, you know, we took a lot of the ideas and we're trying to do this at an even bigger scale right now.”
Andrew White Jan 28, 2026 ▶ 13:31
Opinion
White: Existing LLMs Are Already Capable of Automating Much of Science
“We can actually automate so much of the scientific method, because it turns out, especially in a field like biology, which is very empirical limited, you know, the top one percent guesser of, you know, what they think will happen in experiment that, you know, …”
Andrew White Jan 28, 2026 ▶ 15:29
Disclosure
White: Future House Focuses on Discovery Loops Over System Modeling
“We're trying to automate the like cognitive process of scientific discovery, making hypotheses. Choosing experiments to do, analyzing the results from experiments and using it to update your hypotheses or your confidence in those hypotheses, and then leading t…”
Andrew White Jan 28, 2026 ▶ 16:01
Insight
White: AI Science Agents Do Not Require Bespoke Automated Labs
“We don't actually have to hold their hands so much anymore, or like they don't actually need to necessarily have an automated lab. They can like write an email to a CRO or something, or they can like tell you what experiment to do. And you can take a video of …”
Andrew White Jan 28, 2026 ▶ 16:56
Insight
White: Simulating biological systems hits fundamental limits; empirical lab measurement is unavoidable
“Basically the best model whatever, Opus seven or GPT 10, like it really can only propose the first experiment, maybe slightly more clever, but at a certain point you just need information, right? Like some little calculations you can do that, like, there's mor…”
Andrew White Jan 28, 2026 ▶ 17:53
Opinion
Logistics Data Matters More Than Frontier LLM Intelligence in Lab Science
“I think whether GPT, 5.2 codex max or Opus 4.5 is going to do better. It's probably doesn't matter. It's just a matter of like, which one's going to have all the information about what's in the lab and how much will it cost? How long will it take?”
Andrew White Jan 28, 2026 ▶ 18:54
Insight
White: AI models struggle to capture 'scientific taste' and distinguish exciting results
“Models don't capture that so well about knowing what is an exciting result and what is a boring result. So I think that's like a scientific taste.”
Andrew White Jan 28, 2026 ▶ 19:45
Insight
White: RLHF fails on scientific hypotheses by ignoring impact and information gain
“We learned a lot about how bad our LHF is with people, just like people pay really attention to the tone, to the details, to like how many specific facts or figures on the hypothesis, right? Like actionability about like if the experiment is feasible, but what…”
Andrew White Jan 28, 2026 ▶ 20:35
Assertion Supported
White: Human expert hypothesis rankings failed to predict wet-lab success
“And one of the things that came out of the Robin paper is that the hypothesis that people thought was best was not the one that led to success in that paper.”
Andrew White Jan 28, 2026 ▶ 23:18
Insight
White: Verifiers in the loop provide higher signal than hypothesis rankings
“I have a lot more faith in these, like verifier-in-the-loop kind of scenarios where you have either data analysis, literature search, or you're running a unit test, or whatever, you're going and running the experiment. Anything like that, I think, is going to …”
Andrew White Jan 28, 2026 ▶ 24:58
Assertion Supported
White: PaperQA outputs page citations for every generated sentence
“Paper QA was has, like, every sentence that it outputs has a citation to a page, right?”
Andrew White Jan 28, 2026 ▶ 26:03
Assertion Supported
White: Co-Scientist uses LLM ranking while Robin filters with data and literature
“In co-scientists, their filtration process was other LLMs sort of ranking it with rubrics or like personas. And our filtration process was like literature search and data analysis.”
Andrew White Jan 28, 2026 ▶ 26:58
Insight
White: AI's advantage over humans is rapidly testing more enumerated ideas
“And I think that's the easy way to succeed in an AI over humans is you can try more ideas faster.”
Andrew White Jan 28, 2026 ▶ 27:19
Opinion
White: LLMs Filter Scientific Hypotheses as Well as Human Domain Experts
“Nowadays, I actually would argue that if you go to an LLM and you ask it to evaluate, you know, hypotheses, including some garbage ones, it will probably do as good of a job as an expert in the field and filtering them out. That's not always the case.”
Andrew White Jan 28, 2026 ▶ 28:44
Assertion Partly supported
White: AI hits 60-70% on Bixbench, matching human expert agreement
“We're getting to 60%, 70% correctness on Bixbench, and we found that actually we're at the point where humans disagree at this level. Like, humans only agree 70% of the analysis, and so it's true that, like, when it comes to analyzing data, like, humans do not…”
Andrew White Jan 28, 2026 ▶ 30:33
Opinion
White: Medicinal chemistry is an AI-resistant field driven by superstition
“You want to know what the real modern dark arts are that like AI resistant area of the world is like a medicinal chemistry that is like, The spot where, like, you know, there's so much superstition.”
Andrew White Jan 28, 2026 ▶ 31:16
Assertion Supported
White: Cosmos's 50% benchmark score stems from subjective result interpretation
“That number that was 50% that came from cosmos's interpretation of some of the analysis. So like it might go in literature and find this result. And then it would say, wow, this is super exciting. This is amazing. Or it might do data analysis, but this is a no…”
Andrew White Jan 28, 2026 ▶ 33:59
Assertion Not checkable as stated
White: Frontier AI models cannot handle molecules well
“We noticed that the frontier models can't work with molecules very well. So let's make a model with intuition for medicinal chemistry. And that was what led to ether zero.”
Andrew White Jan 28, 2026 ▶ 35:49
Insight
White: A world model serves as unifying glue for scientific agents
“In cosmos, we basically, we had all the pieces sitting around. We working on world models, we working on data analysis agent, working on literature agent. And then we're working on, you know, we built a platform for scientific agents. So we had things that can…”
Andrew White Jan 28, 2026 ▶ 38:21
Insight
White: A scientific agent world model is analogous to a Git repo
“Your Git repo is like a distillation of all of the work that people put into the PRs, into the commits, and so I think there's a nice analogy between a Git repo and what a world model is.”
Andrew White Jan 28, 2026 ▶ 39:32
Opinion
White: Molecular dynamics and DFT are overrated in materials science
“So, I think molecular dynamics is overrated. And DFT is overrated. In fact, DFT may be even more overrated than like the dynamics.”
Andrew White Jan 28, 2026 ▶ 40:12
Assertion Not checkable as stated
Anderson: 20% of pre-ChatGPT global compute went to simulating water
“Once I did an estimate, I think pre like ChatGPT, something like 20% of the world's computing power just went to simulating water.”
Brandon Anderson Jan 28, 2026 ▶ 40:45
Insight
White: Simulations model boring things well but fail on interesting phenomena
“So I think this is one of the fundamental, I don't know, dichotomies of the world is that simulations stimulate really boring things really well. They don't simulate interesting things very well.”
Andrew White Jan 28, 2026 ▶ 42:52
Assertion Supported
White: DE Shaw Research Hardcoded Molecular Dynamics into Custom Silicon
“They built their own silicon. They built their own clusters. They had them taped out all themselves. They burned into the silicon the algorithms to run MD.”
Andrew White Jan 28, 2026 ▶ 43:39
Assertion Supported
White: A good protein folding model can be trained in ~10,000 GPU hours
“I think the numbers are now like 10,000 GPU hours you can train a good protein folding model. It's actually turned out to be barely an inconvenience.”
Andrew White Jan 28, 2026 ▶ 45:01
Assertion Supported
White: ML trained on experimental data beat first-principles simulations by a large margin
“Two very well-resourced groups. They both tried different ideas, and the machine learning on experimental data beat out first principles simulation by You know, a very large margin.”
Andrew White Jan 28, 2026 ▶ 45:21
Assertion Not yet assessed · timeframe Jan 2026
White: Synthesis routes for dangerous compounds are already available on Wikipedia
“You can go find the synthesis route for many dangerous compounds on Wikipedia. People know what are the targets in the human body that, like, are targeted by most biological weapons. It's not really that much of a mystery.”
Andrew White Jan 28, 2026 ▶ 47:05
Opinion
White: New AI laboratory risk scenarios have become realistic
“To some extent, there was like a first wave that we thought this could unlock a lot of stuff and I don't think it came to pass. I think there's now an emerging sort of second wave of like, there are some actually new scenarios that were just too farfetched to …”
Andrew White Jan 28, 2026 ▶ 48:57
Opinion
White: LLMs do not meaningfully accelerate physical CBRN and nuclear threats
“The classical example of nuclear is like, it's a lot of centrifugation, a lot of ultra centrifugation, a lot of high pressure or high RPMs. And so it, It's just, you can maybe get smarter about how to set up, you know, the economy of scale to do that with an L…”
Andrew White Jan 28, 2026 ▶ 49:53
Opinion
White: AI accelerates dual-use logistics like equipment sourcing and KYC
“All those, like, sort of simple logistical things, I think, are accelerated by AI, just as, like, a consequence of AI being an accelerating technology.”
Andrew White Jan 28, 2026 ▶ 51:11
Opinion
Andrew White: Funding $1M+ Cash Salaries With VC Equity Is Insane
“There are some venture backed companies that are having cash salaries over a million dollars. And it's like insane to me that you would use all of your cash from your equity financing You know, in these insane salaries.”
Andrew White Jan 28, 2026 ▶ 52:57
Opinion
Anderson: $1M+ salaries can make sense when GPU spend dominates burn
“In terms of like total spend on GPUs, it can still be a total, a small fraction of your burn. So sometimes it kind of makes sense.”
Brandon Anderson Jan 28, 2026 ▶ 53:14
Prediction Not checkable as stated
Andrew White: Automating scientific discovery will increase overall demand rather than displace jobs
“In science, I don't think there is a finite appetite or a finite capacity for science. I don't think science is like a scarcity thing. Like there's, you know, 100 more discoveries left to be made and then we'll be done. And so like we're displacing jobs. I thi…”
Andrew White Jan 28, 2026 ▶ 54:05
Prediction Not checkable as stated
White: Future scientists will act as AI agent wranglers exploring 100 ideas simultaneously
“My vision for what a scientist would be in the future is that They will be, I don't know, like agent wranglers or cosmos wranglers of like, okay, they're exploring a hundred ideas simultaneously, or they're like working with systems like ours to make an X the …”
Andrew White Jan 28, 2026 ▶ 54:30
Insight
White: Scientists will remain active producers because they are the primary consumers of science
“I think the enjoyers of science are also scientists. And so I think that it's kind of hard to imagine a scenario when there's not scientists as the consumers of science. And so I think if they're going to be consumers of science, they're also going to be some …”
Andrew White Jan 28, 2026 ▶ 55:56
Insight
White: Engineering Can Run Autonomously, but Science Requires Human Interpretation
“Engineering can exist by itself. Like if you give some kind of system, a goal of like making me a material that I can make a space elevator out of, you could be not participating in the beginning, the process of the middle, the process, and you just come by th…”
Andrew White Jan 28, 2026 ▶ 56:51
Insight
White: Natural language is the only way to connect all scientific data
“And so I think that that, for that reason, natural language is the only possible way to connect all the different pieces of data we need in biology, medicine, or any domain for that matter.”
Andrew White Jan 28, 2026 ▶ 1:01:52
Insight
White: Taking strong, imperfect opinions drives faster scientific progress than optionality
“I think in my career, it's actually been better for me to take strong opinions, which in my deepest of hearts, I know that are maybe not correct or not fully correct. But once you take these strong opinions, it just, You can sort of move many steps down the ro…”
Andrew White Jan 28, 2026 ▶ 1:07:17
Disclosure
White: FutureHouse bypassed custom foundation models to focus on scientific agents
“For example, at future house, we took the opinion that scientific agents are the future, and that allows us to skip a lot of steps because a lot of other people were like, we need to build a foundation model for X. Yeah. And we just skipped all that.”
Andrew White Jan 28, 2026 ▶ 1:07:32
Insight
White: Writing bulletproof RL verifiers is far harder than supervised training
“Pre-training or training transformers, you know on just data, like just supervised training where you just have the inputs and the outputs directly, very nice, relaxing, you know, like things are always robust, you know, things go pretty smoothly. When we do t…”
Andrew White Jan 28, 2026 ▶ 1:12:03
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.