Nov 2, 2025 · 27m · latent-space

⚡️Automating Scientific Discovery - Jessica Rumbelow, Leap Labs

Jessica Rumbelow · 19m spoken Shawn Wang · 3m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Dr. Jessica Rumbelow of Leap Labs joins Swyx to discuss the Discovery Engine, demonstrating how mechanistic AI interpretability transforms scientific research from manual hypothesis testing into automated, verifiable pattern discovery.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 16.5% of the talking time here. How this is scored →

The hosts as informed peer 3.4 Guest teaching 4.6 Guest disagreement 1.4 The hosts pushing back 1.6
05100:0010:0020:000:03–5:44 · The hosts as informed peer 5/10 Jessica Rumbelow's Journey into AI Interpretability Swyx demonstrates domain awareness regarding mechanistic interpretability and steering vectors at Anthropic and Goodfire. Jessica gently pushes back on Swyx's theory about isolated reasoning vectors, arguing reasoning is likely suffused across network layers.5:45–12:15 · The hosts as informed peer 1/10 Discovery Engine and Real-World Scientific Case Studies Jessica leads the walkthrough of Discovery Engine case studies across plant biology, immunology, and meteorological surface layer theory. Swyx acts almost entirely as an engaged listener prompting transitions.12:15–17:56 · The hosts as informed peer 4/10 Benchmarking LLMs vs. Discovery Engine in Frontier Science Jessica details why standard LLMs fail at hypothesis-free discovery on complex datasets due to path dependence and hallucinations. Swyx brings up Latent Space framing on tool-augmented thinking and next-gen frontier model optimization.17:56–24:18 · The hosts as informed peer 3/10 Live Demonstration: Concrete Compressive Strength Dashboard During the live concrete strength dashboard demo, Swyx compares the output to exploratory data analysis (EDA), prompting Jessica to firmly clarify that it is systematic extraction rather than exploratory. She explains how the engine catches multi-feature combinatorial effects humans miss.24:18–27:17 · The hosts as informed peer 4/10 Roadmap, Multimodal Expansion, and Conclusion When Jessica claims a 100x acceleration over manual analysis, Swyx immediately presses for the factual basis and asks if that assumes simulation. Jessica clarifies the gain stems from bypassing human iterative hypothesis cycles.0:03–5:44 · Guest teaching 3/10 Jessica Rumbelow's Journey into AI Interpretability Swyx demonstrates domain awareness regarding mechanistic interpretability and steering vectors at Anthropic and Goodfire. Jessica gently pushes back on Swyx's theory about isolated reasoning vectors, arguing reasoning is likely suffused across network layers.5:45–12:15 · Guest teaching 6/10 Discovery Engine and Real-World Scientific Case Studies Jessica leads the walkthrough of Discovery Engine case studies across plant biology, immunology, and meteorological surface layer theory. Swyx acts almost entirely as an engaged listener prompting transitions.12:15–17:56 · Guest teaching 5/10 Benchmarking LLMs vs. Discovery Engine in Frontier Science Jessica details why standard LLMs fail at hypothesis-free discovery on complex datasets due to path dependence and hallucinations. Swyx brings up Latent Space framing on tool-augmented thinking and next-gen frontier model optimization.17:56–24:18 · Guest teaching 5/10 Live Demonstration: Concrete Compressive Strength Dashboard During the live concrete strength dashboard demo, Swyx compares the output to exploratory data analysis (EDA), prompting Jessica to firmly clarify that it is systematic extraction rather than exploratory. She explains how the engine catches multi-feature combinatorial effects humans miss.24:18–27:17 · Guest teaching 4/10 Roadmap, Multimodal Expansion, and Conclusion When Jessica claims a 100x acceleration over manual analysis, Swyx immediately presses for the factual basis and asks if that assumes simulation. Jessica clarifies the gain stems from bypassing human iterative hypothesis cycles.0:03–5:44 · Guest disagreement 2/10 Jessica Rumbelow's Journey into AI Interpretability Swyx demonstrates domain awareness regarding mechanistic interpretability and steering vectors at Anthropic and Goodfire. Jessica gently pushes back on Swyx's theory about isolated reasoning vectors, arguing reasoning is likely suffused across network layers.5:45–12:15 · Guest disagreement 0/10 Discovery Engine and Real-World Scientific Case Studies Jessica leads the walkthrough of Discovery Engine case studies across plant biology, immunology, and meteorological surface layer theory. Swyx acts almost entirely as an engaged listener prompting transitions.12:15–17:56 · Guest disagreement 1/10 Benchmarking LLMs vs. Discovery Engine in Frontier Science Jessica details why standard LLMs fail at hypothesis-free discovery on complex datasets due to path dependence and hallucinations. Swyx brings up Latent Space framing on tool-augmented thinking and next-gen frontier model optimization.17:56–24:18 · Guest disagreement 2/10 Live Demonstration: Concrete Compressive Strength Dashboard During the live concrete strength dashboard demo, Swyx compares the output to exploratory data analysis (EDA), prompting Jessica to firmly clarify that it is systematic extraction rather than exploratory. She explains how the engine catches multi-feature combinatorial effects humans miss.24:18–27:17 · Guest disagreement 2/10 Roadmap, Multimodal Expansion, and Conclusion When Jessica claims a 100x acceleration over manual analysis, Swyx immediately presses for the factual basis and asks if that assumes simulation. Jessica clarifies the gain stems from bypassing human iterative hypothesis cycles.0:03–5:44 · The hosts pushing back 2/10 Jessica Rumbelow's Journey into AI Interpretability Swyx demonstrates domain awareness regarding mechanistic interpretability and steering vectors at Anthropic and Goodfire. Jessica gently pushes back on Swyx's theory about isolated reasoning vectors, arguing reasoning is likely suffused across network layers.5:45–12:15 · The hosts pushing back 0/10 Discovery Engine and Real-World Scientific Case Studies Jessica leads the walkthrough of Discovery Engine case studies across plant biology, immunology, and meteorological surface layer theory. Swyx acts almost entirely as an engaged listener prompting transitions.12:15–17:56 · The hosts pushing back 1/10 Benchmarking LLMs vs. Discovery Engine in Frontier Science Jessica details why standard LLMs fail at hypothesis-free discovery on complex datasets due to path dependence and hallucinations. Swyx brings up Latent Space framing on tool-augmented thinking and next-gen frontier model optimization.17:56–24:18 · The hosts pushing back 1/10 Live Demonstration: Concrete Compressive Strength Dashboard During the live concrete strength dashboard demo, Swyx compares the output to exploratory data analysis (EDA), prompting Jessica to firmly clarify that it is systematic extraction rather than exploratory. She explains how the engine catches multi-feature combinatorial effects humans miss.24:18–27:17 · The hosts pushing back 4/10 Roadmap, Multimodal Expansion, and Conclusion When Jessica claims a 100x acceleration over manual analysis, Swyx immediately presses for the factual basis and asks if that assumes simulation. Jessica clarifies the gain stems from bypassing human iterative hypothesis cycles.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 21% · guest 79%0:00 · the hosts 21% · guest 79%3:00 · the hosts 36.1% · guest 63.9%3:00 · the hosts 36.1% · guest 63.9%6:00 · the hosts 5% · guest 95%6:00 · the hosts 5% · guest 95%9:00 · the hosts 0.9% · guest 99.1%9:00 · the hosts 0.9% · guest 99.1%12:00 · the hosts 15.9% · guest 84.1%12:00 · the hosts 15.9% · guest 84.1%15:00 · the hosts 20% · guest 80%15:00 · the hosts 20% · guest 80%18:00 · the hosts 10.2% · guest 89.8%18:00 · the hosts 10.2% · guest 89.8%21:00 · the hosts 4.2% · guest 95.8%21:00 · the hosts 4.2% · guest 95.8%24:00 · the hosts 29.3% · guest 70.7%24:00 · the hosts 29.3% · guest 70.7%27:00 · the hosts 90.3% · guest 9.7%27:00 · the hosts 90.3% · guest 9.7%
Sharpest disagreement ▶ 21:19 Correcting the EDA characterization

Jessica directly rejects Swyx's framing of the dashboard as exploratory data analysis, firmly asserting it is exhaustive and systematic.

Hardest push from the hosts ▶ 25:38 Challenging the 100x speedup claim

Swyx halts the pitch to challenge where the 100x number comes from, asking whether it relies on synthetic simulations.

Biggest teaching moment ▶ 22:30 Combinatorial synergy in concrete strength

Jessica breaks down how individual linear correlations miss high-order combinations of conditions that dramatically boost concrete strength.

The host holds their own ▶ 4:14 Mechanistic interpretability and reasoning activations

Swyx articulates a technical theory on how decomposing and steering reasoning vectors could allow AI labs to leapfrog competitors.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Jessica Rumbelow's Journey into AI Interpretability 5322 Swyx demonstrates domain awareness regarding mechanistic interpretability and steering vectors at Anthropic and Goodfire. Jessica gently pushes back on Swyx's theory about isolated reasoning vectors, arguing reasoning is likely suffused across network layers.
Discovery Engine and Real-World Scientific Case Studies 1600 Jessica leads the walkthrough of Discovery Engine case studies across plant biology, immunology, and meteorological surface layer theory. Swyx acts almost entirely as an engaged listener prompting transitions.
Benchmarking LLMs vs. Discovery Engine in Frontier Science 4511 Jessica details why standard LLMs fail at hypothesis-free discovery on complex datasets due to path dependence and hallucinations. Swyx brings up Latent Space framing on tool-augmented thinking and next-gen frontier model optimization.
Live Demonstration: Concrete Compressive Strength Dashboard 3521 During the live concrete strength dashboard demo, Swyx compares the output to exploratory data analysis (EDA), prompting Jessica to firmly clarify that it is systematic extraction rather than exploratory. She explains how the engine catches multi-feature combinatorial effects humans miss.
Roadmap, Multimodal Expansion, and Conclusion 4424 When Jessica claims a 100x acceleration over manual analysis, Swyx immediately presses for the factual basis and asks if that assumes simulation. Jessica clarifies the gain stems from bypassing human iterative hypothesis cycles.

Statements from this episode (11)

Insight
Rumbelow: Interpretability turns neural networks into scientific discovery tools
“If you've got really good interp, you can start to reframe neural networks, not as just a tool for automating things that we already know how to do, but as a tool for discovery, as like a lens through which you can see patterns in data that would otherwise Be …”
Jessica Rumbelow Nov 2, 2025 ▶ 2:29
Prediction Not checkable as stated
Swyx: Tuning reasoning activations could let Anthropic leapfrog OpenAI
“If Anthropic ever found The activations for reasoning and could break down the different kinds of reasoning and turn, tune them properly. I think that's the thing that takes Anthropic to leapfrog OpenAI.”
Shawn Wang Nov 2, 2025 ▶ 4:33
Assertion Supported
Rumbelow: Leap Labs' Discovery Engine Automates Novel Scientific Discovery via Interpretability
“Discovery Engine is an end-to-end system, takes in arbitrary scientific data set, automatically trains a bunch of neural networks on it, and then We systematically, with our interpretability methods, which is the real secret extract the patterns that have been…”
Jessica Rumbelow Nov 2, 2025 ▶ 6:16
Assertion Supported
Rumbelow: Leap Labs Discovered Novel Predictive Markers for Tumor-Reactive T Cells
“And we found basically novel markers that are quite predictive of this. And these novel markers were not what we expected them to be.”
Jessica Rumbelow Nov 2, 2025 ▶ 8:12
Assertion Supported
Rumbelow: Leap Labs Found Soil Manganese and Genotype Synergy Impacting Root Architecture
“We found a bunch of new things actually. This is one of them, which is a particular like synergistic effect between manganese content in the soil and a particular genotype. Of this test crop, which actually has a really profound effect on the root architecture…”
Jessica Rumbelow Nov 2, 2025 ▶ 9:57
Insight
Rumbelow: LLMs Are Ill-Suited for Large Numeric Datasets
“Language models are just that, right? They're models of language. They are not particularly well suited for understanding arbitrary, you know, like big numeric data sets.”
Jessica Rumbelow Nov 2, 2025 ▶ 12:40
Assertion Not checkable as stated
Rumbelow: Standalone Claude Opus Hallucinated Materials Science Data Findings
“So, so Claude, lovely Claude. I'm sorry Claude, but you did a terrible job. It hallucinated some stuff. It made some like big sweeping over, over generalizations. It like over indexed the few outliers. Like, it's fine. It's not Claude's fault. Like Claude is j…”
Jessica Rumbelow Nov 2, 2025 ▶ 15:04
Assertion Not checkable as stated
Rumbelow: arXiv Is Filling With Plausible but Unverified AI Papers
“I'm actually really worried about this because I think we're already seeing archive and other online repositories and... Submissions too, full of these very, very plausible papers. That may or may not be true. And like at that point, what good is the, is our s…”
Jessica Rumbelow Nov 2, 2025 ▶ 16:21
Disclosure
Rumbelow: Leap Labs is launching free platform for open-source data
“We're launching a basically free. Dashboard, service, web, app, platform, free for academics, or free for anybody who will publish their data, basically free for open source data.”
Jessica Rumbelow Nov 2, 2025 ▶ 18:19
Insight
Rumbelow: Multimodal analysis is impossible without fine labeling or interpretability
“It's largely impossible to do good data analysis on multimodal data of this kind, unless you have really fine grained labeling of your images. For example, which is just very, very burdensome, but obviously deep learning, we can let the model figure out its ow…”
Jessica Rumbelow Nov 2, 2025 ▶ 25:04
Assertion Not checkable as stated
Rumbelow: Leap Labs is roughly 100x faster than manual analysis
“We're about a hundred times faster than manual analysis.”
Jessica Rumbelow Nov 2, 2025 ▶ 25:36
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.