Jessica Rumbelow of Leap Labs discusses the architectural limitations of LLMs when analyzing scientific data versus using mechanistic interpretability tools.
Assertion Supported
Rumbelow: Leap Labs' Discovery Engine Automates Novel Scientific Discovery via Interpretability
“Discovery Engine is an end-to-end system, takes in arbitrary scientific data set, automatically trains a bunch of neural networks on it, and then We systematically, with our interpretability methods, which is the real secret extract the patterns that have been…”
Assertion Not checkable as stated
Rumbelow: arXiv Is Filling With Plausible but Unverified AI Papers
“I'm actually really worried about this because I think we're already seeing archive and other online repositories and... Submissions too, full of these very, very plausible papers. That may or may not be true. And like at that point, what good is the, is our s…”
Insight
Rumbelow: Interpretability turns neural networks into scientific discovery tools
“If you've got really good interp, you can start to reframe neural networks, not as just a tool for automating things that we already know how to do, but as a tool for discovery, as like a lens through which you can see patterns in data that would otherwise
Be …”
Assertion Supported
Rumbelow: Leap Labs Discovered Novel Predictive Markers for Tumor-Reactive T Cells
“And we found basically novel markers that are quite predictive of this. And these novel markers were not what we expected them to be.”
Assertion Not checkable as stated
Rumbelow: Standalone Claude Opus Hallucinated Materials Science Data Findings
“So, so Claude, lovely Claude. I'm sorry Claude, but you did a terrible job. It hallucinated some stuff. It made some like big sweeping over, over generalizations. It like over indexed the few outliers. Like, it's fine. It's not Claude's fault. Like Claude is j…”
Insight
Rumbelow: Multimodal analysis is impossible without fine labeling or interpretability
“It's largely impossible to do good data analysis on multimodal data of this kind, unless you have really fine grained labeling of your images. For example, which is just very, very burdensome, but obviously deep learning, we can let the model figure out its ow…”