Periodic Labs co-founder Doge Chubuk discusses the value of physical wet-lab data versus LLM-generated synthetic data with host Shail Khan.
0:00 / 0:47exact quote · 47.0s
720p mp4 · rendered on demand · StarZero watermark
“What's interesting about scientific data is it's not just a few bits or numbers, right? Like, for example, there are certain experiments you can run where the result you get from it is just, say, three floating point numbers. But the implications of those could be tremendous, right? It's not just going to be, like, a few bytes. It will actually be potentially an incredible amount of understanding just from a few experiments. And this has been how it is in human history, right? Like, there are certain experiments that told us so much about how we understand about the universe. And the way to do this with synthetic data can, of course, be, you know, you run simulations that relate to that experiment. And when you get the experimental result, that actually validates or Refuse so much of the simulations you ran, and then that is a lot of information in itself.”
quote is from the automated transcript, cleaned for reading:
filler sounds and stutters are removed, nothing is rephrased. names can be misheard
(the analysis reads context, assessments check outside sources). how →
More from Doge Chubuk
Insight
Chubuk: AI cannot reason to breakthrough superconductors from training data alone
“I think it's still true that it would be difficult to just reason your way into a much better superconductor. I actually would guess that there's a law out there that we haven't discovered yet that says that you can't just look at your training set that's diff…”
Doge ChubukNov 6, 2025▶ 12:56Inside a $300 million bet on AI for physical R&D
Insight
Chubuk: LLMs already bridge solid-state chemistry and physics better than human specialists
“Like, there was probably a time when a physicist could contribute and be one of the best in the world on many fields of physics, but it's definitely not true today, and this is one of the reasons I think we are very excited about LLMs, because when you talk to…”
Doge ChubukNov 6, 2025▶ 26:37Inside a $300 million bet on AI for physical R&D
“So what O-one showed is if you spend test time compute, you can get better results. So that was very exciting to me because there was one way of investing resources that was beyond the training set.”
Doge ChubukNov 6, 2025▶ 6:30Inside a $300 million bet on AI for physical R&D
AssertionNot checkable as stated
Chubuk: AI is currently not better than humans at hypothesis generation
“It does seem like today there are things that ML, AI is better than humans, but one of those things is not hypothesis generation.”
Doge ChubukNov 6, 2025▶ 22:54Inside a $300 million bet on AI for physical R&D
Disclosure
Chubuk: GPU compute and training costs drove Periodic Labs' $300M seed
“We are going to train LLMs, we are going to use GPUs to run simulations, so that does end up being a large part of the cost. Yeah, it's funny, like, before, you know, if you asked me this question 10 years ago, I would have thought that the biggest part of the…”
Doge ChubukNov 6, 2025▶ 24:05Inside a $300 million bet on AI for physical R&D
Insight
Chubuk: Science requires out-of-domain generalization unlike standard ML
“Machine learning works best on the training set distribution. But in science and technology, we almost only care about auto-domain generalization, right?”
Doge ChubukNov 6, 2025▶ 6:14Inside a $300 million bet on AI for physical R&D
Made with StarZero
Turn any episode into a week of clips.
This entire site, over 200 episodes transcribed, diarized, checked and made playable,
runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the
moments worth sharing, cuts them, captions them, and reframes them for every feed.
We use essential cookies to make the site work. With your permission we
also use analytics cookies (Google Analytics and Mixpanel) to understand
usage and improve StarZero. See our Cookie Policy.