Sep 30, 2025 · 51m · a16z

Building an AI Physicist: ChatGPT Co-Creator’s Next Venture

Liam Fedus · 20m spoken Doge Cubuk · 16m spoken Anjney Midha · 9m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of the a16z Podcast, Periodic Labs co-founders Liam Fedus and Doge Cubuk discuss their mission to build an AI physicist by combining large language models with automated physical lab experiments. They explain how physically grounded feedback, high-temperature superconductivity targets, and interdisciplinary collaboration can accelerate scientific discovery and material design.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. How this is scored →

The host as informed peer 2.8 Guest teaching 3.3 Guest disagreement 0.5 The host pushing back 0.6
05100:0015:0030:0045:000:48–3:53 · The host as informed peer 1/10 Origin Story: Flipping Tires at Google Brain The host opens with light podcast origin questions about how the co-founders met and transitioned from flipping tires to LLMs in physics. The guests share an anecdote and explain their early discussions on physics and AI.3:53–8:40 · The host as informed peer 2/10 Defining Periodic Labs and Physical Reward Functions The host prompts the guests to define Periodic Labs and explain physical verification versus standard AI training. Liam elaborates on replacing digital math/code reward functions with physically grounded reward functions in real-world labs.8:40–12:36 · The host as informed peer 2/10 Scientific Inquiry and the Need for Physical Iteration The host asks why existing deployed models cannot perform physical discovery. Doge and Liam explain the necessity of physical iteration over pure logic, highlighting epistemic uncertainty and missing negative results in literature.12:36–16:35 · The host as informed peer 2/10 Measuring Progress Through Material Discovery The host asks for specific progress metrics for Periodic Labs. Doge and Liam set concrete benchmarks such as discovering superconductors beyond 135 Kelvin and direct property measurements.16:35–22:23 · The host as informed peer 6/10 Out-of-Domain Generalization and Data Bottlenecks The host demonstrates strong technical knowledge by citing GPT-3 and scaling laws papers to question why pure compute and data scaling wouldn't automatically solve physics out-of-domain. Liam and Doge counter by explaining out-of-domain power law slopes and data bottlenecks.22:23–26:08 · The host as informed peer 5/10 High-Temperature Superconductivity as a Unifying Mission The host invokes Sutton's 'bitter lesson' concept and questions if focusing on specific domain pipelines like superconductivity creates off-ramps from true AGI. Doge explains superconductivity as a strategic North Star rich in sub-goals.26:08–28:49 · The host as informed peer 4/10 Unlocking R&D Value in Advanced Industries The host draws a structural analogy between white-collar software copilots and physical R&D copilots. Liam confirms that commercial copilots for advanced physical industries represent their intermediate business model.28:49–32:48 · The host as informed peer 3/10 Fostering Synergy Between ML and Physical Scientists The host inquires about organizational design and uniting machine learning scientists with physical scientists. The guests discuss cross-teaching sessions and mapping physical concepts into machine learning APIs.32:48–35:34 · The host as informed peer 1/10 Mission-Driven Culture and Non-Traditional Backgrounds The host asks about hiring criteria and whether advanced physics degrees are required. Doge uses a humorous LeBron James analogy to explain the vastness of scientific knowledge and the necessity of interdisciplinary collaboration.35:34–38:23 · The host as informed peer 2/10 Land-and-Expand Strategy for Industrial Adoption The host asks about deployment strategies into conservative industries like space and defense. Liam details a land-and-expand approach focused on solving well-scoped, critical evaluation problems.38:23–40:57 · The host as informed peer 2/10 Replacing Basic Retrieval with Deep Model Weights The host asks about urgent customer problems from recent calls. Liam contrasts basic retrieval-augmented generation (RAG) with deep model weight pre-training and high-compute reinforcement learning.40:57–43:28 · The host as informed peer 2/10 Injecting Knowledge via Scientific Mid-Training The host asks for a definition of 'mid-training' for the audience. Liam breaks down how mid-training continually injects domain knowledge into model weights between standard pre-training and post-training.43:28–49:47 · The host as informed peer 4/10 Modular Integration of Base LLMs and Physics Tools The host cites personal experience evaluating models at the Stanford physics lab to highlight model deficiencies. Liam playfully retorts that models failed because they weren't trained for physics, leading into a discussion on modular tools and academic partnerships.0:48–3:53 · Guest teaching 2/10 Origin Story: Flipping Tires at Google Brain The host opens with light podcast origin questions about how the co-founders met and transitioned from flipping tires to LLMs in physics. The guests share an anecdote and explain their early discussions on physics and AI.3:53–8:40 · Guest teaching 4/10 Defining Periodic Labs and Physical Reward Functions The host prompts the guests to define Periodic Labs and explain physical verification versus standard AI training. Liam elaborates on replacing digital math/code reward functions with physically grounded reward functions in real-world labs.8:40–12:36 · Guest teaching 4/10 Scientific Inquiry and the Need for Physical Iteration The host asks why existing deployed models cannot perform physical discovery. Doge and Liam explain the necessity of physical iteration over pure logic, highlighting epistemic uncertainty and missing negative results in literature.12:36–16:35 · Guest teaching 3/10 Measuring Progress Through Material Discovery The host asks for specific progress metrics for Periodic Labs. Doge and Liam set concrete benchmarks such as discovering superconductors beyond 135 Kelvin and direct property measurements.16:35–22:23 · Guest teaching 5/10 Out-of-Domain Generalization and Data Bottlenecks The host demonstrates strong technical knowledge by citing GPT-3 and scaling laws papers to question why pure compute and data scaling wouldn't automatically solve physics out-of-domain. Liam and Doge counter by explaining out-of-domain power law slopes and data bottlenecks.22:23–26:08 · Guest teaching 4/10 High-Temperature Superconductivity as a Unifying Mission The host invokes Sutton's 'bitter lesson' concept and questions if focusing on specific domain pipelines like superconductivity creates off-ramps from true AGI. Doge explains superconductivity as a strategic North Star rich in sub-goals.26:08–28:49 · Guest teaching 3/10 Unlocking R&D Value in Advanced Industries The host draws a structural analogy between white-collar software copilots and physical R&D copilots. Liam confirms that commercial copilots for advanced physical industries represent their intermediate business model.28:49–32:48 · Guest teaching 3/10 Fostering Synergy Between ML and Physical Scientists The host inquires about organizational design and uniting machine learning scientists with physical scientists. The guests discuss cross-teaching sessions and mapping physical concepts into machine learning APIs.32:48–35:34 · Guest teaching 2/10 Mission-Driven Culture and Non-Traditional Backgrounds The host asks about hiring criteria and whether advanced physics degrees are required. Doge uses a humorous LeBron James analogy to explain the vastness of scientific knowledge and the necessity of interdisciplinary collaboration.35:34–38:23 · Guest teaching 2/10 Land-and-Expand Strategy for Industrial Adoption The host asks about deployment strategies into conservative industries like space and defense. Liam details a land-and-expand approach focused on solving well-scoped, critical evaluation problems.38:23–40:57 · Guest teaching 4/10 Replacing Basic Retrieval with Deep Model Weights The host asks about urgent customer problems from recent calls. Liam contrasts basic retrieval-augmented generation (RAG) with deep model weight pre-training and high-compute reinforcement learning.40:57–43:28 · Guest teaching 3/10 Injecting Knowledge via Scientific Mid-Training The host asks for a definition of 'mid-training' for the audience. Liam breaks down how mid-training continually injects domain knowledge into model weights between standard pre-training and post-training.43:28–49:47 · Guest teaching 4/10 Modular Integration of Base LLMs and Physics Tools The host cites personal experience evaluating models at the Stanford physics lab to highlight model deficiencies. Liam playfully retorts that models failed because they weren't trained for physics, leading into a discussion on modular tools and academic partnerships.0:48–3:53 · Guest disagreement 0/10 Origin Story: Flipping Tires at Google Brain The host opens with light podcast origin questions about how the co-founders met and transitioned from flipping tires to LLMs in physics. The guests share an anecdote and explain their early discussions on physics and AI.3:53–8:40 · Guest disagreement 0/10 Defining Periodic Labs and Physical Reward Functions The host prompts the guests to define Periodic Labs and explain physical verification versus standard AI training. Liam elaborates on replacing digital math/code reward functions with physically grounded reward functions in real-world labs.8:40–12:36 · Guest disagreement 1/10 Scientific Inquiry and the Need for Physical Iteration The host asks why existing deployed models cannot perform physical discovery. Doge and Liam explain the necessity of physical iteration over pure logic, highlighting epistemic uncertainty and missing negative results in literature.12:36–16:35 · Guest disagreement 0/10 Measuring Progress Through Material Discovery The host asks for specific progress metrics for Periodic Labs. Doge and Liam set concrete benchmarks such as discovering superconductors beyond 135 Kelvin and direct property measurements.16:35–22:23 · Guest disagreement 2/10 Out-of-Domain Generalization and Data Bottlenecks The host demonstrates strong technical knowledge by citing GPT-3 and scaling laws papers to question why pure compute and data scaling wouldn't automatically solve physics out-of-domain. Liam and Doge counter by explaining out-of-domain power law slopes and data bottlenecks.22:23–26:08 · Guest disagreement 1/10 High-Temperature Superconductivity as a Unifying Mission The host invokes Sutton's 'bitter lesson' concept and questions if focusing on specific domain pipelines like superconductivity creates off-ramps from true AGI. Doge explains superconductivity as a strategic North Star rich in sub-goals.26:08–28:49 · Guest disagreement 0/10 Unlocking R&D Value in Advanced Industries The host draws a structural analogy between white-collar software copilots and physical R&D copilots. Liam confirms that commercial copilots for advanced physical industries represent their intermediate business model.28:49–32:48 · Guest disagreement 0/10 Fostering Synergy Between ML and Physical Scientists The host inquires about organizational design and uniting machine learning scientists with physical scientists. The guests discuss cross-teaching sessions and mapping physical concepts into machine learning APIs.32:48–35:34 · Guest disagreement 0/10 Mission-Driven Culture and Non-Traditional Backgrounds The host asks about hiring criteria and whether advanced physics degrees are required. Doge uses a humorous LeBron James analogy to explain the vastness of scientific knowledge and the necessity of interdisciplinary collaboration.35:34–38:23 · Guest disagreement 0/10 Land-and-Expand Strategy for Industrial Adoption The host asks about deployment strategies into conservative industries like space and defense. Liam details a land-and-expand approach focused on solving well-scoped, critical evaluation problems.38:23–40:57 · Guest disagreement 1/10 Replacing Basic Retrieval with Deep Model Weights The host asks about urgent customer problems from recent calls. Liam contrasts basic retrieval-augmented generation (RAG) with deep model weight pre-training and high-compute reinforcement learning.40:57–43:28 · Guest disagreement 0/10 Injecting Knowledge via Scientific Mid-Training The host asks for a definition of 'mid-training' for the audience. Liam breaks down how mid-training continually injects domain knowledge into model weights between standard pre-training and post-training.43:28–49:47 · Guest disagreement 2/10 Modular Integration of Base LLMs and Physics Tools The host cites personal experience evaluating models at the Stanford physics lab to highlight model deficiencies. Liam playfully retorts that models failed because they weren't trained for physics, leading into a discussion on modular tools and academic partnerships.0:48–3:53 · The host pushing back 0/10 Origin Story: Flipping Tires at Google Brain The host opens with light podcast origin questions about how the co-founders met and transitioned from flipping tires to LLMs in physics. The guests share an anecdote and explain their early discussions on physics and AI.3:53–8:40 · The host pushing back 0/10 Defining Periodic Labs and Physical Reward Functions The host prompts the guests to define Periodic Labs and explain physical verification versus standard AI training. Liam elaborates on replacing digital math/code reward functions with physically grounded reward functions in real-world labs.8:40–12:36 · The host pushing back 0/10 Scientific Inquiry and the Need for Physical Iteration The host asks why existing deployed models cannot perform physical discovery. Doge and Liam explain the necessity of physical iteration over pure logic, highlighting epistemic uncertainty and missing negative results in literature.12:36–16:35 · The host pushing back 0/10 Measuring Progress Through Material Discovery The host asks for specific progress metrics for Periodic Labs. Doge and Liam set concrete benchmarks such as discovering superconductors beyond 135 Kelvin and direct property measurements.16:35–22:23 · The host pushing back 3/10 Out-of-Domain Generalization and Data Bottlenecks The host demonstrates strong technical knowledge by citing GPT-3 and scaling laws papers to question why pure compute and data scaling wouldn't automatically solve physics out-of-domain. Liam and Doge counter by explaining out-of-domain power law slopes and data bottlenecks.22:23–26:08 · The host pushing back 2/10 High-Temperature Superconductivity as a Unifying Mission The host invokes Sutton's 'bitter lesson' concept and questions if focusing on specific domain pipelines like superconductivity creates off-ramps from true AGI. Doge explains superconductivity as a strategic North Star rich in sub-goals.26:08–28:49 · The host pushing back 1/10 Unlocking R&D Value in Advanced Industries The host draws a structural analogy between white-collar software copilots and physical R&D copilots. Liam confirms that commercial copilots for advanced physical industries represent their intermediate business model.28:49–32:48 · The host pushing back 0/10 Fostering Synergy Between ML and Physical Scientists The host inquires about organizational design and uniting machine learning scientists with physical scientists. The guests discuss cross-teaching sessions and mapping physical concepts into machine learning APIs.32:48–35:34 · The host pushing back 0/10 Mission-Driven Culture and Non-Traditional Backgrounds The host asks about hiring criteria and whether advanced physics degrees are required. Doge uses a humorous LeBron James analogy to explain the vastness of scientific knowledge and the necessity of interdisciplinary collaboration.35:34–38:23 · The host pushing back 0/10 Land-and-Expand Strategy for Industrial Adoption The host asks about deployment strategies into conservative industries like space and defense. Liam details a land-and-expand approach focused on solving well-scoped, critical evaluation problems.38:23–40:57 · The host pushing back 0/10 Replacing Basic Retrieval with Deep Model Weights The host asks about urgent customer problems from recent calls. Liam contrasts basic retrieval-augmented generation (RAG) with deep model weight pre-training and high-compute reinforcement learning.40:57–43:28 · The host pushing back 0/10 Injecting Knowledge via Scientific Mid-Training The host asks for a definition of 'mid-training' for the audience. Liam breaks down how mid-training continually injects domain knowledge into model weights between standard pre-training and post-training.43:28–49:47 · The host pushing back 2/10 Modular Integration of Base LLMs and Physics Tools The host cites personal experience evaluating models at the Stanford physics lab to highlight model deficiencies. Liam playfully retorts that models failed because they weren't trained for physics, leading into a discussion on modular tools and academic partnerships.

speaking balance: gold is the host, purple is the guest (3 minute bins)

0:00 · the host 0% · guest 100%0:00 · the host 0% · guest 100%3:00 · the host 0% · guest 100%3:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%6:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%9:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%12:00 · the host 0% · guest 100%15:00 · the host 0% · guest 100%15:00 · the host 0% · guest 100%18:00 · the host 0% · guest 100%18:00 · the host 0% · guest 100%21:00 · the host 0% · guest 100%21:00 · the host 0% · guest 100%24:00 · the host 0% · guest 100%24:00 · the host 0% · guest 100%27:00 · the host 0% · guest 100%27:00 · the host 0% · guest 100%30:00 · the host 0% · guest 100%30:00 · the host 0% · guest 100%33:00 · the host 0% · guest 100%33:00 · the host 0% · guest 100%36:00 · the host 0% · guest 100%36:00 · the host 0% · guest 100%39:00 · the host 0% · guest 100%39:00 · the host 0% · guest 100%42:00 · the host 0% · guest 100%42:00 · the host 0% · guest 100%45:00 · the host 0% · guest 100%45:00 · the host 0% · guest 100%48:00 · the host 0% · guest 100%48:00 · the host 0% · guest 100%51:00 · the host 0% · guest 100%51:00 · the host 0% · guest 100%
Sharpest disagreement ▶ 18:25 Liam reframing model limits on physical domain tasks

Liam directly rejects the premise that standard coding or general LLM scaling will naturally allow a model to solve out-of-domain physical tasks like curing cancer without targeted environment optimization.

Hardest push from the host ▶ 16:36 Host challenging physical lab necessity via scaling literature

The host actively pushes back against the premise that specialized physical verification labs are necessary, using classic scaling law research to argue that general compute scaling might crack physics automatically.

Biggest teaching moment ▶ 19:47 Doge breaking down out-of-domain power law slopes

Doge educates the host on the mathematical nuances of scaling laws, explaining how out-of-domain power law slopes can be so flat that scaling compute without dataset adjustment requires centuries of computation.

The host holds their own ▶ 16:36 Host citing GPT-3 and scaling law papers

The host demonstrates deep domain fluency by referencing specific seminal papers on few-shot learning and generative model scaling to construct a rigorous counter-argument.

the scores for every segment, with the reasoning behind each
ChapterTopicThe host as informed peerGuest teachingGuest disagreementThe host pushing backWhy
Origin Story: Flipping Tires at Google Brain 1200 The host opens with light podcast origin questions about how the co-founders met and transitioned from flipping tires to LLMs in physics. The guests share an anecdote and explain their early discussions on physics and AI.
Defining Periodic Labs and Physical Reward Functions 2400 The host prompts the guests to define Periodic Labs and explain physical verification versus standard AI training. Liam elaborates on replacing digital math/code reward functions with physically grounded reward functions in real-world labs.
Scientific Inquiry and the Need for Physical Iteration 2410 The host asks why existing deployed models cannot perform physical discovery. Doge and Liam explain the necessity of physical iteration over pure logic, highlighting epistemic uncertainty and missing negative results in literature.
Measuring Progress Through Material Discovery 2300 The host asks for specific progress metrics for Periodic Labs. Doge and Liam set concrete benchmarks such as discovering superconductors beyond 135 Kelvin and direct property measurements.
Out-of-Domain Generalization and Data Bottlenecks 6523 The host demonstrates strong technical knowledge by citing GPT-3 and scaling laws papers to question why pure compute and data scaling wouldn't automatically solve physics out-of-domain. Liam and Doge counter by explaining out-of-domain power law slopes and data bottlenecks.
High-Temperature Superconductivity as a Unifying Mission 5412 The host invokes Sutton's 'bitter lesson' concept and questions if focusing on specific domain pipelines like superconductivity creates off-ramps from true AGI. Doge explains superconductivity as a strategic North Star rich in sub-goals.
Unlocking R&D Value in Advanced Industries 4301 The host draws a structural analogy between white-collar software copilots and physical R&D copilots. Liam confirms that commercial copilots for advanced physical industries represent their intermediate business model.
Fostering Synergy Between ML and Physical Scientists 3300 The host inquires about organizational design and uniting machine learning scientists with physical scientists. The guests discuss cross-teaching sessions and mapping physical concepts into machine learning APIs.
Mission-Driven Culture and Non-Traditional Backgrounds 1200 The host asks about hiring criteria and whether advanced physics degrees are required. Doge uses a humorous LeBron James analogy to explain the vastness of scientific knowledge and the necessity of interdisciplinary collaboration.
Land-and-Expand Strategy for Industrial Adoption 2200 The host asks about deployment strategies into conservative industries like space and defense. Liam details a land-and-expand approach focused on solving well-scoped, critical evaluation problems.
Replacing Basic Retrieval with Deep Model Weights 2410 The host asks about urgent customer problems from recent calls. Liam contrasts basic retrieval-augmented generation (RAG) with deep model weight pre-training and high-compute reinforcement learning.
Injecting Knowledge via Scientific Mid-Training 2300 The host asks for a definition of 'mid-training' for the audience. Liam breaks down how mid-training continually injects domain knowledge into model weights between standard pre-training and post-training.
Modular Integration of Base LLMs and Physics Tools 4422 The host cites personal experience evaluating models at the Stanford physics lab to highlight model deficiencies. Liam playfully retorts that models failed because they weren't trained for physics, leading into a discussion on modular tools and academic partnerships.

Statements from this episode (21)

Opinion
Cubuk: A 200K superconductor would reshape fundamental quantum physics
“For example, if we could find a 200 Kelvin superconductor, even before we make any product with it, to be able to see such quantum effects at such high temperatures, I think would be such an update to people's view of how they see the universe.”
Doge Cubuk Sep 30, 2025 ▶ 0:33
Assertion Not checkable as stated
Fedus: Physics and chemistry demonstrate scaling laws similar to AI
“On the material science side, we're seeing scaling laws within physics, within chemistry both with respect to simulations, with respect to experiment, and it's like the same kind of principles at play and ML.”
Liam Fedus Sep 30, 2025 ▶ 3:02
Insight
Fedus: Physics provides ideal verifiable reward functions for AI
“Physics is very verifiable. It's a great reward function, fairly fast iteration loop. You have simulators for large classes of physical systems.”
Liam Fedus Sep 30, 2025 ▶ 3:35
Disclosure
Fedus: Periodic Labs uses physical experiments as RL reward functions
“And what we're doing, and by having the lab, is we create a physically grounded reward function. That becomes the basis on which we're optimizing against. And so, If a simulator has some deficiencies or some issues, we always error correct, because for us, the…”
Liam Fedus Sep 30, 2025 ▶ 5:01
Disclosure
Fedus: Early ChatGPT was mathematically weak due to friendliness rewards
“The reward functions that we were using originally couldn't determine whether you were mathematically correct or not. So early versions of Chachapiti were mathematically not particularly strong, and it sort of results from the reward function. What did you opt…”
Liam Fedus Sep 30, 2025 ▶ 7:02
Insight
Fedus: AI science requires real-world experimental feedback loops
“Ultimately science is driven against experiment in the real world. And so that's what we're doing with periodic labs. We're taking these precursor technologies and we're saying, okay, if you care about advancing science, we need to have experiment in the loop.”
Liam Fedus Sep 30, 2025 ▶ 7:46
Disclosure
Cubuk: Periodic Labs is building an automated powder synthesis lab
“We're going to have a powder synthesis lab, and turns out this is one of those methods where robots can do it, like very cheap, simple methods.”
Doge Cubuk Sep 30, 2025 ▶ 9:36
Prediction Not checkable as stated
Cubuk: Quantum mechanics foundation models are AI's next frontier
“And we feel like teaching these LLMs to be foundation models, but for quantum mechanics will be the next frontier for LLMs.”
Doge Cubuk Sep 30, 2025 ▶ 10:05
Insight
Fedus: Unpublished negative scientific results are uniquely valuable for AI
“Then another point is, it's very uncommon to publish negative results. All of the results are basically positive, and a valid negative result is very valuable.”
Liam Fedus Sep 30, 2025 ▶ 12:03
Assertion Supported
Cubuk: Ambient-pressure superconductivity record stands at roughly 135 Kelvin
“Today the best number for ambient pressure is 135 Kelvin or so”
Doge Cubuk Sep 30, 2025 ▶ 12:51
Insight
Cubuk: Physical experiments provide unhackable training signals for LLMs
“As we measure it, the LLM will get very clear signal. It's hard to hack, you know, unless, unlike these other LLM training techniques, it's like really what you see in real life is the signal that's going to the LLM.”
Doge Cubuk Sep 30, 2025 ▶ 13:10
Insight
Fedus: AI physics requires generating new experimental data, not web scrapes
“The technology that we think is necessary to do it has really just emerged in the last couple of years, and this data Isn't like on a Reddit forum or something like you need to actually go produce experimental data, simulation data.”
Liam Fedus Sep 30, 2025 ▶ 16:08
Assertion Supported
Cubuk: Out-of-domain AI scaling power-law slopes can be practically useless
“We published a paper where we saw that as you increase the size of your training set, the IID performance, the in-domain performance improves the power law. Out-of-domain performance also improves the power law, but depending on what the outdoor domain is, lik…”
Doge Cubuk Sep 30, 2025 ▶ 20:29
Assertion Not checkable as stated
Cubuk: Existing superconductivity datasets are too noisy for AI training
“For superconductivity, there is a lot of data sets you can look at, but the noise floor on them is so high that training on them usually doesn't help.”
Doge Cubuk Sep 30, 2025 ▶ 22:00
Prediction Open · timeframe Sep 2028
Fedus: Periodic Labs will build AI co-pilots for space and defense
“Basically co-pilots for engineers, researchers in advanced industries. So maybe perhaps just being in Silicon Valley, we, you know, we really think about like computer oriented work. Everything is digital. Everything is bits, but there's so many industries. Li…”
Liam Fedus Sep 30, 2025 ▶ 27:32
Assertion Not checkable as stated
Cubuk: Frontier AI labs haven't trained LLMs on physics or chemistry
“The frontier AI labs have figured out how to train them on math and logic, but not yet on physics chemistry.”
Doge Cubuk Sep 30, 2025 ▶ 29:53
Insight
Fedus: Pre-training on domain data outperforms retrieval-augmented generation
“However, as we've seen with things like ChatGPT and other things, when you pre-train on the data, when you actually encode the knowledge into the weights, it's not just a retrieval system, you have a richer, deeper understanding of the material.”
Liam Fedus Sep 30, 2025 ▶ 39:20
Insight
Fedus: High-compute reinforcement learning is essential for AI tool use
“High compute reinforcement learning is really effective. This is how you should think about the strategies it's using. This is how you create effective tool using towards those problems, and this is how you optimize it effectively.”
Liam Fedus Sep 30, 2025 ▶ 40:35
Insight
Fedus: Mixing data distributions does not guarantee AI model generalization
“If you just sort of mix together distribution A, B, and C, there's no guarantee of generalization. What you want to hope to see from these systems is the inclusion of this other data set is improving performance on the other data sets.”
Liam Fedus Sep 30, 2025 ▶ 43:02
Assertion Not checkable as stated
Midha: AI models performed poorly on Stanford physics evaluations
“I spent some time running evals on a bunch of these models at the Stanford physics lab earlier this year, and the results were that the models are terrible at scientific analysis.”
Anjney Midha Sep 30, 2025 ▶ 43:29
Disclosure
Cubuk: Periodic Labs mid-trains existing LLMs rather than building from scratch
“We take a pre-trained model and then mid-train it, you know, high computer.”
Doge Cubuk Sep 30, 2025 ▶ 44:06
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.