Disclosure certainty 4/5 debate potential 2/5

Fedus: Early ChatGPT was mathematically weak due to friendliness rewards

Liam Fedus · Building an AI Physicist: ChatGPT Co-Creator’s Next Venture · Sep 30, 2025 · at 7:02

Liam Fedus, co-creator of ChatGPT and co-founder of Periodic Labs, describes the architectural limitations of early RLHF post-training at OpenAI.

0:00 / 0:27exact quote · 27.1s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“The reward functions that we were using originally couldn't determine whether you were mathematically correct or not. So early versions of Chachapiti were mathematically not particularly strong, and it sort of results from the reward function. What did you optimize against? You, the reward function basically encoded, be a friendly assistant, try to help people get to their thing, but it had no sense of, is this mathematically correct or not? Is this code valid or not?”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Liam Fedus

Insight
Fedus: Pre-training on domain data outperforms retrieval-augmented generation
“However, as we've seen with things like ChatGPT and other things, when you pre-train on the data, when you actually encode the knowledge into the weights, it's not just a retrieval system, you have a richer, deeper understanding of the material.”
Liam Fedus Sep 30, 2025 ▶ 39:20 Building an AI Physicist: ChatGPT Co-Creator’s Next Venture
Assertion Not checkable as stated
Fedus: Physics and chemistry demonstrate scaling laws similar to AI
“On the material science side, we're seeing scaling laws within physics, within chemistry both with respect to simulations, with respect to experiment, and it's like the same kind of principles at play and ML.”
Liam Fedus Sep 30, 2025 ▶ 3:02 Building an AI Physicist: ChatGPT Co-Creator’s Next Venture
Insight
Fedus: Physics provides ideal verifiable reward functions for AI
“Physics is very verifiable. It's a great reward function, fairly fast iteration loop. You have simulators for large classes of physical systems.”
Liam Fedus Sep 30, 2025 ▶ 3:35 Building an AI Physicist: ChatGPT Co-Creator’s Next Venture
Disclosure
Fedus: Periodic Labs uses physical experiments as RL reward functions
“And what we're doing, and by having the lab, is we create a physically grounded reward function. That becomes the basis on which we're optimizing against. And so, If a simulator has some deficiencies or some issues, we always error correct, because for us, the…”
Liam Fedus Sep 30, 2025 ▶ 5:01 Building an AI Physicist: ChatGPT Co-Creator’s Next Venture
Insight
Fedus: AI physics requires generating new experimental data, not web scrapes
“The technology that we think is necessary to do it has really just emerged in the last couple of years, and this data Isn't like on a Reddit forum or something like you need to actually go produce experimental data, simulation data.”
Liam Fedus Sep 30, 2025 ▶ 16:08 Building an AI Physicist: ChatGPT Co-Creator’s Next Venture
Insight
Fedus: AI science requires real-world experimental feedback loops
“Ultimately science is driven against experiment in the real world. And so that's what we're doing with periodic labs. We're taking these precursor technologies and we're saying, okay, if you care about advancing science, we need to have experiment in the loop.”
Liam Fedus Sep 30, 2025 ▶ 7:46 Building an AI Physicist: ChatGPT Co-Creator’s Next Venture
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 1,000 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.