Direct Preference Optimization

also referred to as: dpo

8 statements across 4 episodes · 4 bullish · 1 bearish · 4 people on the record · first statement Jan 11, 2024 by Nathan Lambert · said 4 times in 3 episodes since 2024 · across every show →

Mentions by year

brought up most by Soumith Chintala (2)

tap a year for its mentions
00213220242025episodesmentions
01220242025episodes it came up in
000.811.5220242025episodesmentions per episode

every mention, scene by scene, with the transcript →

Everything said about Direct Preference Optimization, oldest first

Jan 11, 2024 neutral
Assertion Not checkable as stated
Lambert: AI2 trained 70B TÜLU 2 on the first run without ablations
“Let's just try the Zephyr recipe on seventy billion parameters, and it's literally, like, the first run. It's like, we did no ablations, didn't change any parameters, we just copied them all over. And like, that's the model that people have been working with”
Nathan Lambert Jan 11, 2024 ▶ 1:19:31 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Jan 11, 2024
Insight
Lambert: DPO is closer to RLHF than RLHF is to RL
“I think DPO is closer to RLHF than RLHF is to RL.”
Nathan Lambert Jan 11, 2024 ▶ 1:12:56 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Jan 11, 2024 positive
Assertion Supported
Lambert: DPO Has Become the Standard Release Expectation for Open-Source LLMs
“I think DPO releases are kind of becoming expected because Mistral released a DPO model as well. I think the slide after this is just like, there's a ton. It's like Intel releases DPO models, Stability releases DPO models. At some point, you just have to accep…”
Nathan Lambert Jan 11, 2024 ▶ 1:20:43 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Jan 11, 2024 bullish
Prediction Held up
Lambert: More DPO models will emerge than any other method
“I expect to see more DPO models than anything else in the next six months.”
Nathan Lambert Jan 11, 2024 ▶ 1:15:07 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
Jul 5, 2024 negative
Insight
Yi Tay: Distilled open-source model variants disappeared after failing to climb LMSYS
“When people realize that, like, this, like, turning on the GPT-IV tab and running some DPO is not going to give them the reward signal that they want anymore, right? Then all these variants gone, right? You know, there was this era where there's, wow, there's …”
Yi Tay Jul 5, 2024 ▶ 2:01:00 The 10,000x Yolo Researcher Metagame — with Yi Tay of Reka
Dec 23, 2024 positive
Disclosure
Soldani: Ai2 Trains OLMo Using Community Datasets and Other Models' Outputs
“We see a lot of these even in our own work of like, you know, as we iterate in the various version of Olmo it's not just like every time we collect from scratch all the data. No, the first step is like, okay, what are the cool data sources and datasets people …”
Luca Soldani Dec 23, 2024 ▶ 4:55 Best of 2024: Open Models [LS LIVE! at NeurIPS 2024]
Mar 30, 2026 positive
What-if
Lample: Post-training breakthroughs like DPO were impossible without open LLaMA
“And if you look at many of the techniques that were developed after, for instance, Temma was open source, like all these post-training approaches like even DPOD, like performance optimization, all of this were done by people that had access to this model, and …”
Guillaume Lample Mar 30, 2026 ▶ 38:05 Mistral: Voxtral TTS, Forge, Leanstral, & Mistral 4 — w/ Pavan Kumar Reddy & Guillaume Lample
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.