Direct Preference Optimization

topic on 2 shows · 8 statements across 4 episodes · said 7 times in 4 episodes since 2024

Latent Space 4 the MAD Podcast 3

Mentions by year, every show

tap a year for its mentions
00214220242025episodesmentions
01220242025episodes it came up in
00112220242025episodesmentions per episode

Latent Space 4the MAD Podcast 3

2025 4 mentions in 2 episodes 2 per episode
2024 3 mentions in 2 episodes 2 per episode

every mention on every show, scene by scene, with the transcript →

7 statements about Direct Preference Optimization, every show

Lample: Post-training breakthroughs like DPO were impossible without open LLaMA
“And if you look at many of the techniques that were developed after, for instance, Temma was open source, like all these post-training approaches like even DPOD, like performance optimization, all of this were done by people that had access to this model, and …”
Guillaume Lample Mar 30, 2026 ▶ 38:05 Mistral: Voxtral TTS, Forge, Leanstral, & Mistral 4 — w/ Pavan Kumar Reddy & Guillaume Lample
LATENT SPACE Disclosure
Soldani: Ai2 Trains OLMo Using Community Datasets and Other Models' Outputs
“We see a lot of these even in our own work of like, you know, as we iterate in the various version of Olmo it's not just like every time we collect from scratch all the data. No, the first step is like, okay, what are the cool data sources and datasets people …”
Luca Soldani Dec 23, 2024 ▶ 4:55 Best of 2024: Open Models [LS LIVE! at NeurIPS 2024]
Yi Tay: Distilled open-source model variants disappeared after failing to climb LMSYS
“When people realize that, like, this, like, turning on the GPT-IV tab and running some DPO is not going to give them the reward signal that they want anymore, right? Then all these variants gone, right? You know, there was this era where there's, wow, there's …”
Yi Tay Jul 5, 2024 ▶ 2:01:00 The 10,000x Yolo Researcher Metagame — with Yi Tay of Reka
Lambert: DPO is closer to RLHF than RLHF is to RL
“I think DPO is closer to RLHF than RLHF is to RL.”
Nathan Lambert Jan 11, 2024 ▶ 1:12:56 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
LATENT SPACE Prediction Held up
Lambert: More DPO models will emerge than any other method
“I expect to see more DPO models than anything else in the next six months.”
Nathan Lambert Jan 11, 2024 ▶ 1:15:07 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
LATENT SPACE Assertion Not checkable as stated
Lambert: AI2 trained 70B TÜLU 2 on the first run without ablations
“Let's just try the Zephyr recipe on seventy billion parameters, and it's literally, like, the first run. It's like, we did no ablations, didn't change any parameters, we just copied them all over. And like, that's the model that people have been working with”
Nathan Lambert Jan 11, 2024 ▶ 1:19:31 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert
LATENT SPACE Assertion Supported
Lambert: DPO Has Become the Standard Release Expectation for Open-Source LLMs
“I think DPO releases are kind of becoming expected because Mistral released a DPO model as well. I think the slide after this is just like, there's a ton. It's like Intel releases DPO models, Stability releases DPO models. At some point, you just have to accep…”
Nathan Lambert Jan 11, 2024 ▶ 1:20:43 The Origin and Future of RLHF: the secret ingredient for ChatGPT - with Nathan Lambert

← every entity, every show

Made with StarZero

Turn any episode into a week of clips.

This entire site, thousands of episodes across every show transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.