Lample: Post-training breakthroughs like DPO were impossible without open LLaMA
“And if you look at many of the techniques that were developed after, for instance, Temma was open source, like all these post-training approaches like even DPOD, like performance optimization, all of this were done by people that had access to this model, and …”
Soldani: Ai2 Trains OLMo Using Community Datasets and Other Models' Outputs
“We see a lot of these even in our own work of like, you know, as we iterate in the various version of Olmo it's not just like every time we collect from scratch all the data. No, the first step is like, okay, what are the cool data sources and datasets people …”
Yi Tay: Distilled open-source model variants disappeared after failing to climb LMSYS
“When people realize that, like, this, like, turning on the GPT-IV tab and running some DPO is not going to give them the reward signal that they want anymore, right? Then all these variants gone, right? You know, there was this era where there's, wow, there's …”
Lambert: DPO is closer to RLHF than RLHF is to RL
“I think
DPO is closer to RLHF than RLHF is to RL.”
Lambert: More DPO models will emerge than any other method
“I expect to see more DPO models than anything else in the next six months.”
Lambert: AI2 trained 70B TÜLU 2 on the first run without ablations
“Let's just try the Zephyr recipe on seventy billion parameters, and it's literally, like, the first run. It's like, we did no ablations, didn't change any parameters, we just copied them all over. And like, that's the model that people have been working with”
Lambert: DPO Has Become the Standard Release Expectation for Open-Source LLMs
“I think DPO releases are kind of becoming expected because Mistral released a DPO model as well. I think the slide after this is just like, there's a ton. It's like Intel releases DPO models, Stability releases DPO models. At some point, you just have to accep…”