Soldani: Ai2 Trains OLMo Using Community Datasets and Other Models' Outputs
Luca Soldani · Best of 2024: Open Models [LS LIVE! at NeurIPS 2024] · Dec 23, 2024 · at 4:55
Luca Soldani of the Allen Institute for AI (Ai2) explains how the open-source AI community accelerates research by sharing datasets and model outputs during training.
“We see a lot of these even in our own work of like, you know, as we iterate in the various version of Olmo it's not just like every time we collect from scratch all the data. No, the first step is like, okay, what are the cool data sources and datasets people have put together for language model for training? Or when it comes to like our post training pipeline we one of the steps is You want to do some DPO, and you use a lot of outputs of other models to improve your preference model.”
quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →