Thomas Scialom, technical lead at Meta AI, describes the pre-training data curation pipeline for Llama 3 and how prior models were used as data classifiers.
“LAMA was the best, at the time, before LAMA Free, the best model we had access to legally, to labelize the web and select what are the good tokens and the bad tokens. The additional thing is that it also enabled to have a topic tag, Like, is it about law? Is it about politics? Is it about chemistry, math, reasoning? So that you can also adapt a bit the mixture to like balance a bit more the diversity.”
quote is from the automated transcript, cleaned for reading:
filler sounds and stutters are removed, nothing is rephrased. names can be misheard
(the analysis reads context, assessments check outside sources). how →
More from Thomas Scialom
Insight
Scialom: Overtrain models beyond Chinchilla optimal to minimize inference costs
“And so, to be compute efficient at inference time, it's much better to train it much longer training time, even if it's an effort, an additional effort, than to have a bigger model. That's what I call, like, I refer to the chinchilla trap, Not that Chinchilla …”
Thomas ScialomJul 23, 2024▶ 11:44Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
Disclosure
Meta's Llama 3 post-training uses almost entirely synthetic data
“So what we did is that we generated all the data on the prompts with LAMA-II, and we applied, like, basically the last round of LAMA-II we had to kick off and start LAMA-III post-training. So Lama-free post-training doesn't have any, like, human-written answer…”
Thomas ScialomJul 23, 2024▶ 33:40Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
AssertionSupported
Scialom: Llama 3 405B is the best open-source model ever released
“At a high level, it's the best open source model ever. It's Better than GPT-IV. I mean, what version? But, by far, compared to the version originally released even now, I think there's maybe the last cloud Sonya FF-V and GPT-IV-Zero that are performing it.”
Thomas ScialomJul 23, 2024▶ 37:43Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
PredictionNot checkable as stated
Scialom: Agentic systems will yield order-of-magnitude scaling gains over pre-training
“I expect some incremental and significant progress on pre-training and post-training, but I'm really hopeful that we can gain some order of magnitude of scaling by interconnecting well models into agents as a more complex system that can do planning, that can …”
Thomas ScialomJul 23, 2024▶ 47:14Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
Opinion
Scialom: AI has minted infrastructure unicorns but few successful application companies
“I see like now a lot of fundamental stacks that are like the unicorn of today. Foundational models, foundational like clusters, data notations, things like that. There's a lot, but less successful yet, for now at least, application company. And it's hard to bu…”
Thomas ScialomJul 23, 2024▶ 1:02:32Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
Insight
Scialom: Multilinguality in LLMs emerges naturally with very little data
“Multilinguality almost emerged naturally with very, very few data, which was really surprising and not expected at all for us at the time.”
Thomas ScialomJul 23, 2024▶ 3:48Training Llama 2, 3 & 4: The Path to Open Source AGI — with Thomas Scialom of Meta AI
Made with StarZero
Turn any episode into a week of clips.
This entire site, over 200 episodes transcribed, diarized, checked and made playable,
runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the
moments worth sharing, cuts them, captions them, and reframes them for every feed.
We use essential cookies to make the site work. With your permission we
also use analytics cookies (Google Analytics and Mixpanel) to understand
usage and improve StarZero. See our Cookie Policy.