Thomas Scialom (Meta AI research scientist and Llama 3 post-training lead) clarifies parameter specifications during the launch of the Llama 3 family of models.
Insight
Scialom: Overtrain models beyond Chinchilla optimal to minimize inference costs
“And so, to be compute efficient at inference time, it's much better to train it much longer training time, even if it's an effort, an additional effort, than to have a bigger model. That's what I call, like, I refer to the chinchilla trap, Not that Chinchilla …”
Disclosure
Meta's Llama 3 post-training uses almost entirely synthetic data
“So what we did is that we generated all the data on the prompts with LAMA-II, and we applied, like, basically the last round of LAMA-II we had to kick off and start LAMA-III post-training. So Lama-free post-training doesn't have any, like, human-written answer…”
Assertion Supported
Scialom: Llama 3 405B is the best open-source model ever released
“At a high level, it's the best open source model ever. It's Better than GPT-IV. I mean, what version? But, by far, compared to the version originally released even now, I think there's maybe the last cloud Sonya FF-V and GPT-IV-Zero that are performing it.”
Prediction Not checkable as stated
Scialom: Agentic systems will yield order-of-magnitude scaling gains over pre-training
“I expect some incremental and significant progress on pre-training and post-training, but I'm really hopeful that we can gain some order of magnitude of scaling by interconnecting well models into agents as a more complex system that can do planning, that can …”
Opinion
Scialom: AI has minted infrastructure unicorns but few successful application companies
“I see like now a lot of fundamental stacks that are like the unicorn of today. Foundational models, foundational like clusters, data notations, things like that. There's a lot, but less successful yet, for now at least, application company. And it's hard to bu…”
Insight
Scialom: Multilinguality in LLMs emerges naturally with very little data
“Multilinguality almost emerged naturally with very, very few data, which was really surprising and not expected at all for us at the time.”