why aren't all 13 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Disclosure
Meta's Llama 3 post-training uses almost entirely synthetic data
“So what we did is that we generated all the data on the prompts with LAMA-II, and we applied, like, basically the last round of LAMA-II we had to kick off and start LAMA-III post-training. So Lama-free post-training doesn't have any, like, human-written answer…”
Disclosure
Scialom: Meta Used Llama 2 to Filter and Tag Llama 3 Pre-Training Data
“LAMA was the best, at the time, before LAMA Free, the best model we had access to legally, to labelize the web and select what are the good tokens and the bad tokens. The additional thing is that it also enabled to have a topic tag, Like, is it about law? Is i…”
Assertion Supported
Huang: Training CodeLlama on Llama 2 caused catastrophic language forgetting
“We do have historical precedent where CodeLlama was, you know, trained further from the original CodeLlama was trained further from Lama II, and it just lost, All its language capabilities, basically, right?”
Assertion Partly supported
Firshman: Llama 2 costs $25M to train but $50 to fine-tune
“Lama II as a base model is that, like, yeah, it costs twenty five million dollars to train to start with, but then you can fine tune it for, like, 50 bucks.”
Assertion Supported
Bakouch: DeepSeek-V3 uses the same Adam optimizer parameters as Llama 2
“And for example, a good a good way to view that is that DeepSeq rig three is still using the same Adam parameter than Lama two.”
Assertion Supported
Scialom: Llama 3 Scaled Pre-Training to 15 Trillion Tokens
“It's the same recipe done in terms of architectures and training than LAMA-II, but we put so much effort on scaling the data and the quality of data. There's now 15 triant tokens compared to two triants, so it's another magnitude there as well, including for t…”
Disclosure
Scialom: Meta expanded Llama 3's vocabulary size to support multilingual capabilities
“Lama III compared to Lama II is multilingual, has multilingual capabilities. We worked on that. And so, because you have languages that are not just Latin languages like English, there's a lot of different characters you want to include them to represent, like…”
Disclosure
Meta skipped coding, reasoning, and multilingual annotations for Llama 2
“And we didn't annotate at all for code, neither for reasoning or multinguity.”
Assertion Supported
Huang: Curriculum context expansion outperforms full-length training from scratch
“If you train a model on a shorter context and you progressively increase that context to, like,
You know, the final limit that you have, like, 32 K is usually the limit of Lama two was that long.
It actually performs better than if you try to train 32 K the …”
Disclosure
Firshman: Llama 2 release was Replicate's biggest week of growth ever
“Llama II was, like, our biggest week of growth ever, because, like, tons of people wanted to tinker with it and run it.”
Assertion Supported
Lambert: Meta Used Rejection Sampling to Bootstrapping Llama 2 RLHF
“Llama started their RLHF process with this to get some signal out of preference data. That preference data went into a reward model, and then the reward model did a good enough ranking that it was, like, essentially superpowered instruction tuning based on rew…”
Assertion Not checkable as stated
Lambert: Meta spent roughly $6M to $8M on Llama 2 preference data
“So I would say, still say, like, six to eight million is safe to say that they're spending, if not more, they're probably also buying other types of data and or throwing out data that they don't like.”
Disclosure
Scialom: Llama 1 and 2 flagship size was chosen to reproduce Chinchilla
“Lama two, maybe I would say it's like Lama one. We had a flagship model, which was seven TB. It's also because the project was taking some routes to reproducing a chinchilla, which was a seven TB.”