why aren't all 6 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Supported
Huang: Training CodeLlama on Llama 2 caused catastrophic language forgetting
“We do have historical precedent where CodeLlama was, you know, trained further from the original CodeLlama was trained further from Lama II, and it just lost, All its language capabilities, basically, right?”
Assertion Partly supported
Firshman: Llama 2 costs $25M to train but $50 to fine-tune
“Lama II as a base model is that, like, yeah, it costs twenty five million dollars to train to start with, but then you can fine tune it for, like, 50 bucks.”
Assertion Supported
Bakouch: DeepSeek-V3 uses the same Adam optimizer parameters as Llama 2
“And for example, a good a good way to view that is that DeepSeq rig three is still using the same Adam parameter than Lama two.”
Assertion Supported
Scialom: Llama 3 Scaled Pre-Training to 15 Trillion Tokens
“It's the same recipe done in terms of architectures and training than LAMA-II, but we put so much effort on scaling the data and the quality of data. There's now 15 triant tokens compared to two triants, so it's another magnitude there as well, including for t…”
Assertion Supported
Huang: Curriculum context expansion outperforms full-length training from scratch
“If you train a model on a shorter context and you progressively increase that context to, like,
You know, the final limit that you have, like, 32 K is usually the limit of Lama two was that long.
It actually performs better than if you try to train 32 K the …”
Assertion Supported
Lambert: Meta Used Rejection Sampling to Bootstrapping Llama 2 RLHF
“Llama started their RLHF process with this to get some signal out of preference data. That preference data went into a reward model, and then the reward model did a good enough ranking that it was, like, essentially superpowered instruction tuning based on rew…”