why aren't all 65 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 2 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Supported
Huang: Curriculum context expansion outperforms full-length training from scratch
“If you train a model on a shorter context and you progressively increase that context to, like,
You know, the final limit that you have, like, 32 K is usually the limit of Lama two was that long.
It actually performs better than if you try to train 32 K the …”
Opinion
Chintala: Time and data constrain Meta LLM releases more than GPUs
“So, I think the, it's all a matter of time. I think time is the biggest bottleneck. It's like, when do you stop training the previous one, and when do you start training the next one? And how do you make those decisions? The data, do you have net new data, bet…”
Prediction Held up
Chintala: Meta will have over 600k H100 GPU equivalents by end of 2024
“That is by the end of this year, and 600 K H-One hundred equivalents. With 250 K H-one hundreds and including all of the other GPU or accelerator stuff, it would be 600 and something K aggregate capacity.”
Opinion
Lambert: Only 20% to 40% of Meta's Llama RLHF data is useful
“I do think that if we had all the llama data, we wouldn't know what to do with all of it. Like, probably, like, 20 to 40% would be pretty useful for people, but not the whole data set. Like, a lot of it's probably kind of gibberish, because they had a lot of d…”
Assertion Supported
Lambert: Meta Used Rejection Sampling to Bootstrapping Llama 2 RLHF
“Llama started their RLHF process with this to get some signal out of preference data. That preference data went into a reward model, and then the reward model did a good enough ranking that it was, like, essentially superpowered instruction tuning based on rew…”
Assertion Not checkable as stated
Lambert: Meta spent roughly $6M to $8M on Llama 2 preference data
“So I would say, still say, like, six to eight million is safe to say that they're spending, if not more, they're probably also buying other types of data and or throwing out data that they don't like.”
Assertion Supported
Ravi: Meta launched three separate SAM 3 models, not just one
“We launched actually three separate models this time. It was SAM-III, SAM-III objects, and SAM-III body. Those were two completely separate models and SAM III is just the image and video understanding model.”
Disclosure
Ravi: Meta built SAM 3 using open-source community contributions to SAM 2
“In SAM-III we did leverage many of the open source contributions people have made on top of SAM-II.
There were new data sets, there were new benchmarks,
There were new kind of inference time optimizations.
We adopt a lot of the things that the community builds…”
Disclosure
Zhang: Meta intentionally avoided OCR-heavy images during SAM 3 training data sampling
“In fact, during our data engine, we intentionally do not sample OCR-heavy images.”
Assertion Supported
Ravi: Meta Achieved Fully Automated Annotation in SAM 1, Not SAM 2
“Getting to that fully automated data engine is something that we tried to do in SAM too. We actually didn't get to that fully automated approach. In SAM one, we did, we, you know, But the SA-I-B dataset that we released was fully annotated automatically. We di…”
Disclosure
Speak's pronunciation coach fine-tunes Meta's wav2vec on proprietary phonetic transcripts
“We have for English only right now, a pronunciation coach that is basically like a fine-tuned version of WaveDeVec, which is a meta model, but we basically fine-tune it on a bunch of our own phonetic transcripts, like fine-tuned data.”
Assertion Supported
Ben Allal: LLaMA 3 used 15x more pre-training tokens than original LLaMA
“LAMA was trained on one trillion tokens, but LAMA-III was trained on 15 trillion tokens.”
Disclosure
Scialom: Multimodal Llama 3 will add parameters beyond 405B
“For the text text model only? Yes. A bit of additional parameters for the multimodal version that we come later.”
Assertion Supported
Frankle: OpenAI, Google, Meta, and Apple have data deals with Shutterstock
“And you know, I, at least I've heard in the news, like opening, I Google, Meta Apple have all called Shutterstock and made those deals.”
Assertion Supported
Feinberg: Genesis CTO Sergey Edunov led Meta's Llama 2 research team
“Sergei led the LLAMA II research team at Meta when he was still there.”
Assertion Supported
Swix: Developers at large tech companies like Meta do not run code locally
“That's how it is at most companies, most like big calls, like Facebook, like nobody runs things locally.”
Assertion Supported
Mohan: Codeium's office housed Ghost Autonomy and 'Silicon Valley' exterior shoots
“It previously was being leased by, I think, Facebook slash WhatsApp. And then immediately after that Ghost Autonomy, and then now here we are, and we also, you know, I guess one of the things that the landlord told us was this was the place that they shot all …”