why aren't all 24 resolved? a statement only gets an assessment when the public
record can support or contradict it. opinions and what-ifs never can, and 0 checkable
ones are still open, waiting for their date. predictions held up or didn't;
assertions are supported or contradicted. on every card:
▮▮▮▮▮ certainty ·
▮▮▮▮▮ debate potential. speakers are clickable
Assertion Supported
Lambert: Tulu 3 matches or beats Meta Llama 3.1 on core evals
“On, like, core evals for our Suite of models from, I think, eight, seven D and four or five B is based on llama at the time. It's like it matches or beats meta on these core valves.”
Assertion Supported
Meta Llama 3.3 and Llama 4 perform poorly on agent benchmarks
“Another, of course, the other surprise was that all the Lama models were not performing well on our benchmark. 3.3 and even the Lama four all were really performing extremely poor.”
Assertion Supported
DeepMind, Microsoft, and Meta are building or using physical science labs
“You see people like Google DeepMind, Microsoft, other places like Meta, either building their own lab or running experiments at someone else's lab to get that data back.”
Assertion Partly supported
Patel: Meta is taking on $40B in debt for its Louisiana AI cluster
“Meta, they've, they're already taking debt on for their largest AI cluster in Louisiana. You know, they're taking like forty billion dollars of debt on for that”
Assertion Partly supported
Zhang: Fine-Tuned Llama 3.2 Achieved Superhuman Vision Verification Performance
“We kind of fine-tune our, kind of, for example, NAMA's 3.2 with our, kind of, verification, human annotated verification data. We get, kind of, superhuman performance on these two verification tasks, and then we do not need human on these two tasks. Let's furt…”
Assertion Supported
Meta used stepwise reward models and Monte Carlo Tree Search for Llama 3.1
“They actually went the extra step to, no pun intended, to actually train stepwise reward models. That's kind of crazy, no? I mean, they wanted each step in the chain of thought to be so good that they actually took the extra effort to train step, to train step…”
Assertion Partly supported
Firshman: Llama 2 costs $25M to train but $50 to fine-tune
“Lama II as a base model is that, like, yeah, it costs twenty five million dollars to train to start with, but then you can fine tune it for, like, 50 bucks.”
Assertion Supported
Patel: Broadcom is working on Meta's second-generation internal AI chip
“So Broadcom is, is, is, is working with, on Google's chip right now, but of course on Meta's, Meta's internal AI chip, which they're on the second generation of, working on that.”
Assertion Supported
Ravi: Meta's SA-Co Benchmark Has Over 200,000 Unique Concepts
“If you look at the size of these benchmarks, the previous benchmark, Peng Chuan mentioned, Elvis, that everyone uses, it has about 1.2 K unique concepts and the benchmark that we created, which we're calling segment anything with concepts or Seiko, COCO for sh…”
Assertion Supported
Ravi: Over 70% of SAM 3 Dataset Annotations Are Negative Phrases
“We have about 70, more than 70% of the annotations are these like negative phrases that are not present in the image.”
Assertion Supported
Robinson: Original DiT Research Showed Compute Scaling Outweighed Hyperparameters
“This is so interesting because the original diffusion transformer paper from Facebook actually showed that, in fact, the specific hyperparameters of the transformer didn't really matter that much. What mattered was that you were just increasing the amount of c…”
Assertion Supported
PyTorch was originally built for researchers without considering production requirements
“PyTorch actually started as the framework for researchers. Don't care about production at all.”
Assertion Supported
Scialom: Llama 3 Scaled Pre-Training to 15 Trillion Tokens
“It's the same recipe done in terms of architectures and training than LAMA-II, but we put so much effort on scaling the data and the quality of data. There's now 15 triant tokens compared to two triants, so it's another magnitude there as well, including for t…”
Assertion Supported
Scialom: Meta had to reinvent scaling RLHF without published frontier research
“You have just the basics, but then when it comes to, like, ChatGPT or GPT Instruct or Cloud, No one published the details there. And so we had to reinvent the wheel there in a very short amount of time.”
Assertion Supported
Huang: Curriculum context expansion outperforms full-length training from scratch
“If you train a model on a shorter context and you progressively increase that context to, like,
You know, the final limit that you have, like, 32 K is usually the limit of Lama two was that long.
It actually performs better than if you try to train 32 K the …”
Prediction Held up
Chintala: Meta will have over 600k H100 GPU equivalents by end of 2024
“That is by the end of this year, and 600 K H-One hundred equivalents. With 250 K H-one hundreds and including all of the other GPU or accelerator stuff, it would be 600 and something K aggregate capacity.”
Assertion Supported
Lambert: Meta Used Rejection Sampling to Bootstrapping Llama 2 RLHF
“Llama started their RLHF process with this to get some signal out of preference data. That preference data went into a reward model, and then the reward model did a good enough ranking that it was, like, essentially superpowered instruction tuning based on rew…”
Assertion Supported
Ravi: Meta launched three separate SAM 3 models, not just one
“We launched actually three separate models this time. It was SAM-III, SAM-III objects, and SAM-III body. Those were two completely separate models and SAM III is just the image and video understanding model.”
Assertion Supported
Ravi: Meta Achieved Fully Automated Annotation in SAM 1, Not SAM 2
“Getting to that fully automated data engine is something that we tried to do in SAM too. We actually didn't get to that fully automated approach. In SAM one, we did, we, you know, But the SA-I-B dataset that we released was fully annotated automatically. We di…”
Assertion Supported
Ben Allal: LLaMA 3 used 15x more pre-training tokens than original LLaMA
“LAMA was trained on one trillion tokens, but LAMA-III was trained on 15 trillion tokens.”
Assertion Supported
Frankle: OpenAI, Google, Meta, and Apple have data deals with Shutterstock
“And you know, I, at least I've heard in the news, like opening, I Google, Meta Apple have all called Shutterstock and made those deals.”
Assertion Supported
Feinberg: Genesis CTO Sergey Edunov led Meta's Llama 2 research team
“Sergei led the LLAMA II research team at Meta when he was still there.”
Assertion Supported
Swix: Developers at large tech companies like Meta do not run code locally
“That's how it is at most companies, most like big calls, like Facebook, like nobody runs things locally.”
Assertion Supported
Mohan: Codeium's office housed Ghost Autonomy and 'Silicon Valley' exterior shoots
“It previously was being leased by, I think, Facebook slash WhatsApp. And then immediately after that Ghost Autonomy, and then now here we are, and we also, you know, I guess one of the things that the landlord told us was this was the place that they shot all …”