The Ledger, every show
Every statement that passed quotation and attribution checks, across all 44 shows. Pick shows below, then mix any filter with any other.
shows 




every show 44 of 44
Morcos: Soft inductive biases become harmful past 1M data points in vision
“Turns out in the small data regime, and when I say small data here, I mean, say less than 500,000 data points. And this was in the context of image self-supervised learning. So in that small data regime, this is super helpful. And where this paper's actually b…”
Morcos: Kaplan and Chinchilla scaling laws incorrectly assume all data is equal
“And even if you go and you look at the scaling laws work from Kaplan and Chinchilla and all these other things, they all assume IID data which is insane. We know that all data are not created equal, that garbage in garbage out is like the oldest adage in compu…”
Morcos: DCLM researchers could not predict their own classifiers' filtering decisions above chance
“These are nominally the best experts you could ever hire to do this. These are students who have just spent all of their time looking at NLP data for two years. They could not predict what the DCLM classifiers would say above chance.”
Morcos: Proper data curation can bend neural scaling laws
“And what that paper showed was that if you use your data correctly, you can actually bend the scaling laws themselves.”
Morcos: Datology matches DCLM performance 12x faster with under 10% tokens
“We're able to now get to the same performance as DCLM about 12 x faster. So, you know, in fewer than 10% of the tokens we can match What you get from training to convergence.”
Morcos: Data curation still has at least 100x in performance gains ahead
“You know, we've already been able to get 10 X gains. I think there's at least another hundred X behind this that are still to be done.”
Morcos: Training a specialized frontier model will cost under $1M very soon
“I believe that getting to a frontier model should cost a million dollars or less for most organizations, at least in a specialized domain, right?
And when you think about what enterprises need, that's generally what they need.
They don't need a model that can …”
Morcos: Proper training curricula could reduce model training costs by 10x
“And getting a curriculum right could literally make the difference between, you know, spending 10 times as much on a model training, you know, hundreds of millions of dollars potentially.”
Morcos: Qwen is much easier to align than Llama due to pre-training
“It's much easier to RL Quen than it is to do Lama. Likely that has to do with the fact that Quen put a lot of synthetic reasoning traces into their training data.”
Morcos: AI inference costs will skyrocket, penalizing oversized models
“The inference costs are going to skyrocket with these models. And if you use a general purpose model, then you constrain to say, hey, this model knows about everything, but now only do this one thing. That model is going to have a ton of parameters that do not…”
Morcos: Most AI models used in three years will be under 10B parameters
“Most of the models that the vast majority of people will be using in say three years will be single digit B or smaller.”
Morcos: Yann LeCun was never defining Meta's AI strategy
“I don't think he was ever you know, or at least not since the beginning in a role where he was defining AI strategy for Meta. I don't think that's the role he wanted at any point. You know, I think he really wanted to be doing that research, and I think, so I …”
Morcos: Nemotron dataset quality is similar to DCLM despite token gains
“Nematron is actually pretty similar in quality to DCLM. It's, it came out about six months later. It has more unique tokens. They made a really big deal about it having more unique tokens, but on average, the quality is, is pretty straightforward.”
Morcos: Major AI labs fail to holistically integrate pre-, mid-, and post-training
“Something that you don't see happen even at the big labs because they have entirely separate teams, right? There's a free training team. There's a mid training team. There's a post training team. And like the mid training team is a customer of the free trainin…”
Morcos: Arcee 4.5B beat Gemma before reaching one trillion tokens
“It was beating Gemma pretty consistently before the one trillion mark, which was pretty cool to see.”
Morcos: Meta's metaverse bet will pay off in the long run
“I think the one that's still really up in the air is a metaverse, but I would actually argue that I think that's going to end up paying off in the long run. I think the Ray-Ban glasses pretty darn cool. And a lot of the foundations of what was in reality labs …”
Morcos: Proper data curation enables smaller models with equal or better performance
“Help the folks we work with to train models much faster to much better performance and to also help them train much smaller models to the same or better performance, which I actually think is some of the most exciting stuff going forward. But fundamentally, th…”