AI Training Data
topic on 7 shows · 12 statements across 11 episodes
Mixergy
We Live to Build
Latent Space
Lenny's Podcast
No Priors
the MAD Podcast
20VC
12 statements about AI Training Data, every show
Ries: Training AI models on uncompensated human writings is atrocious
“Speaking of someone who's in the training data without any of us being compensated, by the way, which I think is atrocious.”
Catanzaro: AI Data Prep Workloads Are Far Less Predictable Than BI Analytics
“I think one of the things that we saw with analytics that, you know, was surprising to some of the people in the data infrastructure space was that, like, the workloads were actually quite predictable. They were quite predictable because, like, many of them we…”
Morcos: A data point's value depends on its relationship to the full dataset
“The easiest way to think about this is that the value of a data point is not just a function of that data point itself. It's rather a function of how that data point relates to every other data point in the training set, right?”
Frontier AI labs have practically unlimited demand for high-quality data
“In this market, there's essentially like unlimited demand. Like if you can produce high quality volumes of data you most likely will be able to sell whatever you produce.”
Chip Huyen argues human-generated plans are poor training data for AI agents
“When we ask humans to generate like what they consider the best plan for an actions, for a task, it's actually like not quite the best plan for AI, because what is what is easy or efficient for humans is not the same as easy and efficient for AI, right?”
Narayanan: AI training data quality matters far more than quantity
“What we've learned in the last two years is that the quality of data matters a lot more than the quantity of data.”
Human Data Quantity Matters More Than High-Prestige Journalism Sources
“People frequently mistake what quality of data means because they think, oh, you have to train on the New York Times. And actually, in fact most of it is quantity of human data, not because you've had any particular journalist who's, who's got a particularly n…”
Zhang: Pre-training data evolved from a model byproduct into a standalone asset
“So, so I think one fundamental thing that changed in the last year, essentially, in the beginning when people think about data, is, is always like a byproduct of the model, right? You release the model, you also release the data, right? The data side is there …”
Zhang: The world is not running out of AI training data
“I don't think we are running out of data on earth. Right, so think about it globally... But I do think there are many organizations in the world have enough data to actually train, like, very, very good models, right? So, I mean, they are not public available,…”
Smith: AI clones struggle because most creators lack conversational training data
“In most cases with training data is it doesn't have access to a lot of conversational data. That is a little bit different with interviewers that, that have a lot of like podcasts and YouTube. In that case, there's a lot of training data where there's interact…”
Eiso Kant: AI will recycle world data into synthetic data within years
“And so this is kind of our view and our, if you extend this a couple of years out, and this is where some people will probably have objection with us. We think that this will go so far. That we will kind of recycle all of the world's real data into higher qual…”
Data availability is not the true bottleneck for AI scaling
“It's not clear that data really is the bottleneck on performance here. And I've talked to AI researchers about this, and I think there isn't as much of a worry about this as people might think. Probably that's because there's a lot more data that's out there t…”