Chi: Vals AI committed to never sell training data to labs
“At VALS, one very early decision we made was the decision to never sell training data to labs. It's often a place that we're pushed. When we start working with a new lab to actually source and sell for them a bunch of training data.”
Arora: Enterprise AI Success Depends on Fast Training Data Creation
“I think the enterprise's ability to absorb this or digest this or perhaps leverage this to their advantage depends on their ability to create training data as fast as they can. And I think not enough people are focused on training data.”
Antani: AI Models Are Disposable; Data and Harnesses Are Durable
“At the end of the day, the models in AI are disposable. The weights are going to change. The models are going to come out. That doesn't matter. What's durable in AI is the harness and the training data.”
AI model builders demand perpetual data licenses to enable continuous model retraining
“One of the things we've seen is for high quality data, the model builders don't want to part with it. It's not the, we train once, we don't need the data ever again. It's we train once and if it's good, we want it forever because if we retrain our models from …”
Most Protein Design Validation Uses Targets With Training Data Overlap
“One of the things that, you know, we found, ah, with the field was that a lot of the validation, especially outside of the validation that was done on specific problems, was done on targets that have a lot of, you know, known interactions in, in the training d…”
Lacour: Internal GTM AI Models Will Degrade as Company Data Ages
“At one point, this data will get still. So at one point, the AI models will get worse. So we're going to run into this question like, hey, how do we keep everything fresh?”
Bourgeau: AI development is not running out of training data
“The other part of your question are we running out of data? I don't think so, so there's more.”
Agrawal: Frontier AI Training Data Is Saturated, Shifting the War to Talent
“The ingredients of call it being on the frontier are compute. CapEx, we spoke about it. Data which is sort of access accessible to everybody. We are saturated on, on, on the training data that exists. Talent. This is where the war is now.”
New England Journal of Medicine cases avoid AI training data contamination
“Each week they put out a brand new case, which has never even been digitized, so there's no question that it's not in the training data.”
Schulhoff: Prompts formatted like common training data perform best
“It actually comes empirically from studies that have shown that formats of questions that show up most commonly in the training data are the best formats of questions to actually use when you're prompting it.”
Scott: Claims about proprietary data's value for AI lack scientific backing
“Most of those assertions that people make are like most of the assertions that people make are just unfounded in any kind of science. Like no, no one's gotta, and like the measurements we do have show that there's a pretty big disconnect around what some peopl…”
Bryk: Search engines do not need PhD-level training data
“With search, you're asking, like, simple questions about billions of things, like, is this a startup? Did this person write a blog post about search? You know, those are actually simple questions. You don't need, like, PhD level training data.”
Customer inference workloads rarely align with foundation model training distributions
“The data distribution in their inference workload doesn't align with the data distribution in the training data for the model, right? It's a given, actually. If you think about this, because researchers have to guesstimate what is important, what's not importa…”
Gomez: AI models are hyper-sensitive to single bad data examples
“Like a single bad example, right, amongst, like, billions. Like, it's so sensitive. Like, it is a bit surreal how sensitive The models are to their data. Everyone underrates it.”
Altman: AI will not become an arms race for training data
“I definitely don't think it'll be an arms race for data because when the models get smart enough at some point, it shouldn't be about more data, at least not for training.”
Friedberg: AI models do not store training data, but learn synthesized predictors
“The truth is that the models don't actually hold the data that they're trained on. They develop a bunch of synthesized predictors that they learn what the prediction could or should be from the data.”
Sam Lessin: Generative AI content will be homogenizing, not original
“I think there's a very strong argument that it's very homogenizing. You know, yes, you'll get weirder and weirder porn, but it will all hit the same beats, right, that, like, have been trained on the underlying data.”
Evans: Generative AI Infers Patterns Rather Than Storing Training Data
“What is generative AI trying to do? Which is, is not actually trying to reproduce any individual thing in the training data. So, you know, when we did image recognition models, 10 years ago, you give it a billion pictures of cats. The purpose is not that it ca…”
Palihapitiya: Unique web and app data will yield new licensing revenue
“If you're an entrepreneur building a website or building an app that has really unique training data or really unique data, you'll be able to license and sell that. And that'll be an incremental revenue stream to everything you do in the near future.”
Roberge: Top AI talent beats mediocre teams with trillions of data records
“So if you had to choose between having an A plus AI tech talent team with millions of records of data, and you were going against a mediocre AI tech talent team with trillions of data that the startup wins. And that starts to make sense with the diminishing re…”
Kelly: Biological AI's main bottleneck is availability of training data
“The big limitation in bio is the availability of data to train these things, right? And so you have this tough situation where, like, everyone is doing these models of training on the same data, right?”
Language structure and knowledge are embedded directly within raw text data
“The very structure of language and the way to interpret knowledge is actually embedded in the training data itself rather than requiring labeling.”
Biewald: Yahoo Search Success Depended Entirely On Local Training Data Quality
“The model that I'm building is like the same for each country. It's the training data though is different. So some countries would take the training data collection process really seriously and they'd get a great model. And some would just like really half-ass…”
AI Data Needs Scale With the Square Root of Compute
“The data requirements tend to go up like with the square root of the amount of computation, because you're going to train a bigger model and then you're going to throw more data at it.”
Changing deep learning target categories requires retraining models from scratch
“But if you want to change your columns, your categories, then you have to redo all your multiple choice tests and then retrain the system. And if you want to change the categories, you have to start over.”