Krishnan: Purely model-generated synthetic data hits a wall without human feedback
“There is also, I think, limits to how much value can be added there because ultimately there is just not too much new information. If you're telling the model itself to kind of generate, yeah, There is a, I mean, there's all sorts of things about how beyond th…”
Krishnan: Human-designed prompt and auto-verifier tuples maximize synthetic training data ROI
“The, this method is the one, I think, which has a lot of legs in the, particularly the more you can operate in this particular paradigm of prompt and then these rule or rubric based verifier tuples. The, that is a very nice way for sort of creating synthetic d…”
Ambati: Monthly AI model churn is fueled by synthetic data loops
“Every month, if you will, there is a new model, right? That's beating the old model, and visibly adopted already, and people are moving to the new models and the data is just flowing, like you're generating data, you're sort of getting access to new data creat…”
Hong: Accumulating synthetic AI data is not a moat, just buffer
“I think everyone is trying to accumulate like a data, which is not a mode. It's just time and time mode. It's all about like, you know, whether you can execute fast enough to make sure that you have like a certain buffer because of say your data set, you know,…”
Ethan He: Video models require image foundations and 100% synthetic caption pairs
“Building a video model. You actually need to build a image model first and building, building these two models. The data you need is a hundred percent synthetic pair of language and image or language to video because on the internet, actually the videos Don't …”
Sun: Synthetic data matches real-world data for multimodal model pre-training
“We were actually generating a lot of synthetic data and showing that, hey, you can actually, these synthetic data are actually as useful as real-world data when it comes to multimodal pre-training.”
Mensch: AI synthetic data is efficient but cannot replace human training signal
“It's mostly an efficient way of training models to have bigger models that are used as teachers for smaller models, but it's not enough. And so you also need human signal.”
Patel: Companies will differentiate AI models via proprietary, synthetic, and machine data
“And every company is going to differentiate based on their own proprietary enterprise data being used to train the models, synthetic data and machine data, which is where the most amount of growth is.”
Fitzpatrick: Belief that synthetic data replaces human feedback is wrong
“Look, I think the biggest one is just the view that synthetic data will take over, and you just will not need human feedback.”
Santos: Dream Stories trains consistent models using synthetic photo variations
“We were able to reduce the number of pictures required to literally just one, because then we're like, holy , you know, we can essentially just ask for one picture, generate the synthetic version of it, and then say, do you like this? Great. So we generate a f…”
Suleyman: AI Training Is Not Data-Constrained Due To Synthetic Data
“We are not data constrained right now, we're generating vast amounts of high quality synthetic data, which is proving to be useful.”
Suleyman: Synthetic Data And Human Feedback Will Outweigh Robotics Data
“I don't think in the next few years it's gonna be the big differentiator. I think that more synthetic data, more human feedback and high quality data is gonna be the differentiator.”
Chubuk: Minimal physical experiments carry huge information value by validating synthetic simulations
“What's interesting about scientific data is it's not just a few bits or numbers, right? Like, for example, there are certain experiments you can run where the result you get from it is just, say, three floating point numbers. But the implications of those coul…”
Pineau: Synthetic image and language data causes model degradation
“So in some domains, if you think like images, languages, like LLMs talking to each other at some point, you definitely get the degradation and that degradation is due to essentially like a loss of diversity of your data.”
Joseph: Training purely on raw LLM generations cannot produce a better model
“Theoretically, I shouldn't be able to train a better model than that. Like, I'm just going to get the same thing out. So I think that's-”
Valenzuela: Runway is exploring synthetic data for AI video training
“It is. I think it's becoming more of Thing, I would say. It still has its challenges, mostly to generate diversity of data, but definitely something we're exploring.”
Morcos: Filtering synthetic data between generation cycles prevents model collapse
“If you filter the data after each point, that's now information injection, and that can break all of this and I think can prevent model collapse.”
Lord: Synthetic data will not dominate frontier AI training
“Synthetic data has a role to play and like in verifiable domains, but like what we consistently hear from companies is like, you know, their synthetic data is not going to dominate.”
Cuban: Synthetic AI data cannot invent what human scientists discover
“Cause there's no way to synthesize all that shit. You're not synthesizing. You're not creating synthetic data that all of a sudden is going to, you know, you could tell, you give it all the backstory you want, that you're a doctor, that you invented this, that…”
Fulford: OpenAI bootstraps browsing models to generate synthetic training data
“For initial deep research, there's not really any data sets that exist for browsing in the same way that you have a math data set that already exists. So we have to create all this data. But once you have good browsing models or good computer use models, you c…”
Lambert: Labs will surely use parallel-compute models to generate synthetic data
“Well, I bet people, I mean, they surely will use these for synthetic data. It's just like the marginal gain on synthetic data is always very high.”
Chen: AI RL Environments Are Too Complex to Create Synthetically
“I think one of the things that people really underestimate is how it is, how complicated it is that you can't just synthetically generate it.”
Chen: Surge AI Heavily Uses Synthetic Data to Supplement Human Labelers
“Like we use it like a ton ourselves in order to supplement what the humans do.”
Chelsea Finn: Real robot data cannot be replaced by synthetic data
“I think that at the end of the day, there's going to be no replacement for real data. And so we're like large amounts of real robot data. It's going to be a necessary component of any like system that's going to work in a generalizable way.”
Chelsea Finn: Synthetic data's robotic analog is RL, not simulation
“I think that the analog of synthetic data in language models is actually not necessarily simulation in robotics, but closer to something like reinforcement learning.”
Edwin Chen: Synthetic data makes AI models good at benchmarks, not real problems
“Synthetic data, it's made models good at synthetic problems, not, not real ones.”
Surge AI CEO: 2,000 human data points beat 10 million synthetic ones
“A lot of them tell us that even a thousand or a couple of thousand pieces of really high quality human data that we generated for them, it's actually been worth more than ten million pieces of synthetic data.”
Bornstein: Synthetic data will not produce self-improving AI models
“The question is, like, does this lead to sort of, like, a self-improving utopia of models or not? And I think we have some pretty strong opinions on the not side of that.”
Laskin: Reinforcement learning is the only scalable path for synthetic data
“When we're generating synthetic data there is the only scalable path is really reinforcement learning.”
Morin: Early research proves synthetic data cannot effectively train AI models
“Yeah, I mean, that goes back to that first research paper that came out right after ChatGPT launched, which basically said you can't create Fake data for training models. It just doesn't work, and so there was, like, this very early paper that effectively told…”
Gomez: Synthetic data makes up the majority of Cohere's training data
“Synthetic data is incredibly effective. It's now the majority of the data that we train on for creating something like command A.”
Patel: Human labeling is unscalable, forcing reliance on synthetic AI data
“Using humans to train models is just so expensive, right? So then there's the magic of sort of reinforcement learning and other synthetic data technologies, right? Where the model is helping teach the model, right? So you have many models in, in, in a sort of,…”
Feldman: In five years, almost all AI training data will be synthetic
“Almost all synthetic.”
Kiela: DeepSeek proved frontier AI models can rely on synthetic data
“We have kind of an existence proof now that it's actually not that hard to do this and so you don't need to invest all that much in, in data, and you can use synthetic data and get a pretty good model out of that”
Ross: LLM-generated synthetic data is better for model training
“You could have an LLM generate synthetic data, and when it generates the synthetic data, the data is better. You then train on that synthetic data.”
Nguyen: OpenAI built Canvas and Tasks features mostly via synthetic data
“The way we made Canvas and tasks and, like, new, like, product features for HTTP was mostly done by synthetic training.”
Nguyen: Synthetic Data Outperforms Human Data for AI Product Development
“And the reason why I really love, like, synthetic, like, relying purely on synthetic data instead of, like, collecting Data from humans is because it's, like, much more scalable. It's cheap, less than how, like, you literally sample from the model, and you tea…”
Palafox: Autonomous vehicle companies will not outsource synthetic data to startups
“Like there's just not enough enterprises that need that. And the ones that actually need it, it's so core. Like think think like Cruz or like Waymo. It's core to them. Like they're not gonna like outsource that to like a shitty startup.”
Ben Allal: Properly curated synthetic data prevents model collapse
“And I think there's a lot of concerns about model collapse, and I'm going to talk about that later, but we'll see that like, if we use synthetic data properly and we curate it carefully that shouldn't happen.”
Ben Allal: Recent web dumps improve model benchmarks despite synthetic data
“So what we did is we trained different models on these different dumps, and we then computed their performance on popular like NLP benchmarks, and then we computed the aggregated score. And surprisingly, you can see that the latest dumps are actually even bett…”
Ben Allal: Synthetic data may enrich the web rather than pollute it
“So personally, I wouldn't say the web is posted with synthetic data. Maybe it's even making it more rich.”
Ben Allal: Prompt diversity is essential for scaling synthetic training data
“The key ingredient to getting a good data set that is synthetic is trying as much as possible to keep it diverse, because if you just throw the same prompts as your model, like generate, like, a textbook about linear algebra, and even if you change the tempera…”
Synthetic college text boosts MMLU; middle school text boosts OpenBookQA
“College textbooks are really good for MLU or middle school textbooks are good for benchmarks like open book, UA and Pico.”
Ben Allal: Small models can generate synthetic data by rephrasing web pages
“The interesting thing in this approach is that you can use a model that is small Because it doesn't, rewriting doesn't require knowledge. It's just rewriting a page into a different style. So the model doesn't need to have like knowledge that is like extensive…”
Ben Allal: Pre-training on rewritten C4 web data outperforms raw C4
“They rewrite some samples from C four into Q and A into Wikipedia, and they find that doing this works better than training just on C four.”
Ben Allal: Pooling multiple teacher models produces superior synthetic datasets
“Synthetic data, it doesn't have to come from a single model. And because we have so many good models now, you could like pull these models together and get like a dataset that's over really high quality and that's diverse and that's covers all your needs.”
Patel: Synthetic data generation enables continued AI scaling despite data limits
“You can create data out of thin air almost, right? In certain domains, right? And so this is the whole, the debate around scaling laws is how can we create data?”
Patel: AI industry is in early days of synthetic data
“Where have we gone on synthetic data? Oh, we're still like very early days, right? We've spent tens of millions of dollars maybe on synthetic data.”
Patel: Synthetic training only works in functionally verifiable domains like math
“We can't teach it what good art is. Because we have no way to functionally prove what good art is. We can teach it to write really good software. We can teach it how to do mathematical proofs. We can teach it how to engineer systems, because there are, while t…”
Hoffman: LLM scaling is not out of data thanks to synthetic data
“It's like, well, actually, in fact, we can create synthetic data, and there's a ton of data that's out there that's not part of the standard internet training corpus. So, the scale game is still playing.”
Wang: Pure synthetic AI training data has underperformed industry expectations
“One of the things that we've seen over the past few past year in particular is that synthetic data has not worked as well as I think everybody had hoped. You know, pure synthetic data, just using data generated from the models to try to train future models, th…”
Pre-training on human data is hitting limits; synthetic data is required
“So I think on the data side, we're approaching the limit and the only data to increase that is synthetic generated data.”
Gomez: Synthetic data struggles outside verifiable domains like math
“Synthetic data Probably doesn't get us out of that, that issue. I actually, I don't know if synthetic data outside of easily verifiable domains like math, it's hard to use synthetic data to drive outcomes.”
Gomez: Synthetic Data Makes Up a Growing Portion of Cohere's Training
“More and more synthetic data is becoming a huge chunk of the data that we train on.”
Most Open Vision Models Rely on Synthetic Data From Proprietary Models
“Most VLMs are distillations of proprietary closed source models, right? So if you need to generate synthetic data, like most open weight models rely heavily on synthetic data from private models.”
Kant: Synthetic data only improves AI when evaluated by an objective oracle
“If you have something that can determine an oracle of truth that can help say, this is better and this is worse, or this is correct and this is wrong, that's when you can actually use synthetic data.”
Jeff Schmidt: Hermes pioneered synthetic data training before it was standard
“So Hermes was very early to the idea that you could have synthetic data, which is that you could actually make, you could make a better model by taking an AI model, having it generate words and text, and then training a new a model on that output. This is now …”
Karpathy: AI will not run out of training data due to synthetic data
“So I think basically synthetic data is absolutely the future. We're not going to run out of data, is my impression. I just think you have to be careful.”
Desai: Synthetic healthcare data fails to mimic true patients
“A lot of solutions out there, especially in healthcare, use synthetic data. Synthetic data does not mimic a true patient.”
Synthetic data will make individual media training data irrelevant
“Don't try to hold out for money on the training side of things, because, you know, we're going to create synthetic data, we're going to do all kinds of other things that are going to mean that no one's particular data is really going to matter.”