Nguyen: Synthetic Data Outperforms Human Data for AI Product Development
Karina Nguyen · OpenAI researcher on why soft skills are the future of work | Karina Nguyen · Feb 9, 2025 · at 31:55
Karina Nguyen, researcher at OpenAI, explains why she prefers training AI models purely on synthetic data rather than human-annotated datasets.
“And the reason why I really love, like, synthetic, like, relying purely on synthetic data instead of, like, collecting Data from humans is because it's, like, much more scalable. It's cheap, less than how, like, you literally sample from the model, and you teach the core behaviors of the models, and that will generalize to all sorts of diverse coverage. And when you launch the beta feature, you learn so much from the users that you can, like, all your synthetic sets can be shifted in the distribution of how the users behave in Find the product behavior, and this is how you improve.”
quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →