OpenAI researcher Karina Nguyen describes the synthetic data pipeline used to train interactive behaviors in OpenAI Canvas.
Insight
Nguyen: Post-training scaling avoids data walls through infinite learnable tasks
“The scaling in post-chaining itself is not hitting the wall, and that's because Basically, we went from, like, raw data sets from pre-trained models to infinite amount of tasks that you can teach the model in the post-training world via reinforcement learning.…”
Opinion
Nguyen: AI bottleneck is evaluations rather than data as benchmarks saturate
“We are actually getting saturated in all benchmarks. So I think the bottleneck is actually in evaluations that we don't have all the frontier, like evals, like, I don't know GPGA, which is, like, A Google-proof question answering, like, PhD-level intelligence …”
Insight
Nguyen: Synthetic Data Outperforms Human Data for AI Product Development
“And the reason why I really love, like, synthetic, like, relying purely on synthetic data instead of, like, collecting Data from humans is because it's, like, much more scalable. It's cheap, less than how, like, you literally sample from the model, and you tea…”
Opinion
Nguyen: ChatGPT still struggles with writing due to creative reasoning limits
“I think it's actually really, really hard to teach the model how to be aesthetic or, like, do, like, visual, really good, like, visual design or, like, how to be extremely creative in the way they write. I think, like, I still think, like, Chai GP kind of suck…”
Insight
Nguyen: AI research progress is bottlenecked by research management
“I actually, like, AI research progress is bottlenecked by, like, management. Like, research management is because you have, like, constrained set of compute, and you need to, like, allocate the compute to the research path that you feel the most Commenced abou…”
Prediction Not checkable as stated
Nguyen: AI is not far from autonomous self-improving product development
“And I don't think, like, we are far away from that kind of, like, self-improvement, models becoming, like, self-improved via, like, then, like, the product development is basically kind of, like, self-improving, like, it's kind of, like, its own, like, organis…”