Everything Karina Nguyen said on any show that made the record, most notable first. Each card names its show and opens the statement there.
Nguyen: Post-training scaling avoids data walls through infinite learnable tasks
“The scaling in post-chaining itself is not hitting the wall, and that's because Basically, we went from, like, raw data sets from pre-trained models to infinite amount of tasks that you can teach the model in the post-training world via reinforcement learning.…”
Nguyen: AI bottleneck is evaluations rather than data as benchmarks saturate
“We are actually getting saturated in all benchmarks. So I think the bottleneck is actually in evaluations that we don't have all the frontier, like evals, like, I don't know GPGA, which is, like, A Google-proof question answering, like, PhD-level intelligence …”
Nguyen: Synthetic Data Outperforms Human Data for AI Product Development
“And the reason why I really love, like, synthetic, like, relying purely on synthetic data instead of, like, collecting Data from humans is because it's, like, much more scalable. It's cheap, less than how, like, you literally sample from the model, and you tea…”
Nguyen: ChatGPT still struggles with writing due to creative reasoning limits
“I think it's actually really, really hard to teach the model how to be aesthetic or, like, do, like, visual, really good, like, visual design or, like, how to be extremely creative in the way they write. I think, like, I still think, like, Chai GP kind of suck…”
Nguyen: AI research progress is bottlenecked by research management
“I actually, like, AI research progress is bottlenecked by, like, management. Like, research management is because you have, like, constrained set of compute, and you need to, like, allocate the compute to the research path that you feel the most Commenced abou…”
Nguyen: AI is not far from autonomous self-improving product development
“And I don't think, like, we are far away from that kind of, like, self-improvement, models becoming, like, self-improved via, like, then, like, the product development is basically kind of, like, self-improving, like, it's kind of, like, its own, like, organis…”
Nguyen: Anthropic excels at prioritization, while OpenAI takes more product risks
“I would say, like, Antarctic, I learned from Antarctic that, like, They're much better at, like, focusing and, like, prioritization or, like, very, very hard, like, very hardcore prioritization, I guess, and they need to do it. Like, but I think, like, OpenAI …”
Nguyen: Stanford HELM benchmark under-reported Claude performance due to improper prompting
“This has happened with, like, Stanford, I remember, like, when Stanford had lists also, like, they were, like, running benchmarks. Yeah, Helm. And somehow, like, Claude was, like, always, like, not performing well, and that's because, like, the way they prompt…”
Nguyen: User collaboration is the key milestone before full AI delegation
“Sometimes I feel like a lot of researchers or, like, people in the AI community are, like, so into, like, yeah, agents, delegate everything, like, blah, blah. But, like, on the way towards that, I think, like, collaboration is actually one of the main roadbloc…”
Nguyen: Website clicks will drop as internet access shifts to AI models
“In my opinion, like, people in, like, few years will click On, like, websites way less. I want to see the plot of, like, website clicks over time, but then my prediction is, like, it will go down and, like, people's access to the internet will be through the m…”
Nguyen: Small distilled models like Claude 3 Haiku outperform larger predecessors
“Smart, small models are becoming even smarter than, like, large models. And that's because of, like, the distillation research. This happened with, like, Cloud Tree Haiku. I was like working on like post-chaining of like Cloudy Haiku, and I realized it was muc…”
Nguyen: AI will excel at strategy by synthesizing disparate data sources
“Strategy is, like, it's more, like, data analysis and, like coming up with, like, I think what models are really good at is, like, connecting the dots, I think. It's like, okay, if you have user feedback from this source, but you also have an internal, like, d…”
Nguyen: Pixel-based perception is much harder to scale than language in AI
“Much of it is, like because right now the models operating on, like, pixels instead of, like, language or whatnot, like, pixels is actually really, really hard for the models because, like, perception or visual perception. I think there's still, like, a lot of…”
Anthropic delayed web UI due to Claude 1.3 hallucinations
“And I think, like, at that time, Cloud 1.3 I.E. Had a lot of hallucinations, actually. So I think there was, like, one of the concerns is, like, I don't think, like, the leadership was convinced, had a conviction that this is the model that you need to, like, …”
Nguyen: AI model card benchmark numbers are never apples-to-apples across labs
“None of the numbers are, like, apples to apples. So you actually need to, like, go back to, like, I don't know, like, GPT-E for model card and, like, read the appendix just to, like, make sure that, like, The settings were the same as you're running the settin…”
Karina Nguyen: OpenAI o1 excels when given explicit hard constraints
“If you give a one like hard, like constraints of like what you're looking for, basically the model would be, we'll have a much easier time to like, kind of like select the candidates and match like the candidate that is most like, fulfill the criteria that you…”
Nguyen: Claude 2's distinct personality was unintentional until Claude 3
“People said, like, Cloud II is, like, so much better at, like, writing and, like, has a certain personality, even though it was, like, unintentional at all. And we did not pay that much attention and didn't know even how to, like, productionize this property o…”
Nguyen: ChatGPT Will Evolve Into an Interface That Morphs Based on User Intent
“Chat CPT evolves into this Blank interface, which can morph itself in whatever you trying, like the model should try to like derive your true intent and then modify the interface based on your intent. And then if you like writing, it should become like the mos…”
Nguyen: Full Document Rewrites Yield Higher Model Accuracy Than Code Diffs
“We didn't know that, like, code diffs was very difficult for a model, for example. Again, it's like, do we go back to, like, fundamentally improve, like, code diffs as a model capability? Or do you, like, do a workaround where the model will just, like, rewrit…”
Nguyen: Model Training Is More Art Than Science, Debugged Like Software
“Model training is more an art than a science, and in a lot of ways, like, we as, like, model trainers think a lot about, like, data quality. So, like, it's one of the most important things in model training is, like how do you ensure the highest quality data f…”
Nguyen: Claude Refused Setting Alarms After Realizing It Lacked a Body
“One of the things that I've learned early days at Anthropic was, like, we've discovered, especially with, like, cloud three training, when you taught the model some of the self-knowledge of, like, hey, like, you actually don't have a physical body to operate, …”
Nguyen: OpenAI built Canvas and Tasks features mostly via synthetic data
“The way we made Canvas and tasks and, like, new, like, product features for HTTP was mostly done by synthetic training.”
OpenAI used o1 synthetic data to train Canvas commenting behaviors
“The way we used it is, like, we would use a one model to produce, to, like, simulate, like, use a conversation. Let's say, like, write me a document about XYZ, but then we used a one to, like, produce the document, and then we kind of injected, like, user prom…”
Nguyen: Optimizing AI models constantly causes capability regressions across all labs
“If you optimize the model for this behavior, like, you kind of don't want to, like, brain damage in, like, other areas of intelligence, or, and this is happening, like, all the time in every lab and every, like, research team.”