Assertion Not checkable as stated
Nguyen: Stanford HELM benchmark under-reported Claude performance due to improper prompting
“This has happened with, like, Stanford, I remember, like, when Stanford had lists also, like, they were, like, running benchmarks. Yeah, Helm. And somehow, like, Claude was, like, always, like, not performing well, and that's because, like, the way they prompt…”
Insight
Nguyen: User collaboration is the key milestone before full AI delegation
“Sometimes I feel like a lot of researchers or, like, people in the AI community are, like, so into, like, yeah, agents, delegate everything, like, blah, blah. But, like, on the way towards that, I think, like, collaboration is actually one of the main roadbloc…”
Prediction Not checkable as stated
Nguyen: Website clicks will drop as internet access shifts to AI models
“In my opinion, like, people in, like, few years will click On, like, websites way less. I want to see the plot of, like, website clicks over time, but then my prediction is, like, it will go down and, like, people's access to the internet will be through the m…”
Assertion Not checkable as stated
Anthropic delayed web UI due to Claude 1.3 hallucinations
“And I think, like, at that time, Cloud 1.3 I.E. Had a lot of hallucinations, actually. So I think there was, like, one of the concerns is, like, I don't think, like, the leadership was convinced, had a conviction that this is the model that you need to, like, …”
Insight
Nguyen: AI model card benchmark numbers are never apples-to-apples across labs
“None of the numbers are, like, apples to apples. So you actually need to, like, go back to, like, I don't know, like, GPT-E for model card and, like, read the appendix just to, like, make sure that, like, The settings were the same as you're running the settin…”
Insight
Karina Nguyen: OpenAI o1 excels when given explicit hard constraints
“If you give a one like hard, like constraints of like what you're looking for, basically the model would be, we'll have a much easier time to like, kind of like select the candidates and match like the candidate that is most like, fulfill the criteria that you…”
Disclosure
Nguyen: Claude 2's distinct personality was unintentional until Claude 3
“People said, like, Cloud II is, like, so much better at, like, writing and, like, has a certain personality, even though it was, like, unintentional at all. And we did not pay that much attention and didn't know even how to, like, productionize this property o…”
Prediction Open · timeframe Feb 2028
Nguyen: ChatGPT Will Evolve Into an Interface That Morphs Based on User Intent
“Chat CPT evolves into this Blank interface, which can morph itself in whatever you trying, like the model should try to like derive your true intent and then modify the interface based on your intent. And then if you like writing, it should become like the mos…”
Insight
Nguyen: Full Document Rewrites Yield Higher Model Accuracy Than Code Diffs
“We didn't know that, like, code diffs was very difficult for a model, for example. Again, it's like, do we go back to, like, fundamentally improve, like, code diffs as a model capability? Or do you, like, do a workaround where the model will just, like, rewrit…”
Prediction Not checkable as stated
Nguyen: Canvas and Tasks Will Evolve ChatGPT into Something Completely New
“There are different types of like. Features like Canvas, tasks, but all those components that go, they compose together to evolve ChatGPT into something completely new, I think, in the new year.”
Assertion Not checkable as stated
Anthropic built commercial products to self-fund AI safety research
“And that was a time when Antarctic Anthropik really decided to, like, do more product-y related things, and the vision was like, we need to, like, fund research, and, like, building product is, like, the best way to, like, fund safety research”
Assertion Not checkable as stated
Nguyen wrote Claude.ai's first 50,000 lines of code unreviewed
“Yeah, like I think like the first like 50,000 code of lines without any reviews at that time because there's no one. Yeah, it was like very small team. It was like six, seven team who we were called a deployment team.”
What-if
Nguyen: AI canvas interfaces could have happened two years earlier
“I think like those ideas could have happened like two years ago. Just like maybe, I don't think it was like a priority at that time. It was like very unclear. I think like AI landscape at that time was very nascent, if that makes sense.”
Assertion Not checkable as stated
Nguyen: Anthropic Claude 3 post-training team had only 10-12 people
“I was a part of the post-training fine-tuning team. We only had, like, what, like, 10, 12 people involved”
Assertion Supported
Nguyen: Anthropic was first AI lab to publish GPQA benchmark numbers
“I think it was like the first, I think we were the first lab, like, Antarctica was the first lab to, like, run. Publish GPQA, like, numbers”
Opinion
Karina Nguyen: Verification difficulty makes alignment crucial for reasoning models
“The question of like alignment is actually more important for this like complex reasoning models to like, how do we help humans to like verify the outputs of these models is quite important.”
Disclosure
Karina Nguyen: OpenAI Retrained GPT-4o to Handle Canvas Edge Cases
“The only way to like fix some of the edge cases is actually through post training. So we actually, what we did was actually retrain the entire full O plus our canvas stuff.”
Opinion
Nguyen: Original ChatGPT Canvas Beta Model Was More Creative Than GPT-4o Canvas
“I would say like the original better model that we released this canvas was actually much more creative than even right now when I use like for, oh, this canvas”
Insight
Nguyen: AI model training requires evals where prompted baselines fail
“Prototype was prompted baseline. It's all, all, everything starts with, like, prompted baseline, and then, like, we craft, like, certain, like, evaluations that we want to, like, capture, that we want to, like, measure progress, at least, for the model, and th…”
Prediction Not checkable as stated
Nguyen: AI models will evolve to proactively suggest recurring user workflows
“I think that ideally we learn from like the user behavior and ideally the model will just be more proactive in suggesting of like Oh, I can either do this for you every day because I've observed that you do that every day or something. So it's like more become…”
Opinion
Nguyen: OpenAI takes bigger product risks while Anthropic focuses on enterprise
“OpenAI and Anthropik is different in terms of like more like maybe like product mindset. Maybe OpenAI is much more willing to take some of the product risks and explore different bets. And I think Anthropik is much more focused and they have, I think it's fine…”
Insight
Nguyen: AI progress is bottlenecked by human interface creativity
“I feel like we are bottlenecked by like human creativity on like completely changing the way we think about the internet or like some of the way we think about software, like AI right now pushes us to like rethink everything that we've done before in my view.”
Assertion Not checkable as stated
Nguyen: Writing and Coding Are Most Common Use Cases for Canvas
“So for Canvas, for example, one of the most common use cases is basically writing and coding”
Disclosure
Nguyen: Explored a Claude collaborative workspace concept at Anthropic in 2023
“I was working on something similar to, like, Canvas-y, but for Claude at that time, in, like, twenty-twenty-three, it was the same similar idea of, like, Claude workspace where a human and a Claude could have, like, a shared workspace which is like a document.”