Karina Nguyen explains the difficulties of standardizing benchmark evaluations across different frontier AI models like Claude and GPT-4.
Assertion Not checkable as stated
Nguyen: Stanford HELM benchmark under-reported Claude performance due to improper prompting
“This has happened with, like, Stanford, I remember, like, when Stanford had lists also, like, they were, like, running benchmarks. Yeah, Helm. And somehow, like, Claude was, like, always, like, not performing well, and that's because, like, the way they prompt…”
Insight
Nguyen: User collaboration is the key milestone before full AI delegation
“Sometimes I feel like a lot of researchers or, like, people in the AI community are, like, so into, like, yeah, agents, delegate everything, like, blah, blah. But, like, on the way towards that, I think, like, collaboration is actually one of the main roadbloc…”
Prediction Not checkable as stated
Nguyen: Website clicks will drop as internet access shifts to AI models
“In my opinion, like, people in, like, few years will click On, like, websites way less. I want to see the plot of, like, website clicks over time, but then my prediction is, like, it will go down and, like, people's access to the internet will be through the m…”
Assertion Not checkable as stated
Anthropic delayed web UI due to Claude 1.3 hallucinations
“And I think, like, at that time, Cloud 1.3 I.E. Had a lot of hallucinations, actually. So I think there was, like, one of the concerns is, like, I don't think, like, the leadership was convinced, had a conviction that this is the model that you need to, like, …”
Insight
Karina Nguyen: OpenAI o1 excels when given explicit hard constraints
“If you give a one like hard, like constraints of like what you're looking for, basically the model would be, we'll have a much easier time to like, kind of like select the candidates and match like the candidate that is most like, fulfill the criteria that you…”
Disclosure
Nguyen: Claude 2's distinct personality was unintentional until Claude 3
“People said, like, Cloud II is, like, so much better at, like, writing and, like, has a certain personality, even though it was, like, unintentional at all. And we did not pay that much attention and didn't know even how to, like, productionize this property o…”