Everything Chip Huyen said on any show that made the record, most notable first. Each card names its show and opens the statement there.
Huyen: Internet data is maxed out, making post-training the key AI differentiator
“At some point, we are actually, like, have kind of maxed out on, like, internet data, right? And then people, like, text data, people max out. I think a lot of people are doing, like, with other data, like audios and videos, and, like, everyone's trying to thi…”
Huyen: AI data labeling startups face severe customer concentration risk
“It's very lopsided, right? Because like, is there only like a very small numbers of frontier labs, right? And they want a lot of data. And there's like a massive amount of like startups or companies providing data. So like, you can see these companies, like th…”
Huyen: Engineering managers prefer headcount while VPs prefer AI coding agent subscriptions
“Would you rather have access could you rather have, give everyone on the team, like very expensive. Coding agent subscriptions, or you get an extra head count, right? Let's say it's like maybe like and almost everyone could say the managers could say head coun…”
Huyen: A randomized engineering trial found top performers gained most from Cursor
“He was like, okay, here's more like currently like best performing, average performing, and lowest performing. And then there's a randomized trial. So like they give like half of each group, like access to like cursor. And then he was noticed like over time, I…”
Chip Huyen argues post-training is what differentiates frontier AI models
“So, so I do think that post-training is what makes this, like, really big lab models are, like, different.”
Evaluation is the single biggest bottleneck holding back enterprise AI adoption
“So, so I do think that evaluation is the biggest bottleneck for AI adoptions, because unless, like, if we can, like, if we can, like, develop a more reliable way to evaluate the application, that application is not going to get adopted. Like, or maybe, maybe i…”
Chip Huyen warns million-token context capacity does not imply efficient processing
“Just because a model, I think the second reason may be actually more important at least for now is that just because a model can fit in a million con token context doesn't mean that it can process that million token efficiently.”
Chip Huyen argues human-generated plans are poor training data for AI agents
“When we ask humans to generate like what they consider the best plan for an actions, for a task, it's actually like not quite the best plan for AI, because what is what is easy or efficient for humans is not the same as easy and efficient for AI, right?”
Huyen: Legacy systems prevent US companies from matching Chinese online learning
“So a lot of American, American internet companies are like a lot older than the average, like the new, like Chinese internet company. So it means it's like American internet companies have legacy systems that you from like, 20 years ago. And just have to build…”
Chip Huyen: AI becomes harder to evaluate as intelligence increases
“As a more intelligent AI becomes like the harder it is to evaluate it.”
Chip Huyen: Machine translation is largely solved for major languages
“Now it was like pretty much like people are saying that machine translation is like pretty much sold for like major languages.”
Chip Huyen: Lack of labeled data requirements makes language modeling uniquely scalable
“You don't need to curate, like, labels, like, reference data, so that, that you can use a train models that make language modeling, like, so, so much easier to scale than other types of tasks.”
Chip Huyen predicts users will always expand data to fill available context
“I think we always expand our usage to fit in whatever context length. That's going to be available.”
Chip Huyen: Machine learning projects should start with problems, not models
“So so the main idea is you go backward from the problems. So I think a lot of approach machine is saying it's like, you start with the solutions and it tries to have like five problems when machine can I be applied. So, and it's like, so tend to be like, oh, H…”
Chip Huyen: MLOps startups ignore real-time learning for low-hanging fruit
“We don't have tools for it yet. And I see very, very, very little tools focusing on it because most people are like targeting on focusing on low hanging fruit.”
Huyen: Alibaba and ByteDance lead US firms in online ML scale
“When I was looking into online learning and I realized it's like on the examples I felt were by Chinese companies. And it could be I think I've heard some American companies doing that, but they are doing a much smaller scale, like a lot less complex models th…”
Huyen: Evaluate Marginal Gains and Switching Costs Before Adopting New AI Tech
“And I think it's a specific question you should ask them is like, first, Like if how much of the improvements could you get, like from like optimal solutions versus non-optimal solutions, right? And sometimes they were like, actually it's not much, right? And …”
Chip Huyen: Sampling strategy is very underrated for boosting model performance
“Sampling strategy, I think is something extremely important. It can have you boost the performance in a huge way and very, very underrated.”
Huyen: Data preparation drives bigger RAG gains than database choice
“Data preparations for Rack is extremely important. And I would say that's like in the, a lot of the companies that I have seen, that's like the biggest performance in their Rack solutions coming from like better data, data preparations, not agonizing over like…”
Huyen: Companies Buy Customer-Facing AI Because Outcomes Are Measurable
“A lot of applications companies pursue because they can't measure the concrete outcome. And I feel like booking on a sales chatbot is very clear, right? Like what's a conversion rate right now with that chatbot, with human operators and what could be a convers…”
Huyen: Companies restructure engineering orgs for seniors to review AI-generated code
“Or when our company have seen us the way they work, as they told me is they work completely different now. I like, so they actually restructured engineering org so that like they get more senior engineers should be more in the peer review. Because they like to…”
Huyen: AI coding tools struggle with multi-component existing codebases
“So I'm not sure you use a lot of effort coding, but like or something I've noticed and also seen from my friends, it's like, it is pretty good when you have very clear, well defined tasks, maybe write documentations, fix the specific features, or like build an…”
Huyen: Hyper-specialization prevents tech workers from generating product ideas
“Because I, we have gone through we have gone into this phase of like specializations, like people like very highly specialized and people are supposed to do like focus on one thing really well, instead of being a big picture. And we don't have a big picture of…”
Chip Huyen: AI engineering requires developers to have stronger product sense
“It requires engineers to have a much, much better product sense.”