Everything Edwin Chen said on any show that made the record, most notable first. Each card names its show and opens the statement there.
Edwin Chen considers company acquisition an admission of failure
“Because yeah, getting acquired would be really limiting. It would be this admission of failure and jumping ship because you can't make it on your own anymore.”
Chen: Big Tech Could Fire 90% of Staff and Move Faster
“Like I used to work at a bunch of the big tech companies, and I always felt that we could fire 90% of people and we would move faster because the best people wouldn't have all these distractions.”
Chen: Researchers Degrade AI Factuality Just to Boost LMSYS Rankings
“A lot of researchers, they'll tell us that their VPs make them focus on increasing their rank on LMSYS. And so I've had researchers explicitly tell me that they're okay with making their models worse. Add factuality works at following instructions as long as i…”
Edwin Chen says he wouldn't sell Surge for $100 billion
“No, I mean, I definitely wouldn't sell for thirty billion or even a hundred billion.”
Surge AI CEO: 2,000 human data points beat 10 million synthetic ones
“A lot of them tell us that even a thousand or a couple of thousand pieces of really high quality human data that we generated for them, it's actually been worth more than ten million pieces of synthetic data.”
Surge AI Surpassed $1B in Revenue With Under 100 Employees
“Yeah, so we hit over a billion of revenue last year with under a hundred people.”
Chen: AI Efficiency Will Enable $100 Billion Revenue-Per-Employee Ratios
“And I think we're going to see companies with even crazier ratios, like a hundred billion per employee in the next few years. AI is just going to get better and better and make things more efficient. So that ratio just becomes inevitable.”
Chen: Frontier AI Labs Game Benchmarks via Prompt Tweaking and Test Leaks
“Sometimes, yeah, these benchmarks, they accidentally leak in certain ways, or the frontier labs will tweak the way they evaluate their models on these benchmarks. Like they'll tweak their system prompt. Or they'll tweak the number of times they run their model…”
Chen: LMSYS Chatbot Arena rewards bolding, emojis, and length over accuracy
“The easiest way to climb Alamarina, it's adding crazy boating. It's doubling the number of emojis. It's tripling the length of your model responses. Even if your model starts hallucinating and getting the answer completely wrong.”
Chen: Vibe coding is overhyped and will make codebases unmaintainable
“I definitely think that Vibe coding is overhyped.
I think people don't realize,
How much it's going to make your systems unmaintainable in the long term and decently dump this code into your code bases.”
Chen: Many YC Founders Only Care About Fundraising Brags and Headlines
“Like, if you talk to a bunch of YC founders or whoever it is, like, what is their goal? It really is to tell all their friends that they raised ten million dollars and show their parents they got a headline on TechCrunch. Like, that is their goal.”
Chen: LMSYS Chatbot Arena is a giant plague on AI
“One of the things I think is a giant plague on AI is Elimsis, Elimarena.”
Edwin Chen: Most AI data competitors are body shops masquerading as tech
“I think a lot of the other companies in our space, they're just not technology companies at the end of the day. They are either body shops or they are body shops masquerading as technology companies.”
Edwin Chen: 100x engineers exist through multiplicative compounding of small advantages
“Some people are simply two to three times better, like two to three times faster than anybody else, right? They just code faster. There are some people who simply have two to three times more better ideas. There are people who simply work two to three times as…”
Edwin Chen: Silicon Valley fundraising is mostly a status game
“I think one of the things that's always driven me crazy about Silicon Valley is that it really is just a status game for most people. Like, People are just raising for the sake of raising. Their goal isn't to build some great product that solves an idea that t…”
Chen: Competitors treat AI data purely as supply, ignoring technology
“There are some companies in this space who will simply think of it as a pure supply problem, and they don't give any consideration to the technology, like both the technology, the underlying technology, like, how do you identify these people? How do you make s…”
Edwin Chen: AI teams are abandoning body shops for high-quality data
“I mean, so I would say I'm pretty sure that a lot of these other companies, they are Like at the end of the day, people want high quality data and they don't want to be working with body shops. And so. I think we've seen, like, a massive wave interest because,…”
Chen: Surge AI has more daily working PhDs than Big Tech combined
“Like if you think of all the PhDs, even at Google or Meta or Microsoft, we have way more than all of them combined doing work for us in a single day.”
Edwin Chen: LMSYS Chatbot Arena is the equivalent of clickbait
“So LM Arena is this popular leaderboard of LM models, and it's basically the equivalent of clickbait.”
Edwin Chen: Dismissing AI safety ignores current issues like benchmark hacking
“So I think a lot of people think AI safety is overblown, but I think they ignore the paperclip maximizer problem where you have AI models that are accidentally trained towards the wrong objectives, even though this is a big problem that all the models face tod…”
Chen expects top future AI model developers to emerge within years
“Don't think so yet. I can actually see big, new, even more powerful model developers appearing in the next few years.”
Chen: AI post-training is an art driven by taste, not pure science
“One of the things I often think about is that there's a, it's almost like there's an art to post training. It's not purely a science. Like when you were deciding what kind of model you're trying to create and what it's good at. There's this notion of taste and…”
Chen: Public AI Benchmarks Are Unreliable and Often Contain Wrong Answers
“I don't trust the benchmarks at all. And I think that's for two reasons. So one is, I think a lot of people don't realize, even researchers within the community, they don't realize that the benchmarks themselves are often honestly just wrong. Like they have wr…”
Chen: AI Will Automate 80% of L6 Engineer Tasks Within 2 Years
“In my head, I probably bet that within the next one or two years, yeah, the models are going to automate 80% of, you know, the average L six software engineer's job. But it's going to take another few years, do you move to 90%, and another few years to 99%, an…”