Opinion certainty 5/5 debate potential 3/5

Chen: Public AI Benchmarks Are Unreliable and Often Contain Wrong Answers

Edwin Chen · The $1B Al company training ChatGPT, Claude & Gemini on the path to responsible AGI | Edwin Chen · Dec 7, 2025 · at 18:01

Edwin Chen, founder and CEO of Surge AI, explains to Lenny Rachitsky why standard AI benchmarks do not accurately reflect real-world AI model capability.

0:00 / 0:14exact quote · 14.0s
▶ Watch the full episode on YouTube → 720p mp4 · rendered on demand · StarZero watermark
“I don't trust the benchmarks at all. And I think that's for two reasons. So one is, I think a lot of people don't realize, even researchers within the community, they don't realize that the benchmarks themselves are often honestly just wrong. Like they have wrong answers.”

quote is from the automated transcript, cleaned for reading: filler sounds and stutters are removed, nothing is rephrased. names can be misheard (the analysis reads context, assessments check outside sources). how →

More from Edwin Chen

Opinion
Chen: Big Tech Could Fire 90% of Staff and Move Faster
“Like I used to work at a bunch of the big tech companies, and I always felt that we could fire 90% of people and we would move faster because the best people wouldn't have all these distractions.”
Edwin Chen Dec 7, 2025 ▶ 5:57 The $1B Al company training ChatGPT, Claude & Gemini on the path to responsible AGI | Edwin Chen
Assertion Supported
Surge AI Surpassed $1B in Revenue With Under 100 Employees
“Yeah, so we hit over a billion of revenue last year with under a hundred people.”
Edwin Chen Dec 7, 2025 ▶ 5:40 The $1B Al company training ChatGPT, Claude & Gemini on the path to responsible AGI | Edwin Chen
Prediction Open · timeframe Dec 2028
Chen: AI Efficiency Will Enable $100 Billion Revenue-Per-Employee Ratios
“And I think we're going to see companies with even crazier ratios, like a hundred billion per employee in the next few years. AI is just going to get better and better and make things more efficient. So that ratio just becomes inevitable.”
Edwin Chen Dec 7, 2025 ▶ 5:44 The $1B Al company training ChatGPT, Claude & Gemini on the path to responsible AGI | Edwin Chen
Assertion Not checkable as stated
Chen: Frontier AI Labs Game Benchmarks via Prompt Tweaking and Test Leaks
“Sometimes, yeah, these benchmarks, they accidentally leak in certain ways, or the frontier labs will tweak the way they evaluate their models on these benchmarks. Like they'll tweak their system prompt. Or they'll tweak the number of times they run their model…”
Edwin Chen Dec 7, 2025 ▶ 19:32 The $1B Al company training ChatGPT, Claude & Gemini on the path to responsible AGI | Edwin Chen
Assertion Not checkable as stated
Chen: LMSYS Chatbot Arena rewards bolding, emojis, and length over accuracy
“The easiest way to climb Alamarina, it's adding crazy boating. It's doubling the number of emojis. It's tripling the length of your model responses. Even if your model starts hallucinating and getting the answer completely wrong.”
Edwin Chen Dec 7, 2025 ▶ 24:15 The $1B Al company training ChatGPT, Claude & Gemini on the path to responsible AGI | Edwin Chen
Opinion
Chen: Vibe coding is overhyped and will make codebases unmaintainable
“I definitely think that Vibe coding is overhyped. I think people don't realize, How much it's going to make your systems unmaintainable in the long term and decently dump this code into your code bases.”
Edwin Chen Dec 7, 2025 ▶ 51:54 The $1B Al company training ChatGPT, Claude & Gemini on the path to responsible AGI | Edwin Chen
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.