Dec 7, 2025 · 1h 10m · lennys-podcast

The $1B Al company training ChatGPT, Claude & Gemini on the path to responsible AGI | Edwin Chen

Edwin Chen · 42m spoken Lenny Rachitsky · 19m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

Surge AI founder and CEO Edwin Chen discusses how his company bootstrapped to over one billion dollars in revenue by rethinking AI data quality, rejecting venture capital conventions, and prioritizing human taste over engagement-driven metrics.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Lenny holds 30.9% of the talking time here. How this is scored →

Lenny as informed peer 3.5 Guest teaching 5.0 Guest disagreement 3.1 Lenny pushing back 0.8
05100:0015:0030:0045:001:00:004:51–9:09 · Lenny as informed peer 3/10 Bootstrapping Surge AI to a Billion-Dollar Run Rate Lenny opens by admiring Surge AI's unprecedented revenue-to-employee ratio and asks how Edwin leveraged AI. Edwin explains his contrarian philosophy of firing 90% of big tech workers and avoiding the Silicon Valley PR hamster wheel.9:10–13:30 · Lenny as informed peer 4/10 Defining Data Quality and Annotation Methodologies Lenny asks how data quality is defined and measured. Edwin dismantles simplistic checkbox approaches, using poetry and Google search ranking analogies to explain how Surge captures nuanced qualitative signals.13:31–17:38 · Lenny as informed peer 4/10 The Art of Post-Training and Model Taste Lenny inquires why Claude was superior at writing and coding for so long. Edwin explains that post-training is an art involving human taste, aesthetic tradeoffs, and prioritizing real tasks over marketing benchmarks.17:38–23:02 · Lenny as informed peer 3/10 The Pitfalls of Benchmarks and Accurate Evals Lenny asks if academic benchmarks correlate with real-world AI capabilities. Edwin strongly dismisses standard benchmarks as flawed and easily gamed, noting models can win IMO math medals but struggle to parse basic PDFs.23:03–28:33 · Lenny as informed peer 4/10 The Risks of AI Slop and Engagement Optimization Edwin launches a fierce critique of LMSYS Chatbot Arena leaderboards and social-media-style engagement optimization, calling it tabloid baiting. Lenny steelmans commercial video models like Sora, prompting Edwin to reiterate the dangers of deceptive shortcuts.28:33–31:50 · Lenny as informed peer 5/10 Contrarian Startup Strategy and Resisting the Pivot Edwin passionately denounces the Silicon Valley playbook of rapid pivoting, blitzscaling, and prestige hiring. Lenny validates Edwin's perspective by sharing findings from his own research with VC Terence Rohan on generational companies.31:51–44:44 · Lenny as informed peer 4/10 Sponsor Break: Coda Following the Coda sponsor segment, Lenny brings up Richard Sutton's Bitter Lesson and RL post-training environments. Edwin explains how reinforcement learning environments simulate real multi-step tasks rather than isolated academic tests.44:46–48:20 · Lenny as informed peer 3/10 Surge AI's Research Lab Culture and Hiring Standards Lenny asks why Surge invests in internal research teams unlike most data vendors. Edwin explains his desire to advance the scientific frontier rather than focus solely on valuation, noting he would rather be mathematician Terence Tao than Warren Buffett.48:20–52:55 · Lenny as informed peer 4/10 Model Differentiation, Mini-Apps, and Vibe Coding Edwin shares how corporate values will create differentiated model personalities and highlights interactive chatbot mini-apps as underhyped while calling vibe coding overhyped. Lenny connects this to previous discussions with product leaders.53:01–57:52 · Lenny as informed peer 3/10 Origins of Surge AI and Edwin Chen's Career Lenny compares Edwin's interdisciplinary background to Coinbase founder Brian Armstrong. Edwin recounts his journey through MIT linguistics, big tech research frustration with simple image labeling, and launching Surge AI right after GPT-3.57:59–1:03:32 · Lenny as informed peer 2/10 The Philosophy of Training AI Like Raising a Child Edwin reflects philosophically on AI training, comparing it to raising a child and teaching values rather than optimizing for simplistic proxy metrics. Lenny expresses appreciation for Edwin's deep philosophical framing.1:03:33–1:08:40 · Lenny as informed peer 3/10 Lightning Round: Recommendations, Sci-Fi, and Soda vs. Pop In the lightning round, Edwin shares favorite sci-fi books about linguistics, praises Waymo, and jokes about his viral Twitter soda versus pop dialect map.4:51–9:09 · Guest teaching 5/10 Bootstrapping Surge AI to a Billion-Dollar Run Rate Lenny opens by admiring Surge AI's unprecedented revenue-to-employee ratio and asks how Edwin leveraged AI. Edwin explains his contrarian philosophy of firing 90% of big tech workers and avoiding the Silicon Valley PR hamster wheel.9:10–13:30 · Guest teaching 6/10 Defining Data Quality and Annotation Methodologies Lenny asks how data quality is defined and measured. Edwin dismantles simplistic checkbox approaches, using poetry and Google search ranking analogies to explain how Surge captures nuanced qualitative signals.13:31–17:38 · Guest teaching 5/10 The Art of Post-Training and Model Taste Lenny inquires why Claude was superior at writing and coding for so long. Edwin explains that post-training is an art involving human taste, aesthetic tradeoffs, and prioritizing real tasks over marketing benchmarks.17:38–23:02 · Guest teaching 7/10 The Pitfalls of Benchmarks and Accurate Evals Lenny asks if academic benchmarks correlate with real-world AI capabilities. Edwin strongly dismisses standard benchmarks as flawed and easily gamed, noting models can win IMO math medals but struggle to parse basic PDFs.23:03–28:33 · Guest teaching 6/10 The Risks of AI Slop and Engagement Optimization Edwin launches a fierce critique of LMSYS Chatbot Arena leaderboards and social-media-style engagement optimization, calling it tabloid baiting. Lenny steelmans commercial video models like Sora, prompting Edwin to reiterate the dangers of deceptive shortcuts.28:33–31:50 · Guest teaching 4/10 Contrarian Startup Strategy and Resisting the Pivot Edwin passionately denounces the Silicon Valley playbook of rapid pivoting, blitzscaling, and prestige hiring. Lenny validates Edwin's perspective by sharing findings from his own research with VC Terence Rohan on generational companies.31:51–44:44 · Guest teaching 6/10 Sponsor Break: Coda Following the Coda sponsor segment, Lenny brings up Richard Sutton's Bitter Lesson and RL post-training environments. Edwin explains how reinforcement learning environments simulate real multi-step tasks rather than isolated academic tests.44:46–48:20 · Guest teaching 5/10 Surge AI's Research Lab Culture and Hiring Standards Lenny asks why Surge invests in internal research teams unlike most data vendors. Edwin explains his desire to advance the scientific frontier rather than focus solely on valuation, noting he would rather be mathematician Terence Tao than Warren Buffett.48:20–52:55 · Guest teaching 4/10 Model Differentiation, Mini-Apps, and Vibe Coding Edwin shares how corporate values will create differentiated model personalities and highlights interactive chatbot mini-apps as underhyped while calling vibe coding overhyped. Lenny connects this to previous discussions with product leaders.53:01–57:52 · Guest teaching 4/10 Origins of Surge AI and Edwin Chen's Career Lenny compares Edwin's interdisciplinary background to Coinbase founder Brian Armstrong. Edwin recounts his journey through MIT linguistics, big tech research frustration with simple image labeling, and launching Surge AI right after GPT-3.57:59–1:03:32 · Guest teaching 5/10 The Philosophy of Training AI Like Raising a Child Edwin reflects philosophically on AI training, comparing it to raising a child and teaching values rather than optimizing for simplistic proxy metrics. Lenny expresses appreciation for Edwin's deep philosophical framing.1:03:33–1:08:40 · Guest teaching 3/10 Lightning Round: Recommendations, Sci-Fi, and Soda vs. Pop In the lightning round, Edwin shares favorite sci-fi books about linguistics, praises Waymo, and jokes about his viral Twitter soda versus pop dialect map.4:51–9:09 · Guest disagreement 4/10 Bootstrapping Surge AI to a Billion-Dollar Run Rate Lenny opens by admiring Surge AI's unprecedented revenue-to-employee ratio and asks how Edwin leveraged AI. Edwin explains his contrarian philosophy of firing 90% of big tech workers and avoiding the Silicon Valley PR hamster wheel.9:10–13:30 · Guest disagreement 3/10 Defining Data Quality and Annotation Methodologies Lenny asks how data quality is defined and measured. Edwin dismantles simplistic checkbox approaches, using poetry and Google search ranking analogies to explain how Surge captures nuanced qualitative signals.13:31–17:38 · Guest disagreement 2/10 The Art of Post-Training and Model Taste Lenny inquires why Claude was superior at writing and coding for so long. Edwin explains that post-training is an art involving human taste, aesthetic tradeoffs, and prioritizing real tasks over marketing benchmarks.17:38–23:02 · Guest disagreement 5/10 The Pitfalls of Benchmarks and Accurate Evals Lenny asks if academic benchmarks correlate with real-world AI capabilities. Edwin strongly dismisses standard benchmarks as flawed and easily gamed, noting models can win IMO math medals but struggle to parse basic PDFs.23:03–28:33 · Guest disagreement 6/10 The Risks of AI Slop and Engagement Optimization Edwin launches a fierce critique of LMSYS Chatbot Arena leaderboards and social-media-style engagement optimization, calling it tabloid baiting. Lenny steelmans commercial video models like Sora, prompting Edwin to reiterate the dangers of deceptive shortcuts.28:33–31:50 · Guest disagreement 5/10 Contrarian Startup Strategy and Resisting the Pivot Edwin passionately denounces the Silicon Valley playbook of rapid pivoting, blitzscaling, and prestige hiring. Lenny validates Edwin's perspective by sharing findings from his own research with VC Terence Rohan on generational companies.31:51–44:44 · Guest disagreement 2/10 Sponsor Break: Coda Following the Coda sponsor segment, Lenny brings up Richard Sutton's Bitter Lesson and RL post-training environments. Edwin explains how reinforcement learning environments simulate real multi-step tasks rather than isolated academic tests.44:46–48:20 · Guest disagreement 2/10 Surge AI's Research Lab Culture and Hiring Standards Lenny asks why Surge invests in internal research teams unlike most data vendors. Edwin explains his desire to advance the scientific frontier rather than focus solely on valuation, noting he would rather be mathematician Terence Tao than Warren Buffett.48:20–52:55 · Guest disagreement 3/10 Model Differentiation, Mini-Apps, and Vibe Coding Edwin shares how corporate values will create differentiated model personalities and highlights interactive chatbot mini-apps as underhyped while calling vibe coding overhyped. Lenny connects this to previous discussions with product leaders.53:01–57:52 · Guest disagreement 1/10 Origins of Surge AI and Edwin Chen's Career Lenny compares Edwin's interdisciplinary background to Coinbase founder Brian Armstrong. Edwin recounts his journey through MIT linguistics, big tech research frustration with simple image labeling, and launching Surge AI right after GPT-3.57:59–1:03:32 · Guest disagreement 3/10 The Philosophy of Training AI Like Raising a Child Edwin reflects philosophically on AI training, comparing it to raising a child and teaching values rather than optimizing for simplistic proxy metrics. Lenny expresses appreciation for Edwin's deep philosophical framing.1:03:33–1:08:40 · Guest disagreement 1/10 Lightning Round: Recommendations, Sci-Fi, and Soda vs. Pop In the lightning round, Edwin shares favorite sci-fi books about linguistics, praises Waymo, and jokes about his viral Twitter soda versus pop dialect map.4:51–9:09 · Lenny pushing back 1/10 Bootstrapping Surge AI to a Billion-Dollar Run Rate Lenny opens by admiring Surge AI's unprecedented revenue-to-employee ratio and asks how Edwin leveraged AI. Edwin explains his contrarian philosophy of firing 90% of big tech workers and avoiding the Silicon Valley PR hamster wheel.9:10–13:30 · Lenny pushing back 1/10 Defining Data Quality and Annotation Methodologies Lenny asks how data quality is defined and measured. Edwin dismantles simplistic checkbox approaches, using poetry and Google search ranking analogies to explain how Surge captures nuanced qualitative signals.13:31–17:38 · Lenny pushing back 0/10 The Art of Post-Training and Model Taste Lenny inquires why Claude was superior at writing and coding for so long. Edwin explains that post-training is an art involving human taste, aesthetic tradeoffs, and prioritizing real tasks over marketing benchmarks.17:38–23:02 · Lenny pushing back 1/10 The Pitfalls of Benchmarks and Accurate Evals Lenny asks if academic benchmarks correlate with real-world AI capabilities. Edwin strongly dismisses standard benchmarks as flawed and easily gamed, noting models can win IMO math medals but struggle to parse basic PDFs.23:03–28:33 · Lenny pushing back 3/10 The Risks of AI Slop and Engagement Optimization Edwin launches a fierce critique of LMSYS Chatbot Arena leaderboards and social-media-style engagement optimization, calling it tabloid baiting. Lenny steelmans commercial video models like Sora, prompting Edwin to reiterate the dangers of deceptive shortcuts.28:33–31:50 · Lenny pushing back 1/10 Contrarian Startup Strategy and Resisting the Pivot Edwin passionately denounces the Silicon Valley playbook of rapid pivoting, blitzscaling, and prestige hiring. Lenny validates Edwin's perspective by sharing findings from his own research with VC Terence Rohan on generational companies.31:51–44:44 · Lenny pushing back 1/10 Sponsor Break: Coda Following the Coda sponsor segment, Lenny brings up Richard Sutton's Bitter Lesson and RL post-training environments. Edwin explains how reinforcement learning environments simulate real multi-step tasks rather than isolated academic tests.44:46–48:20 · Lenny pushing back 0/10 Surge AI's Research Lab Culture and Hiring Standards Lenny asks why Surge invests in internal research teams unlike most data vendors. Edwin explains his desire to advance the scientific frontier rather than focus solely on valuation, noting he would rather be mathematician Terence Tao than Warren Buffett.48:20–52:55 · Lenny pushing back 1/10 Model Differentiation, Mini-Apps, and Vibe Coding Edwin shares how corporate values will create differentiated model personalities and highlights interactive chatbot mini-apps as underhyped while calling vibe coding overhyped. Lenny connects this to previous discussions with product leaders.53:01–57:52 · Lenny pushing back 0/10 Origins of Surge AI and Edwin Chen's Career Lenny compares Edwin's interdisciplinary background to Coinbase founder Brian Armstrong. Edwin recounts his journey through MIT linguistics, big tech research frustration with simple image labeling, and launching Surge AI right after GPT-3.57:59–1:03:32 · Lenny pushing back 0/10 The Philosophy of Training AI Like Raising a Child Edwin reflects philosophically on AI training, comparing it to raising a child and teaching values rather than optimizing for simplistic proxy metrics. Lenny expresses appreciation for Edwin's deep philosophical framing.1:03:33–1:08:40 · Lenny pushing back 0/10 Lightning Round: Recommendations, Sci-Fi, and Soda vs. Pop In the lightning round, Edwin shares favorite sci-fi books about linguistics, praises Waymo, and jokes about his viral Twitter soda versus pop dialect map.

speaking balance: gold is Lenny, purple is the guest (3 minute bins)

0:00 · Lenny 59.7% · guest 40.3%0:00 · Lenny 59.7% · guest 40.3%3:00 · Lenny 88.1% · guest 11.9%3:00 · Lenny 88.1% · guest 11.9%6:00 · Lenny 12.7% · guest 87.3%6:00 · Lenny 12.7% · guest 87.3%9:00 · Lenny 26.1% · guest 73.9%9:00 · Lenny 26.1% · guest 73.9%12:00 · Lenny 16.1% · guest 83.9%12:00 · Lenny 16.1% · guest 83.9%15:00 · Lenny 32.6% · guest 67.4%15:00 · Lenny 32.6% · guest 67.4%18:00 · Lenny 10.5% · guest 89.5%18:00 · Lenny 10.5% · guest 89.5%21:00 · Lenny 20.7% · guest 79.3%21:00 · Lenny 20.7% · guest 79.3%24:00 · Lenny 18% · guest 82%24:00 · Lenny 18% · guest 82%27:00 · Lenny 23.5% · guest 76.5%27:00 · Lenny 23.5% · guest 76.5%30:00 · Lenny 64.9% · guest 35.1%30:00 · Lenny 64.9% · guest 35.1%33:00 · Lenny 32.8% · guest 67.2%33:00 · Lenny 32.8% · guest 67.2%36:00 · Lenny 47.6% · guest 52.4%36:00 · Lenny 47.6% · guest 52.4%39:00 · Lenny 26.5% · guest 73.5%39:00 · Lenny 26.5% · guest 73.5%42:00 · Lenny 30.6% · guest 69.4%42:00 · Lenny 30.6% · guest 69.4%45:00 · Lenny 10.2% · guest 89.8%45:00 · Lenny 10.2% · guest 89.8%48:00 · Lenny 16.3% · guest 83.7%48:00 · Lenny 16.3% · guest 83.7%51:00 · Lenny 38.5% · guest 61.5%51:00 · Lenny 38.5% · guest 61.5%54:00 · Lenny 7.2% · guest 92.8%54:00 · Lenny 7.2% · guest 92.8%57:00 · Lenny 17.5% · guest 82.5%57:00 · Lenny 17.5% · guest 82.5%1:00:00 · Lenny 29% · guest 71%1:00:00 · Lenny 29% · guest 71%1:03:00 · Lenny 30.7% · guest 69.3%1:03:00 · Lenny 30.7% · guest 69.3%1:06:00 · Lenny 47.7% · guest 52.3%1:06:00 · Lenny 47.7% · guest 52.3%1:09:00 · Lenny 40% · guest 60%1:09:00 · Lenny 40% · guest 60%
Sharpest disagreement ▶ 23:30 Denouncing LM Arena as tabloid optimization

Edwin forcefully attacks standard industry leaderboards like LM Arena, calling them superficial dopamine-chasers equivalent to grocery store tabloids.

Hardest push from Lenny ▶ 27:37 Lenny steelmanning entertainment models like Sora

Lenny counters Edwin's skepticism of frivolous models by steelmanning OpenAI Sora as enjoyable, revenue-generating, and valuable for video training.

Biggest teaching moment ▶ 9:59 Explaining the difference between metric-checking and real quality

Edwin completely reframes data annotation from counting lines in a poem to evaluating Nobel-prize-caliber artistic depth and emotion.

Lenny holds their own ▶ 30:52 Lenny citing his research on early employees at generational startups

Lenny demonstrates his domain expertise by validating Edwin's anti-pivot philosophy using qualitative data from his study on early employees at OpenAI and Stripe.

the scores for every segment, with the reasoning behind each
ChapterTopicLenny as informed peerGuest teachingGuest disagreementLenny pushing backWhy
Bootstrapping Surge AI to a Billion-Dollar Run Rate 3541 Lenny opens by admiring Surge AI's unprecedented revenue-to-employee ratio and asks how Edwin leveraged AI. Edwin explains his contrarian philosophy of firing 90% of big tech workers and avoiding the Silicon Valley PR hamster wheel.
Defining Data Quality and Annotation Methodologies 4631 Lenny asks how data quality is defined and measured. Edwin dismantles simplistic checkbox approaches, using poetry and Google search ranking analogies to explain how Surge captures nuanced qualitative signals.
The Art of Post-Training and Model Taste 4520 Lenny inquires why Claude was superior at writing and coding for so long. Edwin explains that post-training is an art involving human taste, aesthetic tradeoffs, and prioritizing real tasks over marketing benchmarks.
The Pitfalls of Benchmarks and Accurate Evals 3751 Lenny asks if academic benchmarks correlate with real-world AI capabilities. Edwin strongly dismisses standard benchmarks as flawed and easily gamed, noting models can win IMO math medals but struggle to parse basic PDFs.
The Risks of AI Slop and Engagement Optimization 4663 Edwin launches a fierce critique of LMSYS Chatbot Arena leaderboards and social-media-style engagement optimization, calling it tabloid baiting. Lenny steelmans commercial video models like Sora, prompting Edwin to reiterate the dangers of deceptive shortcuts.
Contrarian Startup Strategy and Resisting the Pivot 5451 Edwin passionately denounces the Silicon Valley playbook of rapid pivoting, blitzscaling, and prestige hiring. Lenny validates Edwin's perspective by sharing findings from his own research with VC Terence Rohan on generational companies.
Sponsor Break: Coda 4621 Following the Coda sponsor segment, Lenny brings up Richard Sutton's Bitter Lesson and RL post-training environments. Edwin explains how reinforcement learning environments simulate real multi-step tasks rather than isolated academic tests.
Surge AI's Research Lab Culture and Hiring Standards 3520 Lenny asks why Surge invests in internal research teams unlike most data vendors. Edwin explains his desire to advance the scientific frontier rather than focus solely on valuation, noting he would rather be mathematician Terence Tao than Warren Buffett.
Model Differentiation, Mini-Apps, and Vibe Coding 4431 Edwin shares how corporate values will create differentiated model personalities and highlights interactive chatbot mini-apps as underhyped while calling vibe coding overhyped. Lenny connects this to previous discussions with product leaders.
Origins of Surge AI and Edwin Chen's Career 3410 Lenny compares Edwin's interdisciplinary background to Coinbase founder Brian Armstrong. Edwin recounts his journey through MIT linguistics, big tech research frustration with simple image labeling, and launching Surge AI right after GPT-3.
The Philosophy of Training AI Like Raising a Child 2530 Edwin reflects philosophically on AI training, comparing it to raising a child and teaching values rather than optimizing for simplistic proxy metrics. Lenny expresses appreciation for Edwin's deep philosophical framing.
Lightning Round: Recommendations, Sci-Fi, and Soda vs. Pop 3310 In the lightning round, Edwin shares favorite sci-fi books about linguistics, praises Waymo, and jokes about his viral Twitter soda versus pop dialect map.

Statements from this episode (15)

Assertion Supported
Surge AI Surpassed $1B in Revenue With Under 100 Employees
“Yeah, so we hit over a billion of revenue last year with under a hundred people.”
Edwin Chen Dec 7, 2025 ▶ 5:40
Prediction Open · timeframe Dec 2028
Chen: AI Efficiency Will Enable $100 Billion Revenue-Per-Employee Ratios
“And I think we're going to see companies with even crazier ratios, like a hundred billion per employee in the next few years. AI is just going to get better and better and make things more efficient. So that ratio just becomes inevitable.”
Edwin Chen Dec 7, 2025 ▶ 5:44
Opinion
Chen: Big Tech Could Fire 90% of Staff and Move Faster
“Like I used to work at a bunch of the big tech companies, and I always felt that we could fire 90% of people and we would move faster because the best people wouldn't have all these distractions.”
Edwin Chen Dec 7, 2025 ▶ 5:57
Opinion
Chen: Throwing bodies at data labeling fails to create quality AI data
“I think most people don't understand what quality even means in this space. They think you can just throw bodies at a problem and get good data, and that's completely wrong.”
Edwin Chen Dec 7, 2025 ▶ 9:48
Disclosure
Surge AI tracks keystrokes and model gains to evaluate data annotators
“The way it works is we essentially gather thousands of signals about everything that you're doing when you're working on a platform. So we are looking at your keyboard strokes. We are looking how fast you answer things. We are using reviews. We are using code …”
Edwin Chen Dec 7, 2025 ▶ 11:57
Insight
Chen: AI post-training is an art driven by taste, not pure science
“One of the things I often think about is that there's a, it's almost like there's an art to post training. It's not purely a science. Like when you were deciding what kind of model you're trying to create and what it's good at. There's this notion of taste and…”
Edwin Chen Dec 7, 2025 ▶ 15:35
Opinion
Chen: Public AI Benchmarks Are Unreliable and Often Contain Wrong Answers
“I don't trust the benchmarks at all. And I think that's for two reasons. So one is, I think a lot of people don't realize, even researchers within the community, they don't realize that the benchmarks themselves are often honestly just wrong. Like they have wr…”
Edwin Chen Dec 7, 2025 ▶ 18:01
Assertion Not checkable as stated
Chen: Frontier AI Labs Game Benchmarks via Prompt Tweaking and Test Leaks
“Sometimes, yeah, these benchmarks, they accidentally leak in certain ways, or the frontier labs will tweak the way they evaluate their models on these benchmarks. Like they'll tweak their system prompt. Or they'll tweak the number of times they run their model…”
Edwin Chen Dec 7, 2025 ▶ 19:32
Prediction Not checkable as stated
Chen: AI Will Automate 80% of L6 Engineer Tasks Within 2 Years
“In my head, I probably bet that within the next one or two years, yeah, the models are going to automate 80% of, you know, the average L six software engineer's job. But it's going to take another few years, do you move to 90%, and another few years to 99%, an…”
Edwin Chen Dec 7, 2025 ▶ 22:28
Assertion Not checkable as stated
Chen: LMSYS Chatbot Arena rewards bolding, emojis, and length over accuracy
“The easiest way to climb Alamarina, it's adding crazy boating. It's doubling the number of emojis. It's tripling the length of your model responses. Even if your model starts hallucinating and getting the answer completely wrong.”
Edwin Chen Dec 7, 2025 ▶ 24:15
Opinion
Chen: Anthropic stands out among frontier labs for principled model development
“I would say I've always been very, very impressed by Anthropic. Like, I think Anthropic takes a very principled view about what they do and don't care about. And how they want their models to behave in a way that feels a lot more principle to me.”
Edwin Chen Dec 7, 2025 ▶ 26:21
Opinion
Chen: Silicon Valley is just as money-obsessed as Wall Street
“Silicon Valley loves to score on Wall Street for focusing on money. But honestly, most of the Silicon Valley is chasing the same thing.”
Edwin Chen Dec 7, 2025 ▶ 29:47
Prediction Not checkable as stated
Chen: Achieving AGI will require breakthroughs beyond standard LLMs
“I'm in a camp where I do believe that something new will be needed.”
Edwin Chen Dec 7, 2025 ▶ 33:43
Prediction Not checkable as stated
Chen predicts AI models will become increasingly differentiated across creator labs
“I think one of the things that's going to happen in the next few years is that the models are actually going to become increasingly differentiated because of the personalities and behaviors That the different labs have and the kind of objective functions that …”
Edwin Chen Dec 7, 2025 ▶ 48:20
Opinion
Chen: Vibe coding is overhyped and will make codebases unmaintainable
“I definitely think that Vibe coding is overhyped. I think people don't realize, How much it's going to make your systems unmaintainable in the long term and decently dump this code into your code bases.”
Edwin Chen Dec 7, 2025 ▶ 51:54
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 300 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.