Jul 24, 2025 · 32m · no-priors

No Priors Ep. 124 | With SurgeAI Founder and CEO Edwin Chen

Edwin Chen · 22m spoken Sarah Guo · 4m spoken Elad Gil · 3m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of No Priors, Surge AI founder and CEO Edwin Chen discusses how his bootstrapped data company achieved over one billion dollars in revenue by supplying critical human data to frontier AI laboratories. Chen shares key insights on scalable oversight, the limitations of synthetic data, the flaws of benchmark hacking, and the future of specialized foundation models.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 24.4% of the talking time here. How this is scored →

The hosts as informed peer 4.1 Guest teaching 4.6 Guest disagreement 3.2 The hosts pushing back 2.1
05100:0010:0020:0030:000:40–4:53 · The hosts as informed peer 5/10 Surge's Founding Thesis and Origin Story Edwin launches a fierce critique against Silicon Valley founders raising capital for vanity rather than solving problems. Elad pushes back with investor nuance, noting that while SV over-funds, companies outside SV often suffer from under-funding.4:53–7:34 · The hosts as informed peer 5/10 Early Team Building and Hiring Mistakes Sarah and Elad press Edwin on the practical hiring challenges unknown founders face without venture backing. Edwin strongly dismisses early PM and data science hires, arguing founders should stay hands-on.7:35–9:39 · The hosts as informed peer 3/10 Surge's Core Product Offerings and Deliverables Sarah intervenes to anchor the conversation on Surge's actual products and revenue breakdown. Edwin systematically explains SFT data, verifiers, preference data, and failure analysis.9:40–12:24 · The hosts as informed peer 5/10 Differentiating Surge from Body Shops via Technology Edwin disparages legacy data vendors as mere body shops with low quality ceilings. Elad asks a technical question regarding how Surge scales evaluation loops without running out of human evaluators.12:25–17:29 · The hosts as informed peer 4/10 Scalable Oversight and Human-AI Collaboration Sarah asks about RL environments and scalable oversight interfaces. Edwin provides concrete examples of multi-tool enterprise environments simulating real-world salesperson workflows.17:30–21:26 · The hosts as informed peer 7/10 Synthetic Data Limits, Superhuman AI, and Alignment Traps Edwin critiques the synthetic data bubble and labels LMSYS arena a plague on AI due to superficial vibe evaluations. Elad demonstrates deep scientific expertise by connecting model objective collapse to protein evolution selecting odd local maxima.21:27–23:36 · The hosts as informed peer 3/10 Human Evaluation as the True Gold Standard Sarah pushes for realistic alternatives to public benchmarks. Edwin insists rigorous, fact-checked human evaluation remains the only reliable gold standard against clickbait training.23:37–26:25 · The hosts as informed peer 3/10 Competitive Landscape and the Meta-Scale Deal Sarah asks about the competitive fallout of the Meta-Scale deal and future model specialization. Edwin asserts that low-quality vendors burned customers and predicts specialized frontier models rather than single-commodity AGI.26:26–29:29 · The hosts as informed peer 3/10 Benchmark Hacking and Future Public Research Edwin reveals that frontier lab researchers deliberately make models worse to climb LMSYS and IFEval leaderboards by inflating response lengths and emoji counts rather than building real capabilities.29:29–32:26 · The hosts as informed peer 3/10 Defining High-Quality Data Beyond Checkboxes Elad asks Edwin to define true data quality. Edwin rejects checkbox rubrics, Craigslist hires, and English PhD credentials, explaining why first-principles curation is necessary to prevent scaling mediocrity.0:40–4:53 · Guest teaching 2/10 Surge's Founding Thesis and Origin Story Edwin launches a fierce critique against Silicon Valley founders raising capital for vanity rather than solving problems. Elad pushes back with investor nuance, noting that while SV over-funds, companies outside SV often suffer from under-funding.4:53–7:34 · Guest teaching 4/10 Early Team Building and Hiring Mistakes Sarah and Elad press Edwin on the practical hiring challenges unknown founders face without venture backing. Edwin strongly dismisses early PM and data science hires, arguing founders should stay hands-on.7:35–9:39 · Guest teaching 4/10 Surge's Core Product Offerings and Deliverables Sarah intervenes to anchor the conversation on Surge's actual products and revenue breakdown. Edwin systematically explains SFT data, verifiers, preference data, and failure analysis.9:40–12:24 · Guest teaching 5/10 Differentiating Surge from Body Shops via Technology Edwin disparages legacy data vendors as mere body shops with low quality ceilings. Elad asks a technical question regarding how Surge scales evaluation loops without running out of human evaluators.12:25–17:29 · Guest teaching 5/10 Scalable Oversight and Human-AI Collaboration Sarah asks about RL environments and scalable oversight interfaces. Edwin provides concrete examples of multi-tool enterprise environments simulating real-world salesperson workflows.17:30–21:26 · Guest teaching 6/10 Synthetic Data Limits, Superhuman AI, and Alignment Traps Edwin critiques the synthetic data bubble and labels LMSYS arena a plague on AI due to superficial vibe evaluations. Elad demonstrates deep scientific expertise by connecting model objective collapse to protein evolution selecting odd local maxima.21:27–23:36 · Guest teaching 4/10 Human Evaluation as the True Gold Standard Sarah pushes for realistic alternatives to public benchmarks. Edwin insists rigorous, fact-checked human evaluation remains the only reliable gold standard against clickbait training.23:37–26:25 · Guest teaching 4/10 Competitive Landscape and the Meta-Scale Deal Sarah asks about the competitive fallout of the Meta-Scale deal and future model specialization. Edwin asserts that low-quality vendors burned customers and predicts specialized frontier models rather than single-commodity AGI.26:26–29:29 · Guest teaching 6/10 Benchmark Hacking and Future Public Research Edwin reveals that frontier lab researchers deliberately make models worse to climb LMSYS and IFEval leaderboards by inflating response lengths and emoji counts rather than building real capabilities.29:29–32:26 · Guest teaching 6/10 Defining High-Quality Data Beyond Checkboxes Elad asks Edwin to define true data quality. Edwin rejects checkbox rubrics, Craigslist hires, and English PhD credentials, explaining why first-principles curation is necessary to prevent scaling mediocrity.0:40–4:53 · Guest disagreement 4/10 Surge's Founding Thesis and Origin Story Edwin launches a fierce critique against Silicon Valley founders raising capital for vanity rather than solving problems. Elad pushes back with investor nuance, noting that while SV over-funds, companies outside SV often suffer from under-funding.4:53–7:34 · Guest disagreement 4/10 Early Team Building and Hiring Mistakes Sarah and Elad press Edwin on the practical hiring challenges unknown founders face without venture backing. Edwin strongly dismisses early PM and data science hires, arguing founders should stay hands-on.7:35–9:39 · Guest disagreement 1/10 Surge's Core Product Offerings and Deliverables Sarah intervenes to anchor the conversation on Surge's actual products and revenue breakdown. Edwin systematically explains SFT data, verifiers, preference data, and failure analysis.9:40–12:24 · Guest disagreement 3/10 Differentiating Surge from Body Shops via Technology Edwin disparages legacy data vendors as mere body shops with low quality ceilings. Elad asks a technical question regarding how Surge scales evaluation loops without running out of human evaluators.12:25–17:29 · Guest disagreement 1/10 Scalable Oversight and Human-AI Collaboration Sarah asks about RL environments and scalable oversight interfaces. Edwin provides concrete examples of multi-tool enterprise environments simulating real-world salesperson workflows.17:30–21:26 · Guest disagreement 4/10 Synthetic Data Limits, Superhuman AI, and Alignment Traps Edwin critiques the synthetic data bubble and labels LMSYS arena a plague on AI due to superficial vibe evaluations. Elad demonstrates deep scientific expertise by connecting model objective collapse to protein evolution selecting odd local maxima.21:27–23:36 · Guest disagreement 3/10 Human Evaluation as the True Gold Standard Sarah pushes for realistic alternatives to public benchmarks. Edwin insists rigorous, fact-checked human evaluation remains the only reliable gold standard against clickbait training.23:37–26:25 · Guest disagreement 3/10 Competitive Landscape and the Meta-Scale Deal Sarah asks about the competitive fallout of the Meta-Scale deal and future model specialization. Edwin asserts that low-quality vendors burned customers and predicts specialized frontier models rather than single-commodity AGI.26:26–29:29 · Guest disagreement 5/10 Benchmark Hacking and Future Public Research Edwin reveals that frontier lab researchers deliberately make models worse to climb LMSYS and IFEval leaderboards by inflating response lengths and emoji counts rather than building real capabilities.29:29–32:26 · Guest disagreement 4/10 Defining High-Quality Data Beyond Checkboxes Elad asks Edwin to define true data quality. Edwin rejects checkbox rubrics, Craigslist hires, and English PhD credentials, explaining why first-principles curation is necessary to prevent scaling mediocrity.0:40–4:53 · The hosts pushing back 4/10 Surge's Founding Thesis and Origin Story Edwin launches a fierce critique against Silicon Valley founders raising capital for vanity rather than solving problems. Elad pushes back with investor nuance, noting that while SV over-funds, companies outside SV often suffer from under-funding.4:53–7:34 · The hosts pushing back 4/10 Early Team Building and Hiring Mistakes Sarah and Elad press Edwin on the practical hiring challenges unknown founders face without venture backing. Edwin strongly dismisses early PM and data science hires, arguing founders should stay hands-on.7:35–9:39 · The hosts pushing back 2/10 Surge's Core Product Offerings and Deliverables Sarah intervenes to anchor the conversation on Surge's actual products and revenue breakdown. Edwin systematically explains SFT data, verifiers, preference data, and failure analysis.9:40–12:24 · The hosts pushing back 2/10 Differentiating Surge from Body Shops via Technology Edwin disparages legacy data vendors as mere body shops with low quality ceilings. Elad asks a technical question regarding how Surge scales evaluation loops without running out of human evaluators.12:25–17:29 · The hosts pushing back 1/10 Scalable Oversight and Human-AI Collaboration Sarah asks about RL environments and scalable oversight interfaces. Edwin provides concrete examples of multi-tool enterprise environments simulating real-world salesperson workflows.17:30–21:26 · The hosts pushing back 2/10 Synthetic Data Limits, Superhuman AI, and Alignment Traps Edwin critiques the synthetic data bubble and labels LMSYS arena a plague on AI due to superficial vibe evaluations. Elad demonstrates deep scientific expertise by connecting model objective collapse to protein evolution selecting odd local maxima.21:27–23:36 · The hosts pushing back 2/10 Human Evaluation as the True Gold Standard Sarah pushes for realistic alternatives to public benchmarks. Edwin insists rigorous, fact-checked human evaluation remains the only reliable gold standard against clickbait training.23:37–26:25 · The hosts pushing back 1/10 Competitive Landscape and the Meta-Scale Deal Sarah asks about the competitive fallout of the Meta-Scale deal and future model specialization. Edwin asserts that low-quality vendors burned customers and predicts specialized frontier models rather than single-commodity AGI.26:26–29:29 · The hosts pushing back 1/10 Benchmark Hacking and Future Public Research Edwin reveals that frontier lab researchers deliberately make models worse to climb LMSYS and IFEval leaderboards by inflating response lengths and emoji counts rather than building real capabilities.29:29–32:26 · The hosts pushing back 2/10 Defining High-Quality Data Beyond Checkboxes Elad asks Edwin to define true data quality. Edwin rejects checkbox rubrics, Craigslist hires, and English PhD credentials, explaining why first-principles curation is necessary to prevent scaling mediocrity.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 40% · guest 60%0:00 · the hosts 40% · guest 60%3:00 · the hosts 40.3% · guest 59.7%3:00 · the hosts 40.3% · guest 59.7%6:00 · the hosts 22% · guest 78%6:00 · the hosts 22% · guest 78%9:00 · the hosts 16.4% · guest 83.6%9:00 · the hosts 16.4% · guest 83.6%12:00 · the hosts 27.2% · guest 72.8%12:00 · the hosts 27.2% · guest 72.8%15:00 · the hosts 30.7% · guest 69.3%15:00 · the hosts 30.7% · guest 69.3%18:00 · the hosts 6.8% · guest 93.2%18:00 · the hosts 6.8% · guest 93.2%21:00 · the hosts 44.7% · guest 55.3%21:00 · the hosts 44.7% · guest 55.3%24:00 · the hosts 16% · guest 84%24:00 · the hosts 16% · guest 84%27:00 · the hosts 5.1% · guest 94.9%27:00 · the hosts 5.1% · guest 94.9%30:00 · the hosts 19.4% · guest 80.6%30:00 · the hosts 19.4% · guest 80.6%
Sharpest disagreement ▶ 26:48 Exposing metric gaming in frontier labs

Edwin aggressively calls out frontier lab leadership and researchers for intentionally trading model factuality for superficial leaderboard gains on LMSYS.

Hardest push from the hosts ▶ 4:27 Elad counters Edwin's anti-fundraising stance

Elad directly pushes back on Edwin's broad rejection of venture capital by highlighting that founders outside Silicon Valley suffer from severe under-capitalization.

Biggest teaching moment ▶ 29:38 Deconstructing naive definitions of data quality

Edwin breaks down why standard industry practices like checkbox counting and hiring PhDs produce terrible training data, reframing quality from first principles.

The host holds their own ▶ 20:47 Elad analogizes model training to directed evolution

Elad brings technical depth to the discussion by drawing a direct parallel between reinforcement learning objective drift and unexpected artifacts in protein evolution.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Surge's Founding Thesis and Origin Story 5244 Edwin launches a fierce critique against Silicon Valley founders raising capital for vanity rather than solving problems. Elad pushes back with investor nuance, noting that while SV over-funds, companies outside SV often suffer from under-funding.
Early Team Building and Hiring Mistakes 5444 Sarah and Elad press Edwin on the practical hiring challenges unknown founders face without venture backing. Edwin strongly dismisses early PM and data science hires, arguing founders should stay hands-on.
Surge's Core Product Offerings and Deliverables 3412 Sarah intervenes to anchor the conversation on Surge's actual products and revenue breakdown. Edwin systematically explains SFT data, verifiers, preference data, and failure analysis.
Differentiating Surge from Body Shops via Technology 5532 Edwin disparages legacy data vendors as mere body shops with low quality ceilings. Elad asks a technical question regarding how Surge scales evaluation loops without running out of human evaluators.
Scalable Oversight and Human-AI Collaboration 4511 Sarah asks about RL environments and scalable oversight interfaces. Edwin provides concrete examples of multi-tool enterprise environments simulating real-world salesperson workflows.
Synthetic Data Limits, Superhuman AI, and Alignment Traps 7642 Edwin critiques the synthetic data bubble and labels LMSYS arena a plague on AI due to superficial vibe evaluations. Elad demonstrates deep scientific expertise by connecting model objective collapse to protein evolution selecting odd local maxima.
Human Evaluation as the True Gold Standard 3432 Sarah pushes for realistic alternatives to public benchmarks. Edwin insists rigorous, fact-checked human evaluation remains the only reliable gold standard against clickbait training.
Competitive Landscape and the Meta-Scale Deal 3431 Sarah asks about the competitive fallout of the Meta-Scale deal and future model specialization. Edwin asserts that low-quality vendors burned customers and predicts specialized frontier models rather than single-commodity AGI.
Benchmark Hacking and Future Public Research 3651 Edwin reveals that frontier lab researchers deliberately make models worse to climb LMSYS and IFEval leaderboards by inflating response lengths and emoji counts rather than building real capabilities.
Defining High-Quality Data Beyond Checkboxes 3642 Elad asks Edwin to define true data quality. Edwin rejects checkbox rubrics, Craigslist hires, and English PhD credentials, explaining why first-principles curation is necessary to prevent scaling mediocrity.

Statements from this episode (19)

Assertion Not checkable as stated
Chen: Surge AI Is Largest Human Data Player in AI
“We are kind of like the biggest human data player in this space.”
Edwin Chen Jul 24, 2025 ▶ 0:55
Assertion Not checkable as stated
Chen: Surge AI Operates With Just Over 100 Employees
“And we're about a hundred, a little over a hundred people.”
Edwin Chen Jul 24, 2025 ▶ 1:00
Disclosure
Chen: Surge AI Was Profitable From Day One, Avoiding Venture Capital
“I think we were very, very lucky to be profitable from the start. So we didn't need the money. It always felt weird to give up control.”
Edwin Chen Jul 24, 2025 ▶ 2:42
Opinion
Chen: Many YC Founders Only Care About Fundraising Brags and Headlines
“Like, if you talk to a bunch of YC founders or whoever it is, like, what is their goal? It really is to tell all their friends that they raised ten million dollars and show their parents they got a headline on TechCrunch. Like, that is their goal.”
Edwin Chen Jul 24, 2025 ▶ 3:04
Insight
Chen: Early Startups Should Not Hire Data Scientists as First Employees
“Like I would never hire data scientists when the first three people in a company. And I say that because I used to be a data scientist. Like data scientists are great when you want to optimize your product by two percent or five percent, but that's definitely …”
Edwin Chen Jul 24, 2025 ▶ 6:55
Insight
Chen: Early-stage startups should not hire product managers
“Product managers are great when your company gets big enough. But at the beginning, you should be thinking about yourself about what product you want to build. And your engineer should be hands-on. They should be having great ideas as well.”
Edwin Chen Jul 24, 2025 ▶ 7:14
Opinion
Chen: Most AI Data Competitors Are Non-Technical Body Shops
“A lot of other companies in this space, they are essentially just body shops. What they are delivering is not data. They are literally just delivering warm bodies to companies. And so what that means is like at the end of the day, they don't have any technolog…”
Edwin Chen Jul 24, 2025 ▶ 9:55
Insight
Chen: Generative AI Data Has an Almost Unlimited Quality Ceiling
“So there's a very, very low ceiling on the bar of quality, but then take something like writing poetry. Well, I suck at writing poetry. Hemingway's definitely want to write a much better poem than I am. Or imagine, I don't know, a VC pitch deck. You're going t…”
Edwin Chen Jul 24, 2025 ▶ 10:54
Insight
Chen: AI RL Environments Are Too Complex to Create Synthetically
“I think one of the things that people really underestimate is how it is, how complicated it is that you can't just synthetically generate it.”
Edwin Chen Jul 24, 2025 ▶ 14:20
Prediction Not checkable as stated
Chen: Synthetic RL Environments Alone Won't Meet Future Frontier AI Demand
“Like, I don't think our environments alone will suffice just because, I mean, it depends on how you think about our environments, but oftentimes these are very, very rich trajectories are very, very long. And so it's almost like inconceivable that a single rew…”
Edwin Chen Jul 24, 2025 ▶ 16:59
Disclosure
Chen: Surge AI Heavily Uses Synthetic Data to Supplement Human Labelers
“Like we use it like a ton ourselves in order to supplement what the humans do.”
Edwin Chen Jul 24, 2025 ▶ 18:05
Opinion
Chen: LMSYS Chatbot Arena is a giant plague on AI
“One of the things I think is a giant plague on AI is Elimsis, Elimarena.”
Edwin Chen Jul 24, 2025 ▶ 19:40
Disclosure
Chen: Surge AI Works With All Frontier Labs on Model Evaluation
“So internally we do a lot of work actually today with working with all the frontier labs to help them understand their models. So again, we're constantly evaluating them. We're constantly surfacing loss areas for them to improve on and so on and so on.”
Edwin Chen Jul 24, 2025 ▶ 23:02
Opinion
Chen: Meta's Deal With Scale AI Actually Benefited Surge AI
“It's been beneficial because yeah, there were still some legacy teams using scale. Like they just didn't know about us because we were still pretty under the radar.”
Edwin Chen Jul 24, 2025 ▶ 23:52
Opinion
Chen Bets on xAI to Catch Up With Top Frontier Labs
“So I would bet on XAI. I think they're just very hungry and mission-oriented in a way that gives them a lot of really unique advantages.”
Edwin Chen Jul 24, 2025 ▶ 24:43
Prediction Not checkable as stated
Chen: Frontier AI Models Will Proliferate Rather Than Commoditize
“So I actually see more and more frontier models opening up over time because I actually don't think that the models will be commodities.”
Edwin Chen Jul 24, 2025 ▶ 25:04
Assertion Not checkable as stated
Chen: Frontier AI Labs Are No Longer Publishing Research
“Like, I think it is really interesting in that a lot of the, like for obvious reasons, a lot of the frontier labs, they're just not publishing anymore.”
Edwin Chen Jul 24, 2025 ▶ 26:42
Assertion Not checkable as stated
Chen: Researchers Degrade AI Factuality Just to Boost LMSYS Rankings
“A lot of researchers, they'll tell us that their VPs make them focus on increasing their rank on LMSYS. And so I've had researchers explicitly tell me that they're okay with making their models worse. Add factuality works at following instructions as long as i…”
Edwin Chen Jul 24, 2025 ▶ 27:09
Assertion Not checkable as stated
Chen: The Easiest Way to Boost LMSYS Rank Is Longer Responses
“Like one of the things that we found is that the easiest way to improve your rank on album arena is the ability to make your model response longer.”
Edwin Chen Jul 24, 2025 ▶ 27:52
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 100 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.