Dec 7, 2025 · 34m · latent-space

The Great Evals Debate — Ankur Goyal & Malte Ubl

Ankur Goyal · 15m spoken Malte Ubl · 9m spoken Shawn Wang · 7m spoken
0:00 / 0:00
▶ Watch on YouTube →

gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions

In this episode of the Latent Space Lightning Pod, Swyx hosts Braintrust founder Ankur Goyal and Vercel CTO Malte Ubl to debate the transition from 'vibe coding' to robust evaluation methodologies for AI coding agents. They delve into production-driven feedback loops, reinforcement learning pipelines, and the strategic role of open-source framework benchmarks.

How this conversation actually went

Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 23.2% of the talking time here. How this is scored →

The hosts as informed peer 4.6 Guest teaching 4.4 Guest disagreement 3.3 The hosts pushing back 3.7
05100:0010:0020:0030:001:04–5:21 · The hosts as informed peer 2/10 Defining Evals, Vibes, and Feedback Loops in AI Swyx sets the stage by bringing up Boris Cherny's vibe coding comments, after which Malte and Ankur give structured opening philosophies on feedback loops and why vibe checks scale poorly compared to automated evals. The dynamic is collaborative with the host mostly listening.5:22–8:33 · The hosts as informed peer 5/10 RL Loops, AI Lab Privilege, and Public Benchmarks Malte notes that frontier agent teams benefit from unfair lab proximity. Swyx counters with real-world agent bench burnout at Cognition, prompting Ankur to clearly delineate public marketing benchmarks from internal iteration feedback loops.8:33–14:33 · The hosts as informed peer 4/10 Offline Evals as Production Replay vs Brittle Unit Tests Swyx challenges offline evals by claiming open-ended coding outgrows static test suites quickly. Ankur directly pushes back, telling Swyx he is conflating offline running with static golden datasets and explaining modern production replay workflows.14:36–18:11 · The hosts as informed peer 5/10 Vercel's Practical Eval Workflow and Composite Models Malte outlines Vercel's composite model approach to fix trivial syntax errors fast. Swyx tries to summarize it as avoiding agentic loops entirely, which Malte immediately corrects before Swyx cites RL turn optimization from Cognition's SWEGREP work.18:11–22:48 · The hosts as informed peer 5/10 Evals as Product Specs and Domain Knowledge Transfer Ankur describes how evals replace 50-page PRDs for product managers. When Ankur suggests vibes are just another form of eval, Swyx pushes back against diluting the term, arguing that calling everything an eval makes the definition meaningless.22:48–26:10 · The hosts as informed peer 6/10 Reinforcement Learning Environments and Synthetic Feedback Ankur expresses skepticism that average companies can build RL environments without reward hacking. Swyx jumps in with a strong conceptual counterpoint, arguing RL environments scale because they encode domain rules to decouple learning from slow human feedback.26:10–32:34 · The hosts as informed peer 5/10 Inversion of Control: Publishing Evals for Model Labs Malte proposes an inversion of control where framework authors publish evals to force model labs to optimize for them. When Swyx compares this to independent rating agencies like LMSYS, Malte clarifies the self-interested incentive model behind Vercel's strategy.1:04–5:21 · Guest teaching 3/10 Defining Evals, Vibes, and Feedback Loops in AI Swyx sets the stage by bringing up Boris Cherny's vibe coding comments, after which Malte and Ankur give structured opening philosophies on feedback loops and why vibe checks scale poorly compared to automated evals. The dynamic is collaborative with the host mostly listening.5:22–8:33 · Guest teaching 4/10 RL Loops, AI Lab Privilege, and Public Benchmarks Malte notes that frontier agent teams benefit from unfair lab proximity. Swyx counters with real-world agent bench burnout at Cognition, prompting Ankur to clearly delineate public marketing benchmarks from internal iteration feedback loops.8:33–14:33 · Guest teaching 7/10 Offline Evals as Production Replay vs Brittle Unit Tests Swyx challenges offline evals by claiming open-ended coding outgrows static test suites quickly. Ankur directly pushes back, telling Swyx he is conflating offline running with static golden datasets and explaining modern production replay workflows.14:36–18:11 · Guest teaching 5/10 Vercel's Practical Eval Workflow and Composite Models Malte outlines Vercel's composite model approach to fix trivial syntax errors fast. Swyx tries to summarize it as avoiding agentic loops entirely, which Malte immediately corrects before Swyx cites RL turn optimization from Cognition's SWEGREP work.18:11–22:48 · Guest teaching 4/10 Evals as Product Specs and Domain Knowledge Transfer Ankur describes how evals replace 50-page PRDs for product managers. When Ankur suggests vibes are just another form of eval, Swyx pushes back against diluting the term, arguing that calling everything an eval makes the definition meaningless.22:48–26:10 · Guest teaching 3/10 Reinforcement Learning Environments and Synthetic Feedback Ankur expresses skepticism that average companies can build RL environments without reward hacking. Swyx jumps in with a strong conceptual counterpoint, arguing RL environments scale because they encode domain rules to decouple learning from slow human feedback.26:10–32:34 · Guest teaching 5/10 Inversion of Control: Publishing Evals for Model Labs Malte proposes an inversion of control where framework authors publish evals to force model labs to optimize for them. When Swyx compares this to independent rating agencies like LMSYS, Malte clarifies the self-interested incentive model behind Vercel's strategy.1:04–5:21 · Guest disagreement 1/10 Defining Evals, Vibes, and Feedback Loops in AI Swyx sets the stage by bringing up Boris Cherny's vibe coding comments, after which Malte and Ankur give structured opening philosophies on feedback loops and why vibe checks scale poorly compared to automated evals. The dynamic is collaborative with the host mostly listening.5:22–8:33 · Guest disagreement 3/10 RL Loops, AI Lab Privilege, and Public Benchmarks Malte notes that frontier agent teams benefit from unfair lab proximity. Swyx counters with real-world agent bench burnout at Cognition, prompting Ankur to clearly delineate public marketing benchmarks from internal iteration feedback loops.8:33–14:33 · Guest disagreement 5/10 Offline Evals as Production Replay vs Brittle Unit Tests Swyx challenges offline evals by claiming open-ended coding outgrows static test suites quickly. Ankur directly pushes back, telling Swyx he is conflating offline running with static golden datasets and explaining modern production replay workflows.14:36–18:11 · Guest disagreement 4/10 Vercel's Practical Eval Workflow and Composite Models Malte outlines Vercel's composite model approach to fix trivial syntax errors fast. Swyx tries to summarize it as avoiding agentic loops entirely, which Malte immediately corrects before Swyx cites RL turn optimization from Cognition's SWEGREP work.18:11–22:48 · Guest disagreement 4/10 Evals as Product Specs and Domain Knowledge Transfer Ankur describes how evals replace 50-page PRDs for product managers. When Ankur suggests vibes are just another form of eval, Swyx pushes back against diluting the term, arguing that calling everything an eval makes the definition meaningless.22:48–26:10 · Guest disagreement 3/10 Reinforcement Learning Environments and Synthetic Feedback Ankur expresses skepticism that average companies can build RL environments without reward hacking. Swyx jumps in with a strong conceptual counterpoint, arguing RL environments scale because they encode domain rules to decouple learning from slow human feedback.26:10–32:34 · Guest disagreement 3/10 Inversion of Control: Publishing Evals for Model Labs Malte proposes an inversion of control where framework authors publish evals to force model labs to optimize for them. When Swyx compares this to independent rating agencies like LMSYS, Malte clarifies the self-interested incentive model behind Vercel's strategy.1:04–5:21 · The hosts pushing back 1/10 Defining Evals, Vibes, and Feedback Loops in AI Swyx sets the stage by bringing up Boris Cherny's vibe coding comments, after which Malte and Ankur give structured opening philosophies on feedback loops and why vibe checks scale poorly compared to automated evals. The dynamic is collaborative with the host mostly listening.5:22–8:33 · The hosts pushing back 3/10 RL Loops, AI Lab Privilege, and Public Benchmarks Malte notes that frontier agent teams benefit from unfair lab proximity. Swyx counters with real-world agent bench burnout at Cognition, prompting Ankur to clearly delineate public marketing benchmarks from internal iteration feedback loops.8:33–14:33 · The hosts pushing back 4/10 Offline Evals as Production Replay vs Brittle Unit Tests Swyx challenges offline evals by claiming open-ended coding outgrows static test suites quickly. Ankur directly pushes back, telling Swyx he is conflating offline running with static golden datasets and explaining modern production replay workflows.14:36–18:11 · The hosts pushing back 4/10 Vercel's Practical Eval Workflow and Composite Models Malte outlines Vercel's composite model approach to fix trivial syntax errors fast. Swyx tries to summarize it as avoiding agentic loops entirely, which Malte immediately corrects before Swyx cites RL turn optimization from Cognition's SWEGREP work.18:11–22:48 · The hosts pushing back 6/10 Evals as Product Specs and Domain Knowledge Transfer Ankur describes how evals replace 50-page PRDs for product managers. When Ankur suggests vibes are just another form of eval, Swyx pushes back against diluting the term, arguing that calling everything an eval makes the definition meaningless.22:48–26:10 · The hosts pushing back 5/10 Reinforcement Learning Environments and Synthetic Feedback Ankur expresses skepticism that average companies can build RL environments without reward hacking. Swyx jumps in with a strong conceptual counterpoint, arguing RL environments scale because they encode domain rules to decouple learning from slow human feedback.26:10–32:34 · The hosts pushing back 3/10 Inversion of Control: Publishing Evals for Model Labs Malte proposes an inversion of control where framework authors publish evals to force model labs to optimize for them. When Swyx compares this to independent rating agencies like LMSYS, Malte clarifies the self-interested incentive model behind Vercel's strategy.

speaking balance: gold is the hosts, purple is the guest (3 minute bins)

0:00 · the hosts 33.7% · guest 66.3%0:00 · the hosts 33.7% · guest 66.3%3:00 · the hosts 0% · guest 100%3:00 · the hosts 0% · guest 100%6:00 · the hosts 42.9% · guest 57.1%6:00 · the hosts 42.9% · guest 57.1%9:00 · the hosts 22.2% · guest 77.8%9:00 · the hosts 22.2% · guest 77.8%12:00 · the hosts 14% · guest 86%12:00 · the hosts 14% · guest 86%15:00 · the hosts 8.7% · guest 91.3%15:00 · the hosts 8.7% · guest 91.3%18:00 · the hosts 28.6% · guest 71.4%18:00 · the hosts 28.6% · guest 71.4%21:00 · the hosts 29.5% · guest 70.5%21:00 · the hosts 29.5% · guest 70.5%24:00 · the hosts 30.7% · guest 69.3%24:00 · the hosts 30.7% · guest 69.3%27:00 · the hosts 5.4% · guest 94.6%27:00 · the hosts 5.4% · guest 94.6%30:00 · the hosts 22.7% · guest 77.3%30:00 · the hosts 22.7% · guest 77.3%33:00 · the hosts 62.5% · guest 37.5%33:00 · the hosts 62.5% · guest 37.5%
Sharpest disagreement ▶ 9:40 Ankur rejects Swyx's premise on offline evals

Ankur firmly tells Swyx he is conflating offline evals with creating static golden datasets in a room, rejecting Swyx's argument that open-endedness makes offline testing obsolete.

Hardest push from the hosts ▶ 22:11 Swyx rejects broadening 'evals' to include vibes

Swyx directly challenges Ankur's attempt to classify vibes under evals, arguing that if vibes count, the definition loses meaning because vendor bias makes everything look like an eval.

Biggest teaching moment ▶ 9:55 Ankur explains modern production log replay over golden datasets

Ankur educates Swyx on how top engineering teams actually run offline evals by pulling live failures from prod logs with one click rather than authoring brittle golden test suites.

The host holds their own ▶ 25:15 Swyx formulates the scaling thesis for synthetic RL environments

Swyx demonstrates domain expertise by framing RL environments in computer use as encoded human intuition designed to break past human throughput bottlenecks.

the scores for every segment, with the reasoning behind each
ChapterTopicThe hosts as informed peerGuest teachingGuest disagreementThe hosts pushing backWhy
Defining Evals, Vibes, and Feedback Loops in AI 2311 Swyx sets the stage by bringing up Boris Cherny's vibe coding comments, after which Malte and Ankur give structured opening philosophies on feedback loops and why vibe checks scale poorly compared to automated evals. The dynamic is collaborative with the host mostly listening.
RL Loops, AI Lab Privilege, and Public Benchmarks 5433 Malte notes that frontier agent teams benefit from unfair lab proximity. Swyx counters with real-world agent bench burnout at Cognition, prompting Ankur to clearly delineate public marketing benchmarks from internal iteration feedback loops.
Offline Evals as Production Replay vs Brittle Unit Tests 4754 Swyx challenges offline evals by claiming open-ended coding outgrows static test suites quickly. Ankur directly pushes back, telling Swyx he is conflating offline running with static golden datasets and explaining modern production replay workflows.
Vercel's Practical Eval Workflow and Composite Models 5544 Malte outlines Vercel's composite model approach to fix trivial syntax errors fast. Swyx tries to summarize it as avoiding agentic loops entirely, which Malte immediately corrects before Swyx cites RL turn optimization from Cognition's SWEGREP work.
Evals as Product Specs and Domain Knowledge Transfer 5446 Ankur describes how evals replace 50-page PRDs for product managers. When Ankur suggests vibes are just another form of eval, Swyx pushes back against diluting the term, arguing that calling everything an eval makes the definition meaningless.
Reinforcement Learning Environments and Synthetic Feedback 6335 Ankur expresses skepticism that average companies can build RL environments without reward hacking. Swyx jumps in with a strong conceptual counterpoint, arguing RL environments scale because they encode domain rules to decouple learning from slow human feedback.
Inversion of Control: Publishing Evals for Model Labs 5533 Malte proposes an inversion of control where framework authors publish evals to force model labs to optimize for them. When Swyx compares this to independent rating agencies like LMSYS, Malte clarifies the self-interested incentive model behind Vercel's strategy.

Statements from this episode (15)

Insight
Ubl: Evals Function to Tell Developers Overnight Whether a Change Is Good
“The way I think about evals is essentially like, it's the thing that, that can tell me tomorrow whether my change is good. And I can operate without that knowledge, but it's super, super helpful.”
Malte Ubl Dec 7, 2025 ▶ 2:28
Insight
Goyal: Offline Evals Are Hardest to Build but Most Efficient Feedback Loop
“Offline evals, which I think is what a lot of the debate was about are both the most challenging to build feedback loop and also the most efficient once built. And then I think AB tests are a little bit less challenging to build and a little bit less efficient…”
Ankur Goyal Dec 7, 2025 ▶ 2:57
Opinion
Goyal: AI Coding Has the Most Product Market Fit After ChatGPT
“The first is it is the use case that has the most product market fit in AI, I think, other than ChatGPT.”
Ankur Goyal Dec 7, 2025 ▶ 3:53
Assertion Contradicted
Ubl: Cognition and Cursor shipped RL fine-tunes of open-source models
“Just yesterday, I think we saw both Cognition, congrats, SWIX, and Cursor to ship RL fine tunes of unnamed open source models.”
Malte Ubl Dec 7, 2025 ▶ 5:22
Opinion
Ubl: Anthropic's Boris Power Vibe-Codes From a Position of Privilege
“I think he comes from a particular position of extreme unusual privilege, which is that he works at an AI lab where like people in the office next door are like writing the evals and are like training the model like every day in exactly that way.”
Malte Ubl Dec 7, 2025 ▶ 5:44
Insight
Goyal: Publishing Public Benchmarks Is Marketing, Not Product Improvement
“It's just that the value proposition of publishing an eval is completely orthogonal to the value proposition of building evals in service of building a good product. I think the purpose of publishing benchmarks is marketing, and it's good marketing.”
Ankur Goyal Dec 7, 2025 ▶ 7:53
Insight
Goyal: Creating Golden Datasets for AI Evals Is Wasted Effort
“People don't really want to create golden data sets. It's, I think it's often a wasted effort to the point that you're making. I think the best teams view offline evals as a mechanism of reconciling what they see in production with real users who are using the…”
Ankur Goyal Dec 7, 2025 ▶ 10:17
Insight
Malte Ubl: When Vibes and Eval Data Disagree, Vibes Are Right
“I think that the common quip that if the vibes and the data disagree, the vibes are probably right. It's true, right? So you have to like, be honest with yourself, like, do they agree and kind of iterate On them over time.”
Malte Ubl Dec 7, 2025 ▶ 12:13
Insight
Goyal: North Star AI Evals Prevent Test Brittleness
“Like if you construct evals in a way that represent the true north star of the problem that you're solving, then they tend not to break as you change the underlying system. Whereas if you hard code them to a narrow subset of like an implementation detail of yo…”
Ankur Goyal Dec 7, 2025 ▶ 13:47
Disclosure
Ubl: Vercel's Composite Models Are Faster Than Agentic Loops
“Basically what we do is we have this like composite model architecture. We run the frontier model and then we run the fine tune model after to fix its errors. That doesn't perform better than an agentic loop, but it's orders of magnitude faster, right?”
Malte Ubl Dec 7, 2025 ▶ 16:51
Assertion Not checkable as stated
Goyal: Braintrust sees surge in PMs and designers joining eval process
“We've seen like a massive surge of product manager, product managers and designers getting interested in participating in the eval process among our customers.”
Ankur Goyal Dec 7, 2025 ▶ 18:24
Insight
Goyal: Providing eval criteria and examples is more effective than writing specs
“In many ways coming to the table of product building with representative examples and criteria that articulate what good versus bad is for a use case is just a more precise and usable form of product management than writing a spec.”
Ankur Goyal Dec 7, 2025 ▶ 18:36
Insight
Goyal: Average AI companies cannot hire expertise to prevent reward hacking
“You need to have like a pretty specific expertise to design the RL environment in a way that's not vulnerable to reward hacking. And I think that either you'll end up with some fixed number of very well engineered RL environments, or you need to somehow employ…”
Ankur Goyal Dec 7, 2025 ▶ 24:35
Disclosure
Ubl: Vercel publishes evals to influence OpenAI and Anthropic models
“I'm Vercel and I publish at Eval. That I want OpenAI and Anthropic to use to make sure when they ship the next model that they're better at the stuff that I care about.”
Malte Ubl Dec 7, 2025 ▶ 28:12
Assertion Not checkable as stated
Goyal: Commercial AI customers are reticent to give eval data to labs
“The interesting thing is that most customers, or actually I'd say a stronger statement, like all customers are quite afraid and reticent to just hand over the data that they use to do evals on to labs.”
Ankur Goyal Dec 7, 2025 ▶ 29:24
Made with StarZero

Turn any episode into a week of clips.

This entire site, over 200 episodes transcribed, diarized, checked and made playable, runs on the StarZero media pipeline. Drop in your own episode and the podcast clipper finds the moments worth sharing, cuts them, captions them, and reframes them for every feed.