Dec 7, 2025 · 34m · latent-space
The Great Evals Debate — Ankur Goyal & Malte Ubl
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this episode of the Latent Space Lightning Pod, Swyx hosts Braintrust founder Ankur Goyal and Vercel CTO Malte Ubl to debate the transition from 'vibe coding' to robust evaluation methodologies for AI coding agents. They delve into production-driven feedback loops, reinforcement learning pipelines, and the strategic role of open-source framework benchmarks.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. The hosts hold 23.2% of the talking time here. How this is scored →
speaking balance: gold is the hosts, purple is the guest (3 minute bins)
Ankur firmly tells Swyx he is conflating offline evals with creating static golden datasets in a room, rejecting Swyx's argument that open-endedness makes offline testing obsolete.
Hardest push from the hosts ▶ 22:11 Swyx rejects broadening 'evals' to include vibesSwyx directly challenges Ankur's attempt to classify vibes under evals, arguing that if vibes count, the definition loses meaning because vendor bias makes everything look like an eval.
Biggest teaching moment ▶ 9:55 Ankur explains modern production log replay over golden datasetsAnkur educates Swyx on how top engineering teams actually run offline evals by pulling live failures from prod logs with one click rather than authoring brittle golden test suites.
The host holds their own ▶ 25:15 Swyx formulates the scaling thesis for synthetic RL environmentsSwyx demonstrates domain expertise by framing RL environments in computer use as encoded human intuition designed to break past human throughput bottlenecks.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | The hosts as informed peer | Guest teaching | Guest disagreement | The hosts pushing back | Why |
|---|---|---|---|---|---|---|
| Defining Evals, Vibes, and Feedback Loops in AI | 2 | 3 | 1 | 1 | Swyx sets the stage by bringing up Boris Cherny's vibe coding comments, after which Malte and Ankur give structured opening philosophies on feedback loops and why vibe checks scale poorly compared to automated evals. The dynamic is collaborative with the host mostly listening. | |
| RL Loops, AI Lab Privilege, and Public Benchmarks | 5 | 4 | 3 | 3 | Malte notes that frontier agent teams benefit from unfair lab proximity. Swyx counters with real-world agent bench burnout at Cognition, prompting Ankur to clearly delineate public marketing benchmarks from internal iteration feedback loops. | |
| Offline Evals as Production Replay vs Brittle Unit Tests | 4 | 7 | 5 | 4 | Swyx challenges offline evals by claiming open-ended coding outgrows static test suites quickly. Ankur directly pushes back, telling Swyx he is conflating offline running with static golden datasets and explaining modern production replay workflows. | |
| Vercel's Practical Eval Workflow and Composite Models | 5 | 5 | 4 | 4 | Malte outlines Vercel's composite model approach to fix trivial syntax errors fast. Swyx tries to summarize it as avoiding agentic loops entirely, which Malte immediately corrects before Swyx cites RL turn optimization from Cognition's SWEGREP work. | |
| Evals as Product Specs and Domain Knowledge Transfer | 5 | 4 | 4 | 6 | Ankur describes how evals replace 50-page PRDs for product managers. When Ankur suggests vibes are just another form of eval, Swyx pushes back against diluting the term, arguing that calling everything an eval makes the definition meaningless. | |
| Reinforcement Learning Environments and Synthetic Feedback | 6 | 3 | 3 | 5 | Ankur expresses skepticism that average companies can build RL environments without reward hacking. Swyx jumps in with a strong conceptual counterpoint, arguing RL environments scale because they encode domain rules to decouple learning from slow human feedback. | |
| Inversion of Control: Publishing Evals for Model Labs | 5 | 5 | 3 | 3 | Malte proposes an inversion of control where framework authors publish evals to force model labs to optimize for them. When Swyx compares this to independent rating agencies like LMSYS, Malte clarifies the self-interested incentive model behind Vercel's strategy. |