Sep 18, 2025 · 1h 7m · lennys-podcast
Why experts writing AI evals is creating the fastest-growing companies in history | Brendan Foody
gold bands on the timeline = statements, start to end. Hover to read, click to jump. CC turns on captions
In this interview, Mercor CEO Brendan Foody joins Lenny Rachitsky to discuss how the emergence of expert-driven AI evaluations propelled Mercor's historic scaling from $1 million to $500 million in revenue run rate. Foody explores the mechanics of reinforcement learning environments, dispels near-term superintelligence alarmism, and outlines actionable frameworks for founders and professionals navigating the AI economy.
How this conversation actually went
Every chapter scored 0–10 on four independent dynamics. Hover any point for the reasoning behind the score. Lenny holds 32.2% of the talking time here. How this is scored →
speaking balance: gold is Lenny, purple is the guest (3 minute bins)
Brendan counters conventional AI alarmist narratives about imminent superintelligence and job obsolescence, arguing that models remain incapable of basic tasks like drafting emails or scheduling calendars.
Hardest push from Lenny ▶ 22:15 Challenging the definition of elastic job demandLenny presses Brendan to clarify what he means by elasticity in the workforce, challenging whether it refers to generalist skillsets or specific high-demand industry capacities.
Biggest teaching moment ▶ 15:20 Educating on the mechanics of RLAIF vs RLHFBrendan systematically breaks down how human-written rubrics replace slow RLHF human rankings, allowing automated reinforcement learning from AI feedback to scale model training.
Lenny holds their own ▶ 1:01:56 Lenny demos his custom voice-assistant hardware buildLenny demonstrates his technical product experimentation by showing off his custom-wired 'Parrot GPT' hardware project mounted inside a stuffed owl.
the scores for every segment, with the reasoning behind each
| Chapter | Topic | Lenny as informed peer | Guest teaching | Guest disagreement | Lenny pushing back | Why |
|---|---|---|---|---|---|---|
| The Era of Evals: Models as Products and PRDs | 3 | 5 | 1 | 1 | Lenny opens by referencing Brendan's pinned tweet and Sarah Guo's quote about evals, positioning evals as a confusing topic for many. Brendan clarifies the concept with an intuitive product framing, explaining that evals act as PRDs and sales collateral for foundation models. | |
| Mercor's Journey: From Bootstrapped Startup to Human Data Frontier | 4 | 4 | 0 | 1 | Lenny frames the rapid ascent of AI data startups and categorizes the landscape into foundational models, vibe-coding apps, and data curation companies. Brendan details Mercor's transition from international generalist staffing to high-end expert sourcing for top AI labs. | |
| How Domain Experts Write Evals and Enable RLAIF | 4 | 6 | 1 | 2 | Lenny asks concrete questions about what domain experts actually do day-to-day and self-identifies as the layperson asking for the audience. Brendan educates Lenny on why the industry is shifting from supervised fine-tuning and RLHF toward RLAIF using rubric-based verifiers. | |
| The Future of Work and the RL Environment Economy | 3 | 5 | 2 | 1 | Lenny asks whether human evaluators will eventually become obsolete, citing a tweet about humans existing solely to generate RL data. Brendan rejects the near-term displacement narrative, asserting that humans will build RL environments for decades as models struggle with basic tool use and long-horizon tasks. | |
| Navigating AI Careers: Elastic Demand and Tool Fluency | 4 | 4 | 1 | 2 | Lenny probes into what students and young professionals should study, asking Brendan to specify which jobs remain elastic. Brendan contrasts low-elasticity fields like accounting with high-elasticity domains like software engineering where higher productivity spurs greater aggregate demand. | |
| Reimagining Global Labor Markets with AI-Powered Matching | 4 | 3 | 1 | 1 | Lenny shares his own ongoing research on how AI has flooded job applications and necessitated automated filtering on the recruiter side. Brendan agrees, explaining why Mercor views itself fundamentally as a labor marketplace rather than a generic data vendor. | |
| Sponsor: Enterpret Customer Intelligence and Voice of Customer | 3 | 5 | 1 | 1 | Following the sponsor read, Lenny relays an anecdote about medical x-ray analysis in ChatGPT to ask whether experts train pre- or post-training data. Brendan educates Lenny on how pre-training ingests broad tokens while post-training experts provide reasoning rubrics and rewards. | |
| Talent Curation: Power Laws, Creative Domains, and Fast Turnaround | 3 | 4 | 0 | 1 | Lenny inquires about compensation rates, project turnaround times, and whether creative writing expertise is valued alongside hard technical domains. Brendan shares metrics, including their $95/hr median pay and hiring comedy writers from the Harvard Lampoon to improve humor in models. | |
| Hypergrowth Drivers: Finding Market Pull and True Product-Market Fit | 3 | 3 | 0 | 0 | Lenny asks how Mercor uncovered hypergrowth demand before raising institutional venture funding. Brendan recounts pitching the founding xAI team while in college and observing incumbents neglect talent quality and payment reliability. | |
| Mercor's Core Values: Can-Do Attitude, High Standards, and Intensity | 3 | 3 | 1 | 1 | Lenny brings up the controversial '996' startup work culture debate, inviting Brendan to explain Mercor's intensity and high standards. Brendan clarifies that Mercor avoids rigid hourly mandates, focusing instead on mission alignment and hiring top tier talent. | |
| Early Entrepreneurship: Donut Dynasty and the Power of Initiative | 2 | 2 | 0 | 0 | Lenny prompts Brendan to share stories from his earlier entrepreneurial projects to extract lessons on founder initiative. Brendan entertains Lenny with the story of running 'Donut Dynasty' in middle school and dodging school restrictions. | |
| Debunking Near-Term Superintelligence and Envisioning AI Abundance | 4 | 4 | 1 | 1 | Lenny references David Sacks' commentary and questions whether model capabilities are plateauing short of superintelligence. Brendan agrees that 3-year AGI predictions are unrealistic, arguing that genuine capability expansion will depend on rigorous, multi-year post-training evals. | |
| AI Corner: Daily Workflows, Thought Partners, and Hardware Experiments | 5 | 2 | 0 | 0 | In AI Corner, Brendan explains how he uses ChatGPT Voice mode as a thought partner, prompting Lenny to showcase a custom wearable hardware project ('Parrot GPT') built into a stuffed owl. Both exchange enthusiastic notes on voice interfaces. | |
| Lightning Round: Dyslexia, Focusing on Strengths, and Media Favorites | 2 | 2 | 0 | 0 | Lenny wraps up with lightning round questions and invites Brendan to discuss managing dyslexia as a high-growth startup CEO. Brendan describes reframing dyslexia as an asset that forces reliance on personal strengths and big-picture pattern recognition. |